started · updated
South Korean AI foundation model project faces evaluation controversy
The South Korean government’s project to select national AI foundation models, known as the 'Dokpamo' project, is facing scrutiny following the results of its second evaluation phase. While Upstage, SK Telecom, and LG AI Research advanced to the next stage, Motif Technologies was eliminated despite achieving the highest score in the Artificial Analysis Intelligence Index (AAII) benchmark.
Controversy has arisen regarding the transparency of the evaluation process. Critics have pointed out that the Ministry of Science and ICT initially withheld specific scores and rankings, providing only averages and gaps between the top and bottom performers. Although the ministry later released individual category leaders, the total scores and comprehensive rankings for each team remain undisclosed, making it impossible to determine exactly why the top-performing model in global benchmarks was eliminated.
As the project moves toward its final phase, which aims to select two teams early next year, the National Information Society Agency (NIA) is working to expand evaluation criteria. New benchmarks focusing on tool use, agents, and coding are being developed to move beyond simple text-based reasoning. Industry experts are calling for more diverse global indicators and blind testing methods to prevent 'benchmarking'—the practice of optimizing models specifically to score high on known datasets.
Entities
LG AI Research · Ministry of Science and ICT · Motif Technologies · SK Telecom · Upstage