< Back to all clusters
[TECHNOLOGY] · 2 sources

started · updated

Nvidia introduces ACES framework for AI agent skill evaluation

Nvidia has introduced a new research framework called Agentic Continuous Evaluation of Skills (ACES) to address flaws in how AI agent capabilities are currently measured. The company argues that existing industry standards, which often rely on static code checks and LLM-based judging, fail to accurately reflect real-world performance.

According to Nvidia’s research, there is a very low correlation—a Spearman rho of just 0.14—between traditional scan-only metrics and LLM-judge scores. To solve this, ACES utilizes live, head-to-head trials to calculate “Skill Lift,” which measures the quantitative difference in performance when a specific skill is enabled versus when it is not.

The framework employs the Agent Trajectory Interchange Format (ATIF) to allow for standardized comparisons across different agent architectures. Evaluations under ACES are measured across six specific metrics: security, skill execution, skill efficiency, accuracy, goal accuracy, and behavior check.

Entities

Nvidia