< Back to all clusters
[TECHNOLOGY] · 6 sources

Enterprise AI Agents Fail Without Evaluation Frameworks

Analysts warn that most corporate deployments of AI agents are underperforming because organisations skip the essential step of building evaluation infrastructure. Gartner estimates that over 40 % of agent‑centric AI projects will be cancelled by 2027, citing unclear return on investment and runaway costs as the chief drivers. Companies have distributed tools such as Claude Code and Copilot to large workforces without first defining clear tasks, resulting in agents that generate self‑repeating loops, consume tokens without producing output, and require constant human supervision.

A Forrester study of 287 enterprise roll‑outs found an average 540 % ROI within 18 months for the minority of deployments that reached production, but only 41 % achieved a positive ROI within the first year. The lack of measurable objectives mirrors early railroad expansion, where coordination failures caused costly accidents. Coding remains the sole business function delivering consistent AI value because it possesses an inherent evaluation metric – the code either runs or it does not.

These findings suggest that without a structured evaluation framework—akin to modern OKRs—AI agents are likely to remain an expensive, inefficient layer rather than the productivity catalyst promised by vendors.