started · updated
Grok 4.6 reaches parity with GPT-5.6 Sol in intelligence benchmarks
The release of xAI’s Grok 4.6 marks a significant technical milestone, achieving a score of 61 on the Artificial Analysis Intelligence Index, which ties it with OpenAI’s GPT-5.6 Sol. While the model shows strength in agentic tasks—leading on benchmarks like CursorBench 3.2—it continues to trail competitors in foundational coding performance, such as Terminal-Bench and DeepSWE 1.1.
Despite these performance gains, industry data suggests a growing divide between frontier model capabilities and market usage. While state-of-the-art models are improving rapidly, approximately 84% of tokens on OpenRouter are generated by non-frontier models. These high-volume models provide roughly 77% of frontier performance at a fraction of the cost, highlighting significant price elasticity among developers and enterprises.
As the gap between open-weight models and frontier models closes, market trends indicate that enterprises are increasingly prioritizing cost-efficiency for application deployment, even as they rely on top-tier models for complex software architecture and security design.
Entities
Anthropic · GPT-5.6 Sol · Grok 4.6 · OpenAI · xAI