Superhuman/X ArchiveView on X
Artificial Analysis

@ArtificialAnlys

SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost

Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic.

Key takeaways:
➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3

➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on 𝜏³-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models

➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier

➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max)

Other model details:
➤ Context window of 500k tokens (unchanged from Grok 4.5)

➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5’s $0.3 per 1M tokens for cache hits

Congratulations to @SpaceXAI and @elonmusk on the release!
Image from the post
1473413.2K412
Artificial Analysis

@ArtificialAnlys

Grok 4.6 performs strongly on agentic tasks including knowledge work, terminal use, and customer service. Across GDPval-AA v2, 𝜏³-Banking, and Terminal-Bench v2.1, Grok 4.6 sits in line with or ahead of all models except Claude Opus 5, and sits on the cost vs. performance Pareto frontier across the agentic evaluations in the Intelligence Index.
Image from the post
31723615
Artificial Analysis

@ArtificialAnlys

Grok 4.6 is highly cost-effective: with headline pricing unchanged from Grok 4.5, it delivers a 5 point Intelligence Index gain at a cost per task comparable to Kimi K3 and far below Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5. This places Grok 4.6 firmly on the Intelligence vs. Cost per Task Pareto frontier.
Image from the post
71721919
Artificial Analysis

@ArtificialAnlys

Grok 4.6 debuts on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks. It achieves an Elo of 1577, behind Claude Opus 5. Grok 4.6 performs consistently strongly across rubric grading, presentation, and analytical quality.
Image from the post
151227
Artificial Analysis

@ArtificialAnlys

Full Intelligence Index evaluations breakdown:
Image from the post
1612720
End of thread