AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Sep 22, 2026, 04:00 JST

xAI's Grok 4.7 scores 46, trails Claude and GPT-6

xAI's Grok 4.7 scores 46, trails Claude and GPT-6

3 Key Points

  1. What happened

    xAI introduced Grok 4.7, its most capable model yet for coding and knowledge work, priced at $2 per million input tokens and $6 per million output tokens, scoring 46 on the Artificial Analysis Intelligence Index (v4.3.2).

  2. Why it matters

    The rates are closer to Chinese models than Western frontier models, probably for good reason, since the model lands mid-pack on the index.

  3. What to watch

    On Terminal-Bench 4.0, Grok 4.7 hits just 26 percent, versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1, so its low price hinges on whether buyers accept the agentic coding gap.

WHO IT HITSProcurement teams and developers choosing between API providers for coding and knowledge-work tasks may treat Grok 4.7 as a lower-cost option, though its mid-pack index score and 26 percent Terminal-Bench 4.0 result suggest it may not be the first pick where agentic coding matters.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Elon Musk's xAI has introduced Grok 4.7, describing it as its most capable model yet for coding and knowledge work. According to the company, it is built on a larger base model, trained with longer reinforcement learning, and designed to better verify its own output. Those are the design claims; the independent numbers tell a more mixed story.

On the Artificial Analysis Intelligence Index (v4.3.2), which combines ten benchmarks, Grok 4.7 scores 46 and lands mid-pack, while Claude Fable 5.1 and GPT-6 lead with 53 each. The gap widens in agentic coding: on Terminal-Bench 4.0, Grok 4.7 hits just 26 percent, versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1, and even the cheaper DeepSeek V4.1 Flash edges past it at 27 percent. The counterweight is price: at $2 per million input tokens and $6 per million output tokens, its rates sit closer to Chinese models than Western frontier models, probably for good reason.

That combination is what the reader may want to weigh. The article does not say whether two of its highest reasoning levels performing about the same reflects a tuning choice or a ceiling. On the evidence given, the value proposition appears to hinge on whether buyers are willing to trade agentic coding performance for a much lower per-token price. The model is available through the Grok API, Cursor, and Grok Build.

FAQ
How does Grok 4.7 compare to Claude and GPT-6?
On the Artificial Analysis Intelligence Index (v4.3.2), Grok 4.7 scores 46, while Claude Fable 5.1 and GPT-6 lead with 53 each. On Terminal-Bench 4.0, Grok 4.7 hits 26 percent versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1.
Where can I use Grok 4.7?
Grok 4.7 is available through the Grok API, Cursor, and Grok Build.
How much does Grok 4.7 cost?
Pricing sits at $2 per million input tokens and $6 per million output tokens.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • AWS debuts Strands Harness, 26% more efficient than peersSiliconANGLE AI · 1h ago
  • Salesforce unveils AIforce at Dreamforce 2026SiliconANGLE AI · 1h ago
  • Cisco takes Splunk AI behind the firewall with Nvidia-powered PODYahoo Finance AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleU.S. pitches "notification mechanism" for AI risks to China ahead of Trump-Xi talks