Grok 4.6 vs 4.5: Pricing, Evals, Where It Ships
📊 Comparisons Intermediate

Grok 4.6 vs 4.5: Pricing, Evals, Where It Ships

xAI shipped Grok 4.6 on August 12. Official $2/$6 API price, vendor evals vs 4.5, and same-week placement in Build, Cursor, and Copilot.

The AI Dude · August 17, 2026 · 7 min read

A score of 61 on the Artificial Analysis Intelligence Index is the figure xAI's announcement states for Grok 4.6 High, matching GPT-5.6 Sol on the same nine-benchmark composite. The August 12 post puts Grok 4.5 High at 56 on that index, so the vendor-to-vendor step is five points, and the comparison the post leads with is a tie against Sol rather than a win over it.

xAI's announcement states: "Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work." Same-day placement in Cursor and Grok Build, double included usage for the first week, API access plus OpenRouter, Vercel, and Cloudflare, and a list price that starts at $2 per million input tokens and $6 per million output tokens. Two days later, GitHub's changelog says Grok 4.6 is rolling out in GitHub Copilot.

61 versus 56 on the AA Index

The 61 comes from xAI's own evals table, listed under Grok 4.6 High versus Grok 4.5 High. xAI's announcement states that the Artificial Analysis Intelligence Index is a composite score of nine benchmarks, and that Grok 4.6 matches GPT-5.6 Sol on it. Fable 5 Max sits at 62 on the same row, one point above both.

Treat the nine-task blend as a ranking device. It tells you how the vendor's High setting of 4.6 sits against the vendor's High setting of 4.5 on a blended intelligence score, and it tells you the lab wants 4.6 read as Sol-class on that blend. Which of the nine tasks moved, how High was configured, and how a third-party harness would score the same checkpoint are separate questions from the composite. Plan around 61 if your work looks like a blended intelligence exam. Plan around the agent rows if your work looks like a long coding session that has to finish.

xAI's announcement states that competitor figures in the table are drawn from the respective developers' published system cards or benchmark leaderboards, and that third-party model scores are the best of self-reported or publicly available results. That framing matters for the Sol tie. Both 61s sit on the same row, but they are not guaranteed to be the same harness, the same attempt budget, or the same date of run. Use the five-point lift versus Grok 4.5 High as the cleaner within-lab step, because those two columns share a producer.

65.9% versus 54% on DeepSWE

xAI's announcement states the following vendor scores for Grok 4.6 High against Grok 4.5 High.

EvalGrok 4.6 HighGrok 4.5 High
AA Intelligence Index6156
GDPVal-AA v217531526
CursorBench v3.269.9%66.7%
DeepSWE v1.165.9%54%
FrontierCode v1.1 Extended61.3%56.6%
APEX-Agents57.5%47.1%
Terminal-Bench v3.026%15.7%

The largest moves sit on the agent-shaped boards. DeepSWE v1.1 jumps from 54% to 65.9%. Terminal-Bench v3.0 jumps from 15.7% to 26%. APEX-Agents jumps from 47.1% to 57.5%. CursorBench v3.2 and FrontierCode v1.1 Extended move less, 3.2 and 4.7 points respectively. GDPVal-AA v2, a knowledge-work score, moves from 1526 to 1753.

Those deltas line up with the sentence the August 12 post asks you to take as the product claim. Long-running agents and multi-step coding are the work the table says improved most. Terminal-Bench remaining at 26% still describes a hard suite, even after the jump from 15.7%. DeepSWE is the row that most closely matches a hours-long software engineering trajectory. If your loop looks like that, the 11.9-point DeepSWE move is the figure that should sit next to 61, not underneath it.

xAI's announcement states that Grok 4.6 underwent a longer supplemental training run than Grok 4.5, then used Grok 4.5 to regenerate SFT trajectories across reasoning efforts and agent harnesses, and trained on agentic RL tasks that include knowledge work, general coding, and domain-specific environments. Keep that training story next to the agent rows. Keep 61 next to the index. Separate the two when you decide whether to switch a coding agent off 4.5.

$2 and $6 under 200k tokens

xAI's announcement states that pricing starts at $2 per million input tokens and $6 per million output tokens, and that a fast variant is twice that price. xAI's documentation lists the same $2 / $6 pair for prompts under 200k tokens, then $4 / $12 once the prompt reaches 200k, on a 500k context window. Cached input is $0.50 per million below the step and $1.00 at or above it. The long-context rate applies to every token in the request once the prompt crosses the threshold, not only to the tokens past 200k.

Grok 4.5 on the same sheet uses the same $2 / $6 and $4 / $12 list rates. Cached input is where 4.6 costs more: 4.5 cached input is $0.30 / $0.60. A workload that hits the cache hard will pay more per cached token on 4.6 than it did on 4.5, at the same uncached list price. A long agent trace that stays under 200k prompt tokens pays the launch headline. A trace that piles tools, files, and prior steps past 200k pays the doubled sheet for the whole call.

xAI's documentation lists February 1, 2026 as the knowledge cutoff for Grok 4.6, and it points code and chat traffic at this model. For anyone wiring the Grok API, the model id is grok-4.6. Our earlier Grok 4.5 versus GPT-5.6 Sol comparison still describes the prior generation's scoreboard. A 4.6 agent loop bills at $2 / $6 until the prompt crosses 200k, then at the doubled long-context rate for the whole request.

2x included usage in week one

xAI's announcement states that Grok 4.6 is available the same day in Cursor and Grok Build, with 2x included usage in both products for the first week. The Grok Build page now says the CLI is powered by Grok 4.6. Install is a single line:

curl -fsSL https://x.ai/cli/install.sh | bash

The API opened the same day, and the announcement names OpenRouter, Vercel, and Cloudflare as partner routes. If you already pay for Cursor, the model should appear in the picker without a new contract. If you already call Grok through the API, switch the model id. Treat the doubled included usage as a launch-week allowance, not a new price tier. It expires when the first week ends.

Same-week placement is the part of this release that is easy to miss if you only read the 61. Grok 4.6 is already in the two coding surfaces xAI sells, on the public API, and on three partner gateways, with a temporary usage bump that makes a side-by-side against 4.5 cheap to run on your own repo.

8 Copilot pickers, one policy flag

GitHub's changelog says Grok 4.6 began rolling out in Copilot on August 14 for Pro, Pro+, Max, Business, and Enterprise, and puts the fit in one line: "It is designed for agentic coding and complex multi-step workflows." The changelog lists the model picker in Visual Studio Code, Visual Studio, Copilot CLI, the cloud agent, the Copilot app, JetBrains, Xcode, and Eclipse. Rollout is gradual, so a missing row in the picker today can appear later in the week.

Copilot Enterprise and Business administrators have to enable the Grok 4.6 policy in Copilot settings. The policy ships off by default, so an org that never flips it will not see the model no matter which IDE the developer has open. Usage is billed at provider list pricing under usage-based billing, which is the $2 / $6 (or $4 / $12) sheet above, not a separate Copilot sticker.

To turn 4.6 on this week, install the Grok Build CLI with the curl line, pick Grok 4.6 in Cursor, set the API model id to grok-4.6, and, if you are on Copilot Business or Enterprise, ask an admin to flip the Grok 4.6 policy before you look for it in the picker. Then run one long-running coding task through it while the extra included usage is still live.

Grok 4.6Grok 4.5xAIAPI pricingGitHub Copilot
Share 𝕏 / Twitter Reddit LinkedIn

Keep reading

Weekly issue

The 5 AI tools that mattered this week.

One email, Fridays. No spam, unsubscribe anytime.