# Grok 4.6 hit the Pareto frontier. Don't marry a model provider.

> Grok 4.6 topped CursorBench 3.2 at $2.81 per task while the runner-up charged $17.32. The gains are real, and you only capture them if your stack lets you point at a model you did not buy.

Published: 2026-08-13. Last updated: 2026-08-13.

Although Grok started out as an intentionally raunchy alternative to Claude, ChatGPT, and Gemini, Elon and the team at SpaceXAI have turned it into a model at the Pareto frontier.

That is a strange sentence to write about a product whose early differentiator was an unhinged mode. But the benchmarks are the benchmarks.

The Pareto frontier represents the set of optimal choices where you cannot improve one dimension without sacrificing another. Plot every model on a chart with cost on one axis and capability on the other, and most of them land somewhere in the middle of the cloud. There is always something cheaper at the same quality, or better at the same price. The frontier is the outer edge of that cloud, the short list where that is no longer true. To beat a model on the frontier you have to give something up.

Most models never get there. **Grok 4.6 got there by being the cheap one.**

## The cheap one won

[Grok 4.6 was released yesterday](https://x.ai/news/grok-4-6) with the claim that it now sits on that frontier, offering the best combination of cost and intelligence in its class. It jumped past Fable and Sol on [CursorBench](https://cursor.com/cursorbench), which is Cursor's own coding benchmark, built from ambiguous multi-file tasks pulled out of real sessions rather than synthetic puzzles.

The top of that leaderboard is worth sitting with:

- **Grok 4.6** at 70.8%, averaging **$2.81 per task**
- **Fable 5 Max** at 70.5%, averaging **$17.32 per task**
- **Opus 5 Max** at 70.0%, averaging **$8.23 per task**
- **Grok 4.6** at high effort, 69.9%, averaging **$2.34 per task**

Four entries. Nine tenths of a percentage point between the best and the worst of them. And a **7.4x spread in what they cost to run**.

That is not how leaderboards normally look. Usually the expensive model is expensive because it is winning, and the argument is about whether the last few points are worth the premium. Here the model at the top is also nearly the cheapest thing on the board. There is no premium to debate. You are just paying six times more for a slightly lower score.

<Callout type="note">
  Cursor doesn't quote headline API rates. It applies each model's published
  per-million-token pricing to the tokens that model actually burned on each
  task. A model that thrashes gets billed for thrashing.
</Callout>

## The savings aren't a discount, they're fewer steps

That methodology note is the part I find genuinely interesting, and it is buried under the leaderboard where nobody reads it.

If cost per task were just the sticker price, this would be a boring story about xAI undercutting Anthropic on rate card. It isn't. Grok 4.6 finishes the average task in **39 steps and roughly 32,000 tokens**. Opus 5 Max takes 78 steps. Fable 5 Max burns over 100,000 tokens to land six tenths of a point higher.

So the gap is not a discount. The model needs less work to do the work. It flails less, backtracks less, and re-reads the codebase fewer times. On a single task that is a rounding error. Across a team running agents all day, it is most of your bill.

This also quietly reframes what "fast" means. Grok 4.6 is **not** fast in the way people usually mean it. It takes about 32 seconds to produce a first token, and it generates around 86 tokens per second, which is unremarkable. But it gets to done in half the moves. Time to first token is the metric you feel in a chat window. Steps to completion is the metric you pay for in an agent.

## Where the hype needs a haircut

I want to be careful not to oversell this, because the launch coverage already has.

[Artificial Analysis](https://artificialanalysis.ai/models/grok-4-6) runs a broader intelligence index across reasoning, math, and knowledge work rather than coding alone. On theirs, Grok 4.6 comes in **tied for third at 61**, behind Claude Opus 5 at 63 and Fable 5 at 62. Elon called the model "objectively #1." On one benchmark that holds up. On the other it plainly does not.

It is also worth saying out loud that CursorBench is Cursor's benchmark, measuring the kind of work Cursor sells. That is not a knock on the methodology, which is more honest than most, but a coding benchmark is a coding benchmark. If your agents spend their day reconciling invoices rather than refactoring React, the numbers that matter to you may look different.

And Grok is not the cheapest model available. Gemini 3.6 Flash will run you $1.56 a task. It also scores 53.5%, which is a different product for a different job. **The claim worth defending is that Grok 4.6 is the cheapest thing at the frontier, not the cheapest thing.**

## The benchmarks disagree, and that's the argument

Here is the part I keep turning over.

Two credible, well-run benchmarks looked at the same model in the same week and came back with "first place" and "tied for third." Neither is wrong. They are measuring different work.

The instinct is to treat that as a problem to resolve, to find the real ranking. I think that instinct is the mistake. **If the ranking flips depending on which benchmark you trust and what you happen to be doing that day, then picking the right lab was never a strategy.** Not having to pick is the strategy.

That is a much less satisfying conclusion than "use the best model." It is also the only one that survives contact with a leaderboard that reshuffles monthly.

## Intelligence is commoditizing faster than contracts can

This is more evidence that the stream of intelligence is becoming completely commoditized. Competition and innovation are driving prices down while making it easier and easier to switch between models.

Look at the clock on this. Grok 4.6 landed roughly a month after 4.5. It gained five points on the Artificial Analysis index and **did not raise its price**. A 2.1 trillion parameter Grok 4.7 is already announced for a few weeks out. Whatever any leaderboard says today has a shelf life measured in weeks.

Now compare that to the shape of an enterprise AI agreement. Twelve months, sometimes twenty-four, negotiated seats, committed spend. The technology is repricing every few weeks and the contracts are annual. Those two clocks are not close to synchronized, and the gap between them is where your money goes.

Which is why I cannot emphasize this enough: if you marry a single model provider and commit exclusively to its models and products, you will miss out on these gains in cheaper, more capable intelligence.

## Your ceiling is set by someone else's roadmap

ChatGPT serves OpenAI models. Claude serves Anthropic models. That is not a flaw, it is the definition of the product, and both are excellent at what they do.

But it does mean the ceiling on the intelligence available to you is set by one lab's roadmap and one lab's release schedule. Everything above is a thing that happened to other people if your work lives inside a first-party app.

The half-measures deserve more scrutiny than they get. Microsoft 365 Copilot is the flagship "we support model choice" product and genuinely earned credit for [breaking OpenAI exclusivity](https://www.microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot/). Then look at the actual picker. [Five options](https://learn.microsoft.com/en-us/microsoft-365/copilot/cowork/cowork-models): Auto, Claude Opus 4.8, Claude Sonnet 4.6, an advisor mode, and GPT 5.5. No Grok. And it still lists Opus 4.8 rather than Opus 5.

Model choice, as a feature, means a shortlist someone else curates on someone else's schedule. **That is a different thing from being able to point at anything the day it ships**, and the distinction only shows up on days like yesterday.

OpenAI, Google, and Anthropic have no incentive to give you access to better-value agentic options once they have locked you into an annual enterprise contract. Nobody at any of those companies is going to add a competitor to your dropdown the week that competitor starts winning.

## Own the platform, rent the intelligence

Owning your platform and use cases, while treating models as interchangeable infrastructure, will pay off in the long run.

For what it is worth, we route a lot of our own work to Grok 4.6. It is the primary model in two of our three automatic routing tiers, it was routable here the day it shipped, and selections that were pinned to Grok 4.5 upgraded themselves. Nobody migrated anything. When 4.7 lands, that will be a config change rather than a project.

None of that is a bet on Grok. It is a bet against having to bet. Three months ago the same slot was filled by a different model, and three months from now it probably will be again. The routing survives the churn because the model is a setting rather than a foundation.

**The durable advantage was never picking the right model. It is being able to change your mind cheaply.**

Can't wait to see this trend continue.