Models

Grok 4.6 and DeepSeek V4 Pro Just Redefined the AI Ceiling in One Night.

CRAZE CRAZE Summary 3 things to know
  • DeepSeek V4 Pro nearly matches top model Fable 5 on agent benchmarks but costs 1/57th the price, reshaping developer economics.
  • Grok 4.6 ties GPT-5.6 Sol on general intelligence, tops knowledge work benchmarks, and prices at $6 per million output tokens.
  • DeepSeek's $0.87 per million tokens is temporary; planned price hikes mean the ultra-cheap window is closing fast.
Jeff Editorial | · 5 min read
Grok 4.6 and DeepSeek V4 Pro Just Redefined the AI Ceiling in One Night.

The timing was impossible to ignore. Liang Wenfeng's DeepSeek V4 Pro and Elon Musk's Grok 4.6 arrived on the same evening, both targeting the same thing: long-horizon agent workflows where models call tools, write code, and deliver usable results. Two giants, one battlefield.

The Performance Kill

DeepSeek V4 Pro's official release isn't just an incremental update. The model that scored 72.1 on Terminal-Bench 2.1 three months ago now hits 87.9 — a 15.8-point jump. The global leader, Anthropic's Fable 5, sits at 88.0. The gap is 0.1 points.

The benchmarks tell a clear story:

Benchmark

DeepSeek V4 Pro

Fable 5

Opus 4.8

Terminal-Bench 2.1

87.9

88.0

85.0

CyberGym

83.3

83.1

78.3

AutomationBench

31.8

29.1

27.2

DeepSWE

62.7

70.0

58.0

In cybersecurity agent tests, DeepSeek overtook Fable 5. In automation tasks, it did the same. Even its historically weakest area — software engineering — saw DeepSWE jump from 12.8 in the preview to 62.7, nearly a 5x improvement.

Grok 4.6 is playing a similar game. It scored 61 on the General Intelligence Index — tied with GPT-5.6 Sol, just one point behind Fable 5's 62. On CursorBench, it hit 69.9%, surpassing GPT-5.6's 67.2% and closing in on Fable 5's 70.5%. In knowledge work benchmarks, Grok 4.6 took first place outright — GDPVal-AA v2, AA-Briefcase, and Harvey LAB all went to SpaceXAI.

The Price Kill

Performance matters. But the price gap matters more.

For every million output tokens:

Model

Output Price (per 1M tokens)

Fable 5

$50

GPT-5.6 Sol

$30

Claude Opus 5

$25

Grok 4.6

$6

DeepSeek V4 Pro

$0.87

DeepSeek V4 Pro costs roughly one-fifty-seventh of Fable 5's price. Grok 4.6 costs less than half of GPT-5.6 Sol.

This is what a pricing war looks like at the frontier. Grok 4.6 undercuts OpenAI and Anthropic by roughly half while matching benchmark performance. DeepSeek is playing a different game entirely — delivering competitive performance at a price that makes the incumbents' margins look like a rounding error.

The catch: DeepSeek has already announced it will "significantly raise" API pricing soon. The window for $0.87 token pricing is closing.

Grok 4.6 and DeepSeek V4 Pro Just Redefined the AI Ceiling in One Night.
Grok 4.6 and DeepSeek V4 Pro

The Developer Calculus

When developers choose a model, they balance two numbers: performance and cost. On August 12, both numbers shifted.

Fable 5 is still the best — by 0.1 points on Terminal-Bench, by 1 point on the Intelligence Index. But is it worth 57 times the price? For high-stakes tasks with narrow tolerances, the answer is probably yes. For the other 90% of workloads, the calculus has changed.

DeepSeek's API compatibility adds another layer. It supports both OpenAI and Anthropic API formats, meaning developers can switch from Claude to DeepSeek without rewriting code. The switching cost is effectively zero. The cost savings are not.

Grok 4.6 is pursuing the same distribution play — launching first in Cursor and Grok Build, with double usage credits for the first week. Elon Musk's AI revenue hit $2.56 billion in Q2, up 247% year-over-year. Grok 4.6 is designed to accelerate that trajectory.

What happens next is predictable. OpenAI and Anthropic will respond. They always do. But the direction of travel is clear: performance is compressing, prices are falling, and the moat around the top tier is getting thinner by the day.


P.S. DeepSeek V4 Pro's official pricing is 3 yuan ($0.42) per million input tokens and 6 yuan ($0.87) per million output tokens. The "Flash" version is even cheaper. The company says it will "significantly" raise prices soon. The window to run these models at these prices isn't open forever. But the message has already been delivered: the era of $50 output tokens is ending.


Frequently Asked Questions

Q: Which model performs better in real-world testing—DeepSeek V4 Pro or Grok 4.6?

A: It depends on the use case. DeepSeek V4 Pro excels in agentic cybersecurity and automation benchmarks, overtaking Fable 5 on CyberGym (83.3 vs 83.1) and AutomationBench (31.8 vs 29.1). Grok 4.6 leads in knowledge work benchmarks like GDPVal-AA v2 and AA-Briefcase. For software engineering tasks (DeepSWE), Fable 5 still leads at 70%, followed by Grok 4.6 at 65.9% and DeepSeek at 62.7%.

Q: How significant is the price difference between these models?

A: The gap is unprecedented. For every million output tokens, Fable 5 costs $50, GPT-5.6 Sol costs $30, while Grok 4.6 costs $6 and DeepSeek V4 Pro costs just $0.87. That means DeepSeek is roughly 57 times cheaper than Fable 5. For developers running high-volume workloads, the economics are fundamentally different.

Q: What exactly is Terminal-Bench 2.1 and why does it matter?

A: Terminal-Bench 2.1 is a benchmark that tests AI models on real-world agentic tasks—using a terminal to navigate file systems, write scripts, install packages, and complete multi-step workflows. It's considered one of the most practical measures of an AI's ability to act autonomously in developer environments.

Q: Why did DeepSeek and SpaceXAI release their models on the same day?

A: There's no evidence of coordination. DeepSeek's release timeline was likely driven by its own development cycle and the upcoming API price increase. SpaceXAI's Grok 4.6 release follows the company's rapid iteration pace—the model builds on Grok 4.5's foundation with a longer supplemental training run. The same-day timing appears coincidental, but it sharpens the competitive narrative.

Q: Is DeepSeek V4 Pro available outside China?

A: Yes. The model is available via API globally, with pricing in both yuan and US dollars. DeepSeek supports both OpenAI and Anthropic API formats, making it easy for developers to switch.

Q: What's the catch with DeepSeek's low pricing?

A: The catch is the model's planned price increase. DeepSeek has announced it will "significantly raise" API pricing soon. The current $0.87 per million output tokens rate is a temporary window. The company is offering this aggressive pricing to build market share before raising rates.

Q: Where can I access Grok 4.6?

A: Grok 4.6 is available today in Cursor, Grok Build, and Grok Bot, with third-party API access via OpenRouter, Vercel, and Cloudflare. For the first week, Cursor and Grok Build users receive double usage credits. SpaceXAI's Q2 AI revenue hit $2.56 billion, up 247% year-over-year, and the company is racing to grow its developer ecosystem.

Q: What does this mean for OpenAI and Anthropic?

A: Both companies face increased competitive pressure. OpenAI and Anthropic have historically maintained premium pricing justified by performance leadership. With Grok 4.6 matching GPT-5.6 Sol on benchmarks at half the price, and DeepSeek approaching Fable 5's performance at 1/57th the cost, the premium margin is no longer as defensible. Analysts expect both to respond with pricing adjustments or new model releases in the coming weeks.

Q: How does Grok 4.6 compare to Grok 4.5?

A: Grok 4.6 is built on the same V9 foundation as Grok 4.5 but includes a longer supplemental training run, curated model-generated reasoning data, and aggressive reinforcement learning across agentic tasks. The Intelligence Index score jumped from 56 to 61—a five-point gain—and CursorBench improved from 58.9% to 69.9%, an 11-point leap.

Q: What's coming next from either company?

A: DeepSeek has already signaled a significant API price increase is imminent. SpaceXAI is reportedly preparing Grok 4.7—a larger 2.1-trillion-parameter model—which could arrive within weeks. Grok 5, incorporating SpaceX engineering data, is targeted before the end of 2026.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article