Models

Google Priced. OpenAI Sped. Anthropic Hesitated. One Day, Three Strategies.

CRAZE CRAZE Summary 3 things to know
  • Google slashed Gemini 3.7 Flash prices 50% until 2027 to win developers while its promised Pro flagship remains delayed.
  • OpenAI's Ultrafast mode runs GPT-5.6 Sol up to 14x faster on Cerebras chips, eliminating speed-intelligence tradeoffs for latency-critical firms.
  • Anthropic revealed a stronger internal Model 2 but won't release it, positioning trust over shipping products as its strategy.
Jeff Editorial | · 6 min read
Google Priced. OpenAI Sped. Anthropic Hesitated. One Day, Three Strategies.

On August 13, Google, OpenAI, and Anthropic each moved in the AI market within hours of each other. Google cut prices. OpenAI unlocked speed. Anthropic revealed a stronger model it has no plans to release. Three companies, three strategies, one signal: the AI race has moved beyond just who is smarter.

The clustering was not a coincidence. In an industry now averaging a new model roughly every two days, same-day launches signal deliberate competitive positioning, not accidental timing. Each lab chose a different front to fight on.

Google Priced Half Off. The Pro Model Still Isn't Here.

Google released Gemini 3.7 Flash as a direct answer to the question of whether it can still compete. The model is not a flagship; it is a workhorse. It is also a signal.

The numbers show real progress. On DeepSWE v1.1, a benchmark for agentic coding, Gemini 3.7 Flash jumped from 49.0% to 65.3%. FrontierCode 1.1 Main climbed from 34.4% to 43.6%. WebDev Arena Elo rose from 1538 to 1588. AutomationBench, which measures multi-step agent tasks, moved from 17.0% to 30.4%.

But the shift that matters more is pricing. Through the end of 2026, Google is charging $0.75 per million input tokens and $3.75 per million output tokens—half the launch price of Gemini 3.6 Flash. From January 1, 2027, that price doubles to $1.50/$7.50 unless Google extends the promotion.

The offer is a direct appeal to developers: try this model now, at a loss-leader price. The bet is that the gains in coding and agent performance will justify the switch—and that developers will stay when the price returns to normal.

Google is also betting that developers in the Gemini ecosystem stay even if the flagship Pro model remains missing. Gemini 3.5 Pro, promised in June, is still not here. Reuters reported on August 13 that internal testing showed the flagship's coding performance still lagged, and the delay was tied to concerns about keeping pace with rivals. The same day, Google shipped 3.7 Flash. That is not a coincidence; it is a tactical fill.

OpenAI Made Speed a Feature—Not a Trade-Off.

OpenAI's Ultrafast mode is a different bet. It runs GPT-5.6 Sol up to 14 times faster—up to 750 output tokens per second—without sacrificing intelligence. The trade-off between speed and intelligence that developers have accepted for years is, for this tier, gone.

The hardware behind it is the real story. Ultrafast runs on Cerebras' Wafer-Scale Engine chips, which keep 44 GB of model weights entirely on-chip in SRAM instead of shuttling them between memory and compute units. That design eliminates the memory-bandwidth bottleneck that constrains GPU inference.

OpenAI has been planning this for months. In January 2026, it signed a deal with Cerebras to deploy 750 megawatts of accelerators through 2028, valued at over $10 billion. OpenAI infrastructure head Sachin Katti described Ultrafast as an experiment in "what becomes possible when customers can get the intelligence of our most capable models with significantly lower latency."

The benchmark numbers show why speed matters. On Humanity's Last Exam, a 2,500-question set spanning graduate-level chemistry, economics, and literature, GPT-5.6 Sol Ultrafast completed the full set in just over 11 hours. Claude Fable 5 took more than three days. On GDP-Val, a benchmark of economically valuable knowledge-work tasks, Ultrafast delivered a 5.6x end-to-end speedup with no quality loss.

Early customers include Jane Street, Podium, Basis, and Rogo—firms where latency cost is real. Jane Street AI engineer John Crepezzi described the speed as "completely changes the call experience for complex work."

Anthropic Has a Better Model. It Won't Ship It.

Anthropic's move was the least product-like of the three. It did not launch a model. It revealed one—and then said it is not for sale.

Anthropic's latest risk report confirmed the existence of an internal model, codenamed "Model 2," that slightly outperforms the publicly available Mythos 5. On AECI, a composite capability benchmark, Mythos Preview scored 158.91, Mythos 5 scored 161.29, and Model 2 scored 162.79. On CoBench, which measures performance on Anthropic's actual research work, Model 2 hit 62.8%, compared to 50.3% for Mythos Preview. For reference, Anthropic's human researchers score 85% on the same test.

Anthropic is not releasing Model 2. The company also raised its assessment of "misalignment" risk from "very low" to "low"—a move that signals heightened uncertainty without claiming imminent danger.

The more revealing detail is in the methodology. Anthropic acknowledged that some task-based evaluations it uses to measure risk have reached "saturation"—they can no longer distinguish capability gains at the frontier. The tools the company uses to evaluate its own models are themselves running out of resolution. Anthropic described lower confidence in this assessment than in previous reports.

The contrast with Google and OpenAI is sharp. Google is competing on price. OpenAI is competing on speed. Anthropic is competing on safety narrative—and using a model it will not ship to maintain it.

Google Priced. OpenAI Sped. Anthropic Hesitated. One Day, Three Strategies.
One day, three companies, three strategies. AI competition is no longer just about who is smarter.

What It Means

The August 13 cluster is a snapshot of a market in transition. The companies are no longer fighting exclusively over which model is smarter. They are fighting over how developers will use AI in production—whether through low-cost workhorses, high-speed inference, or safety-guaranteed offerings.

Google is betting on volume and developer habit. OpenAI is betting on speed as a feature. Anthropic is betting on trust as a defense. All three are defensible. Only one will prove right.


P.S. The unanswered question across all three launches is this: What happens when all three strategies converge? Google may push Flash prices up in 2027. OpenAI's Ultrafast capacity is still limited, and Anthropic may yet release Model 2. The August 13 cluster was not a conclusion—it was a preview of how this market will be contested for the rest of 2026.


Frequently Asked Questions

Q: What happened on August 13 in the AI industry?

A: Google, OpenAI, and Anthropic all announced product or model news within hours of each other. Google released Gemini 3.7 Flash with aggressive pricing. OpenAI previewed Ultrafast mode—14x speed with no intelligence loss. Anthropic confirmed it has an internal model better than Mythos 5 but won't release it.

Q: Why did all three companies move on the same day?

A: The clustering was not a coincidence. In an industry now averaging a new model roughly every two days, same-day launches signal deliberate competitive positioning, not accidental timing.

Q: What is Google's strategy with Gemini 3.7 Flash?

A: Google is betting on volume and developer habit. It cut prices to half the launch price of Gemini 3.6 Flash through the end of 2026, aiming to get developers to try the model and stay in the Gemini ecosystem. The flagship Pro model is still missing.

Q: What is OpenAI's Ultrafast mode?

A: Ultrafast is a new service tier for GPT-5.6 Sol that delivers up to 750 output tokens per second—14x faster than Standard—without sacrificing intelligence. It runs on Cerebras' Wafer-Scale Engine chips that eliminate the memory bandwidth bottleneck of GPU inference.

Q: What is Anthropic's "Model 2"?

A: Model 2 is an internal Anthropic model that slightly outperforms the publicly available Mythos 5. On CoBench, it scores 62.8% compared to Mythos Preview's 50.3%. Anthropic says it has no plans to release it, citing lower confidence in its risk assessment due to evaluation saturation.

Q: What does "evaluation saturation" mean?

A: Anthropic acknowledged that some task-based evaluations it uses to measure risk have reached "saturation"—they can no longer distinguish capability gains at the frontier. The tools the company uses to evaluate its own models are running out of resolution.

Q: What is Cerebras and why does it matter?

A: Cerebras builds Wafer-Scale Engine chips that keep 44GB of model weights entirely on-chip in SRAM, eliminating the memory bandwidth bottleneck that constrains GPU inference. OpenAI signed a $10+ billion deal with Cerebras in January 2026 to deploy 750 megawatts of accelerators through 2028. Ultrafast is the first product of that partnership.

Q: What does this mean for the AI market?

A: The August 13 cluster signals that the AI race has moved beyond just which model is smarter. Companies are now competing on price (Google), speed (OpenAI), and safety narratives (Anthropic). The market is transitioning from model competition to system competition.

Q: What are the early use cases for Ultrafast?

A: OpenAI's internal teams are using it for incident response and research iteration. Early external customers include Jane Street, Podium, Basis, and Rogo—firms where latency cost is real. Jane Street described the speed as "completely changes the call experience for complex work."

Q: Is Anthropic planning to release Model 2 in the future?

A: Anthropic says it has no current plans to release Model 2, citing safety concerns and lower confidence in its risk assessment due to evaluation saturation. However, the company acknowledges that the decision could change as it develops better evaluation tools.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article