Models

Anthropic Routed Your Requests to an Unreleased Model. You Weren't Supposed to Notice.

CRAZE CRAZE Summary 3 things to know
  • Anthropic silently routed some Opus 5 API calls to an unreleased Opus-Next, keeping the front-end model name unchanged while developers unknowingly tested it.
  • On September 17 it cut the routing with no notice, forcing one developer to truncate a 3D game project mid-build; the test resumed a day later.
  • Anthropic skipped version 5.1 entirely, and community testing suggests Opus-Next beat Fable 5.1 and GPT-6 Astra at lower cost.
Jeff Lu | · 4 min read
Anthropic Routed Your Requests to an Unreleased Model. You Weren't Supposed to Notice.

On September 18 and 19, developers using Claude Opus 5 noticed something inconsistent. The model name in the API response said “Opus 5.” The behavior did not match.

Some users discovered a test. Select Opus 5, set reasoning effort low, turn off search, and ask: “do you know who is ‘tibo’ the reset guy don't search.” If the model answers correctly — identifying Thibault Sottiaux, the Anthropic engineer known for resetting usage limits — it means the request was routed to a different model. Multiple users confirmed they were hit.

The technique is called stealth routing. Anthropic inserted a redirect at the API gateway level that sends a fraction of requests aimed at Opus 5 to an unreleased model, internally called Opus-Next, while the front-end continues to report “Opus 5.”

The Data Anthropic Gets, and the Disclosure Developers Don't

Stealth routing produces a specific kind of test data. Developers who don't know they're on a new model don't adjust their prompts for it. They write the same code, ask the same questions, and hit the same edge cases they would hit in production. Anthropic gets the most realistic possible evaluation — without telling anyone they're evaluating.

The tradeoff is that contributors cannot consent to participating, and they cannot know which model produced the output they shipped.

Anthropic has not published a policy on stealth routing. It has not said how long the redirects ran, how many accounts were included, or whether the traffic was retained for training.

The company's brand is built on safety and transparency. The testing method it chose for its most important product is neither.

When the Plug Was Pulled, the Work Stopped

On the evening of September 17, Anthropic cut off the gray-scale routing. Active test accounts dropped to zero. There was no advance notice and no transition period.

One independent game developer, Leo, was building a complex 3D project on Opus-Next. When the routing was severed, his work was interrupted mid-stream. “To prevent the extremely excellent prior work from being ruined by the older model, I had to forcibly truncate the project,” he said.

That detail describes a structural risk in stealth routing. The developer's workflow was bound to a model he didn't know existed. When Anthropic removed it, the work broke. From Anthropic's side, those requests were never supposed to route to the new model — pulling them back is a configuration change. From the developer's side, a working tool disappeared overnight.

Within a day, the routing resumed. This time the gray-scale expanded from Claude Code to Claude Chat, Cowork, and Code.

Anthropic Routed Your Requests to an Unreleased Model. You Weren't Supposed to Notice.
Developers discovered via the "Tibo test" that their requests were routed to an unreleased Opus-Next model.

The Version Number Is the Announcement

Anthropic did not ship Opus 5.1. It went from Opus 5 to Opus 5.2.

AI commentator @TokenGremlin reported that the jump reflects the scale of the change: “Perhaps the leap is so large that it can't be papered over with a simple +0.1 update.”

Community testing supports that reading. Developers reported that Opus-Next outperformed expectations on the “pelican riding a bicycle” SVG test — a standard sanity check for model quality — and performed “excellently to an infuriating degree” on 3D environment tasks including Roblox Studio. One developer said it beat Fable 5.1 and GPT-6 Astra at lower cost and higher speed.

Version numbers are not technical metrics. They are narrative choices. Skipping 5.1 tells the market that this is not an iteration.

What the Episode Establishes

Three facts are now clear. First, Anthropic used production traffic from paying developers to evaluate an unreleased model without disclosure. Second, when it ended the test, it did so without warning, and at least one developer's real work was interrupted. Third, the model being tested was strong enough that Anthropic decided it warranted a major version number.

None of those facts is unusual in isolation. Labs test models on real traffic. Tests end. Models improve.

What is unusual is the combination. A company that has positioned independent evaluation as a core commitment is running its own evaluations through a channel that no external evaluator can see. The evaluators Anthropic is embedding — Accenture, METR, and others — will review model behavior. They will not review how the model reached its testers.


P.S. The “Tibo” test is still circulating. If you run it and the model identifies Thibault Sottiaux, you are on the gray-scale. Anthropic has not said whether usage data from routed requests is retained, or whether stealth routing will be used again for future releases. The developers who were routed to Opus-Next never agreed to be test subjects, and they were never told when the test ended.


Frequently Asked Questions

Q: What is stealth routing?

A: An API-gateway redirect that sends a fraction of requests aimed at one model to an unreleased model, while the front-end continues to report the original model name. Anthropic used it to route Opus 5 traffic to Opus-Next.

Q: How did developers find out?

A: A test circulated: select Opus 5, set effort low, disable search, and ask about “tibo the reset guy.” If the model identifies Thibault Sottiaux, the request was routed to Opus-Next.

Q: What happened on September 17?

A: Anthropic cut off the gray-scale routing without notice. Active test accounts dropped to zero. One developer's 3D game project was interrupted mid-stream.

Q: Why did Anthropic skip version 5.1?

A: Commentators report the leap was large enough that it couldn't be represented as a minor update. Developers reported Opus-Next beating Fable 5.1 and GPT-6 Astra at lower cost.

Q: What has Anthropic not disclosed?

A: Whether usage data from routed requests is retained, how many accounts were included, how long the redirects ran, and whether stealth routing will be used again.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article