Inherent, a London AI lab founded by Google DeepMind alumni, emerged from stealth weeks ago with a $50 million seed round . On August 21, it released Faraday, an AI agent designed to independently reproduce the findings of published scientific papers .
The benchmark: Faraday runs on Qwen 3.6, a 27-billion-parameter open model . Claude Opus 4.8 is estimated at over 1 trillion parameters. GPT-5.5 is similar. Faraday won .
The company says the win isn't about beating bigger models. It's about how they got there . And that distinction matters more than the benchmark result itself.
27B beats 1T: The math that breaks the scaling argument
The parameter gap is not a small gap. It is a category difference. Inherent's Faraday runs on a model roughly 40 times smaller than the frontier systems it beat .
This is the first time a sub-30B model has publicly outperformed trillion-parameter frontier systems on a complex, multi-step scientific reasoning task. It suggests that the performance gap between "big enough" and "massive" may be shrinking faster than the industry has been willing to admit.
For context: a 27B model can run on a single consumer GPU with quantization. Claude Opus 4.8 requires data center clusters. If Inherent's approach holds up, the cost of frontier-level research agents drops by orders of magnitude.
The reinforcement learning edge: Faraday was trained to have taste, not just accuracy
Inherent's training approach is the real differentiator. The company focused on reinforcement learning rather than teaching Faraday how science is conducted .
Co-founder Edward Hughes described the goal as giving Faraday "research taste"—an instinct for which experiments are worth running and how to design them . The system is rewarded for good outcomes rather than following explicit rules .
This matters because research taste is what separates a helpful lab assistant from a true scientist. The company's long-term goal is to build an "AI scientist"—an agent capable of discovering new knowledge, not just reproducing old results .
Faraday's process mirrors how human scientists work. It took a published paper, ignored the stated result, and ran the experiments independently . This is standard PhD training. The AI did it better than much larger systems.

London's DeepMind diaspora is now producing results
Inherent is one of several startups founded by DeepMind alumni, but it has gotten relatively little attention compared to better-funded rivals . That may be changing.
The company has a dozen employees working in person out of a King's Cross office—the London neighborhood that DeepMind's presence helped turn into a global AI hub . It plans to grow to 20-25 by year-end.
Co-founder Hughes has also called for the end of the UK's "garden leave" restrictions, which prevent departing employees from joining rivals for months after leaving . He said he personally ran into the constraint before starting Inherent .
The company is currently hiring. With Demis Hassabis taking on a new role at DeepMind that has unsettled some staff, Inherent's timing may be strategic .
The open question: What happens when external scrutiny arrives
Inherent's claim comes with caveats . The company has not published a peer-reviewed evaluation. The exact benchmark suite and scoring methodology have not been detailed publicly. Anthropic and OpenAI have not responded to the specific comparison . Faraday also outsources its coding to OpenAI's own GPT-5.5 Codex .
For a $50 million seed-stage lab to claim a win over frontier models is a bold move . Either it holds up under external scrutiny, or it quietly evaporates. If it holds up, the implication is significant: the cost curve for research-grade agents drops sharply, and the moat around trillion-parameter systems narrows to specific capabilities rather than general reasoning .
P.S. Inherent's most provocative design choice is what it didn't build. Faraday doesn't have its own coding tool—it uses OpenAI's GPT-5.5 Codex . That's the opposite of vertical integration. It's a deliberate bet that the agent layer matters more than the tools it uses.
Frequently Asked Questions
Q: What is Inherent?
A: Inherent is a London AI lab founded by four Google DeepMind alumni: Edward Hughes, Louis Kirsch, Kaloyan Aleksiev, and Tantum Collins. The company emerged from stealth in May 2026 with a $50 million seed round led by Index Ventures and Radical Ventures . It currently employs about 12 people and plans to grow to 20-25 by year-end.
Q: What is Faraday, and what does it do?
A: Faraday is Inherent's AI agent designed to independently replicate the findings of published scientific papers without being given the answer in advance . Paper replication is a standard training exercise for human scientists—"Many PhD students actually start by doing this," said co-founder and chief scientist Edward Hughes . The goal is to build toward an "AI scientist" capable of making new discoveries, not just verifying old results.
Q: How did a 27B model beat trillion-parameter systems?
A: Faraday runs on Qwen 3.6, a 27-billion-parameter open model . In comparison, Claude Opus 4.8 and GPT-5.5 are estimated at over 1 trillion parameters. The difference in scale is about 40x. Inherent's claim is that Faraday's reinforcement learning training—which rewards the agent for good outcomes rather than following fixed rules—allowed it to develop "research taste".
Q: What is "research taste"?
A: "Research taste" is Inherent's term for an agent's ability to judge which experiments are worth running, how to design them, and when to pivot . It's an instinct for good science rather than just following procedural rules. The company uses reinforcement learning to train this quality, rather than teaching Faraday how science is conducted step-by-step.
Q: How was Faraday tested?
A: Inherent built a benchmark called Replica, containing 310 tasks drawn from 100 scientific papers across fields like machine learning, structural biology, and materials science . Agents were given limited time and compute to reproduce published figures without seeing the original plots. Faraday outperformed Claude Opus 4.8 and GPT-5.5 on both in-distribution and held-out tasks.
Q: What is the Replica benchmark?
A: Replica is Inherent's own evaluation framework with 310 tasks from 100 papers. It tests not just accuracy but also experimental design, faithfulness to the original research, and resource use. Inherent built an automated judge with a rubric and compared its assessments with human evaluators.
Q: What is the significance of Faraday using GPT-5.5 Codex?
A: Faraday doesn't write its own code—it uses OpenAI's GPT-5.5 Codex as a tool, similar to how human scientists use existing software rather than building everything from scratch . This design choice signals Inherent's bet that the agent layer—the ability to plan, iterate, and decide what experiments to run—matters more than building its own tool stack.
Q: What is Inherent's long-term goal?
A: Inherent's north star is building an AI scientist—an agent capable of discovering new scientific knowledge across multiple fields, not just reproducing existing research . The company is using reinforcement learning to teach Faraday "research taste" as a stepping stone toward that goal .
Q: Who are the founders?
A: Inherent was co-founded by four people: Edward Hughes (co-founder and chief scientist, formerly at Google DeepMind), Louis Kirsch, Kaloyan Aleksiev (formerly at Reka AI and Microsoft), and Tantum Collins (formerly a Biden White House AI policy staffer).
Q: Should these results be trusted?
A: Inherent's claims come with important caveats. The benchmark results are self-reported—Inherent built the Replica benchmark and ran the comparisons internally. No independent third party has verified the claims, and Anthropic and OpenAI have not responded to the specific comparison . However, the company has published its methodology and research page.
Q: What is "garden leave" and why does it matter?
A: Garden leave is a UK employment practice that prevents departing employees from joining rivals for months after resigning . Hughes has been personally affected by it and has called publicly for the practice to end, arguing it gives US startups a hiring advantage that UK labs have to work around.
