A London startup with a dozen employees just published benchmark results showing its 27-billion-parameter AI agent outperformed Claude Opus 4.8 and GPT-5.5 at independently replicating published scientific research — and did it on a base model that is a fraction of the size of either competitor. Inherent, founded by alumni of Google DeepMind and operating out of King's Cross, made the claim public on August 22, 2026, weeks after closing a $50 million seed round and stepping out of stealth.
The Claim
The agent at the center of the announcement is called Faraday. Its task was paper replication: given a published scientific paper, reproduce the experimental findings independently, without access to the original code or prior knowledge of the correct answer. Inherent says Faraday beat both OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.8 on this benchmark. Multiple reports, including coverage from TechCrunch and a summary published by Zetik, confirm those are the two specific models Inherent cited as its comparison points.
What makes the result striking — and the reason it generated attention beyond the usual startup benchmark announcement — is the underlying architecture. Faraday runs on Qwen 3.6, an open model from Alibaba with 27 billion parameters. The models it beat are frontier-scale systems from two of the best-resourced AI labs in the world. Parameter count is an imperfect proxy for capability, but the gap here is large enough that raw scale alone can't explain the win.
Cofounder and chief scientist Edward Hughes has been specific about what Inherent thinks actually drove the result. The company trained Faraday using reinforcement learning — an approach that rewards successful outcomes rather than scripting the agent's behavior through fixed rules. The target wasn't just accuracy on a given step; it was what Hughes describes as "research taste," an instinct for choosing which experiments are worth running and how to structure them. That's a judgment signal rather than a factual one, and it's a meaningfully different training objective than what most fine-tuning pipelines optimize for.
One design choice worth noting: Faraday doesn't try to do everything itself. For programming tasks, it calls on OpenAI's GPT-5.5 Codex rather than building a proprietary code executor. A separate source, Rīga TV, confirms this detail — Inherent's position is that composing specialist tools mirrors how working scientists actually operate, using existing software rather than building from scratch. The company's stated ambition is not to build a model that replaces all tools but a collaborative agent that behaves like a proactive scientific teammate.
What We See
The counter-intuitive headline — small model beats big models — is real, but it's worth being precise about what "beats" means here. The benchmark is paper replication, which is a narrow and specific task. Inherent is not claiming Faraday is a better general-purpose AI than Opus 4.8 or GPT-5.5. The comparison is scoped: on this task, with this training approach, applied to this evaluation. That framing matters because it changes what conclusion developers should draw.
Our read is that the more interesting claim isn't the benchmark win itself; it's the training methodology. Reinforcement learning that rewards experimental outcomes — as opposed to instruction-following or next-token prediction on curated data — is a plausible path to building agents that generalize to novel scientific problems. Hughes has been clear that paper replication is intended as a scaffold, the same way PhD students build skill by reproducing prior work before attempting original contributions. If the RL signal that produces replication skill also produces something like scientific judgment, that's a genuinely different capability than what current frontier models offer out of the box.
The choice to benchmark against Claude Opus 4.8 and GPT-5.5 specifically also carries implicit context. Both are among the most capable publicly available models as of August 2026. The fact that a 27B-parameter fine-tune of Qwen 3.6 can compete on a defined agentic task supports something developers and researchers have suspected for a while: that fine-tuning approach and task scope matter more than scale once you're past a certain capability threshold. Qwen 3.6 is not an obscure base model — it already appears in routing configurations at organizations doing agentic work precisely because of its efficiency profile. This result is independent evidence that it can handle well-scoped agentic tasks, not just chat or retrieval.
A separate analysis from careeraheadonline.com echoes the same central finding without adding new technical detail, but its framing is useful: it positions this as a signal that the frontier of AI research methodology is shifting, not just that one model beat another in one test. That's a fair read of what Inherent is trying to demonstrate.
Where It Falls Short
The benchmark itself hasn't been independently verified. Inherent released its own results, which is standard practice for a company announcing a product — but it also means the comparison hasn't been validated by a third party with access to the full evaluation protocol. We don't know the size or diversity of the paper replication test set, which scientific domains were included, or whether the same prompting strategy was applied consistently across all three models. Any of those variables could materially affect the outcome. Until the methodology is public and reproducible, the specific performance gap should be treated as a vendor claim, not a confirmed benchmark.
There's also the question of what happens next. Inherent has 12 employees. The Zetik summary notes the company plans to scale to 20–25 people by end of year, which remains a small team relative to what it's taking on — building AI capable of original scientific discovery, not just replication. Paper replication is a tractable milestone. The step from reproducing known results to generating new scientific knowledge is an entirely different problem, and no current system has demonstrated that reliably. Inherent hasn't claimed it has, but the company's roadmap depends on the RL signal that works for replication actually generalizing beyond it. That's still an open empirical question.
Sources
techcrunch.com Inherent Says 27 Billion-Parameter Faraday Beats OpenAI, Anthropic in Paper Replication | Zetik Inherent's AI Surpasses OpenAI and Anthropic in Research Rep London AI startup Inherent says its Faraday agent beat Anthropic and OpenAI at replicating research — Rīga TVBased on
https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research/— techcrunch.comThis article is an original, AI-assisted summary and analysis. Credit for the underlying reporting or footage belongs to the source above.

Written by the vybecoding.ai editorial team
Published on August 22, 2026