Skip to main content

Faraday vs GPT-5.5: London's DeepMind alumni claim their AI teammate wins at replicating research

Ilustrační obrázek
In short: London-based Inherent, founded by former Google DeepMind researchers, says its Faraday agent beat OpenAI's GPT-5.5 and Anthropic's Claude at reproducing scientific papers and machine-learning codebases. The startup has $50 million in seed funding to back that claim — and a big reproducibility problem to solve.

There is a dirty secret in AI research that nobody in the marketing department likes to talk about: a large chunk of published machine-learning papers is never fully reproduced. Code is missing, hyperparameters are vague, and re-running GPU-heavy experiments is too expensive for most labs. Independent verification is slow, costly, and ungrateful work. Inherent, a London startup founded by ex-Google DeepMind researchers, wants to change that with a specialised AI agent it calls Faraday.

Faraday is not another chatbot. Inherent describes it as an "AI teammate" designed for a single job: taking a scientific paper and its codebase, then reproducing the results — running experiments, comparing outputs, and flagging discrepancies. This week, the company published benchmark results claiming Faraday outperformed top general-purpose models from OpenAI and Anthropic, including OpenAI's GPT-5.5 flagship, at exactly that task.

What do we actually know about the benchmark?

The honest answer: only what Inherent says. The company reports measured test results showing Faraday beating OpenAI's GPT-5.5 and Anthropic's frontier Claude models at scientific replication tasks. That is better than a demo video, but it is still a vendor-reported benchmark — the kind of thing we at ai-jarvis.eu automatically treat with suspicion until an independent lab runs the same tests in a neutral environment.

There is also a timing problem. The claimed comparison targets GPT-5.5, but OpenAI has since moved to GPT-5.6, and Anthropic released Claude Opus 5 back in July. Inherent's numbers may already be comparing against the previous generation. That does not make the result meaningless — replication is a niche skill, and a specialised agent beating a generalist even one generation older is notable — but it means the "outperforms OpenAI and Anthropic" headline needs a footnote.

Why replication is the real bottleneck in AI science

Scientific reproducibility is in crisis across many fields, and machine learning is no exception. A paper that cannot be reproduced is, in the worst case, expensive noise. For European institutions running Horizon Europe projects, open science and reproducibility are not just ideals — they are contractual expectations. That is where a tool like Faraday could genuinely help: an automated agent that re-runs codebases, checks numerical results, and produces a verification report in hours instead of turning a PhD student's week into a debugging marathon.

None of today's top research-cloud tools does this end-to-end. General-purpose assistants like GPT-5.6, Claude Opus 5, or Gemini 3.7 Flash are excellent at writing code and explaining papers, but they are not built to autonomously execute a full experimental pipeline and verify the outcome. Faraday is the first model we have seen with that specific posture — a specialised "do one job thoroughly" agent rather than a general assistant.

The $50 million question

Inherent has raised $50 million in seed funding (roughly €45 million) to build out this platform. For context, that is a large seed round by European standards — the kind of money that lets a lab hire researchers, rent GPUs, and build the evaluation infrastructure that makes claims like these testable. It also signals that investors see automated research verification as a real market, not a hobby project.

What we do not know yet is the business model. Inherent has not published pricing, API access details, or an EU availability timeline. For European labs, that is the first practical question: is this going to be a cloud API per run, a subscription platform, or something licensed to universities? Until that is announced, Faraday remains a benchmark result, not a product you can integrate into your workflow.

What agentic research actually costs

While Inherent hasn't published Faraday's pricing, we can estimate what agentic research workloads cost at current market rates. A full paper replication is a heavy job: reading the paper, loading the codebase, running training or fine-tuning, evaluating metrics. Easily 5–10 million tokens of input plus meaningful output. At xAI's published Grok 4.6 API prices — $2 per million input tokens and $6 per million output — a single 5-million-token run would cost about $10 in input alone; add one million output tokens and you are at roughly $16 per replication attempt. A budget-tier model like DeepSeek-V4-Flash, at $0.14/$0.28 per million tokens, would do the same for well under a dollar — assuming verification quality holds. This is the cost equation Faraday and its competitors will have to win.

The European angle here is not just about price. London may be outside the EU, but if Inherent sells into the single market, the EU AI Act applies to its models just as it does to anyone else. Chapter V obligations for general-purpose AI models now require technical documentation, copyright compliance policies, and systemic risk evaluations. And Article 50 mandates machine-readable labelling of AI-generated content. If an AI agent autonomously produces a verification report, a research institution deploying it has to know exactly where that content came from and how it was generated. Tools that produce unlabelled scientific output could create compliance headaches for universities, so serious European deployments will demand provenance tracking baked in from day one.

The bottom line

Faraday is the most interesting AI research-tool claim we have seen this month. The AI Arena will be keeping an eye on it: the moment Inherent opens API access, we plan to run our own replication bench against GPT-5.6, Claude Opus 5, and open-weights models like Llama 4 and GLM-5.3. Until then, treat the headlines as a promising vendor claim — not a verified scientific fact.

Is Inherent's Faraday available in Europe?

Not yet. Inherent has announced the benchmark results and the $50 million seed round, but has not published API access, pricing, or an EU rollout timeline. European researchers should expect to wait for the public product announcement.

Could open-weights models do the same job locally?

For simpler replication tasks, yes — a model like DeepSeek-V4-Flash or Llama 4 can run locally at a fraction of the cost. The advantage of a specialised agent like Faraday is the end-to-end orchestration: executing code, comparing results, and producing a verification report without hand-holding.

What does the EU AI Act mean for automated research verification?

In practice, institutions using AI agents to generate research reports must comply with Article 50's transparency rules (labelling AI-generated content) and, for general-purpose models, Chapter V obligations around documentation. Any tool used in academic workflows should make provenance machine-readable.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.