Skip to main content

Claude Opus 5 Tops DeepSWE, but Kimi K3 and Gemini 3.6 Flash Are the Real Stories

Anthropic AI data center TPU compute infrastructure
The first independent benchmark that frontier labs cannot train for is here, and Claude Opus 5 sits on top of it. On DeepSWE — a long-horizon software engineering benchmark released July 25, 2026 by Datacurve — Anthropic's flagship scores 74% pass@1, three points ahead of GPT-5.6 Sol and five points ahead of Claude Fable 5. The story gets more interesting lower down: Kimi K3, our current favourite coding model in the AI Arena rig, takes fifth place with 69% at $4.65 per task, and Gemini 3.6 Flash lands a respectable 49% — for less than a third of what Anthropic charges per Opus 5 task.

What is DeepSWE?

DeepSWE is a new benchmark from Datacurve, released on July 25, 2026. It is built to do one job: separate frontier coding models that have started to cluster too closely on existing public benchmarks. The top of SWE-bench Verified, SWE-bench Lite and similar leaderboards has become a narrow score band where adjacent model configurations often overlap on confidence intervals. DeepSWE widens that band.

The benchmark ships with four design choices that, taken together, make it harder to game than its predecessors:

  • Contamination-free. All 113 tasks are written from scratch — they are not adapted from existing pull requests or commits that a model could have seen during pretraining. Every task carries a canary GUID so dataset leakage can be detected.
  • High diversity. Tasks span 91 repositories across 5 programming languages (TypeScript, Go, JavaScript, Rust and others), drawn from real open-source projects such as prometheus/prometheus, capricorn86/happy-dom, yjs/yjs, wasmi-labs/wasmi and beevik/etree.
  • Real-world complexity. Prompts are roughly half the length of SWE-bench Pro's, yet solutions require 5.5× more code and emit about 2× more output tokens on average. This is the part that puts long-horizon agentic behaviour under real load.
  • Reliable verification. Verifiers are hand-written to test behaviour, not implementation details. A model cannot pass by mimicking the original author's style; it has to actually fix the bug.

All 18 models on the leaderboard run on the same open scaffold — mini-SWE-agent — so differences in scores reflect model capability, not agentic-wrapper engineering. Datacurve's launch blog goes into the methodology in detail.

The leaderboard: where the frontier sits today

Out of 18 frontier models tested, Claude Opus 5 (max effort) leads with 74% ± 4% pass@1, ahead of GPT-5.6 Sol (73% ± 3%) and Claude Fable 5 (70% ± 4%). The full field, ranked by score, looks like this — costs in EUR are converted at the approximate USD/EUR rate of 0.92 and rounded to the nearest cent:

DeepSWE v1.1 — Pass@1 by model (effort level in brackets)

#
Model
Score
Avg cost
1
Claude Opus 5 [max]
74%
€10.89
2
GPT-5.6 Sol [max]
73%
€7.72
3
Claude Fable 5 [max]
70%
€19.90
4
GPT-5.6 Terra [max]
70%
€4.55
5
Kimi K3 [max]
69%
€4.28
6
GPT-5.6 Luna [max]
67%
€2.79
7
GPT-5.5 [xhigh]
67%
€6.65
8
Claude Opus 4.8 [max]
59%
€12.16
9
Claude Sonnet 5 [max]
54%
€24.29
10
Grok 4.5 [high]
54%
€2.23
11
Muse Spark 1.1 [xhigh]
53%
€2.17
12
GPT-5.4 [xhigh]
52%
€5.20
13
Gemini 3.6 Flash [high]
49%
€3.25
14
GLM 5.2 [max]
44%
€3.61
15
Gemini 3.5 Flash [medium]
37%
€6.75
16
Kimi K2.7 Code
31%
€2.59
17
Claude Sonnet 4.6 [high]
30%
€5.08
18
Gemini 3.1 Pro [high]
12%
€8.72

Source: deepswe.datacurve.ai, DeepSWE v1.1, 113 tasks, mini-SWE-agent scaffold. EUR converted from USD at 0.92. Effort levels are set by the model vendor.

Read the chart with care. The effort level in brackets matters: Opus 5, Fable 5, Sol, Terra, Kimi K3, Luna and GLM 5.2 all run at max, while Grok 4.5, Muse Spark 1.1, GPT-5.4 and Gemini 3.6 Flash run at high or xhigh. Comparing a max-effort Opus 5 with a high-effort Gemini 3.6 Flash is not a perfectly apples-to-apples comparison — but it is how customers actually buy these models: at the effort level their vendor recommends for production coding work.

Opus 5 wins on accuracy, loses on price

Claude Opus 5 is the most capable coding model on DeepSWE right now. 74% ± 4% pass@1 means it resolved roughly 84 of the 113 tasks on the first attempt. The confidence interval puts it inside shouting distance of GPT-5.6 Sol (73% ± 3%), so the headline gap between Anthropic and OpenAI is not statistically crushing — but Anthropic wins the top spot, and it wins it at $5 per million input tokens / $25 per million output tokens (about €4.60 / €23.00), which was already a notable cut against Fable 5. The average cost per DeepSWE task comes out to $11.84 (€10.89) against an average of 118k output tokens and 99 agent steps.

Where Opus 5 starts to look less elegant is beside Fable 5, its own stablemate. Fable 5 scored 70% at $21.63 (€19.90) per task — almost exactly twice the cost of Opus 5 — for four points less. In Anthropic's own framing, that is the whole point of the Opus 5 release: it does most of what Fable 5 does, at half the price per task on real engineering workloads. DeepSWE is an external confirmation of the same claim we previously saw on Frontier-Bench and CursorBench.

What is genuinely new in the DeepSWE data: Opus 4.8 (the model Opus 5 replaced) only scored 59% ± 2%. So the generational jump from Opus 4.8 to Opus 5, measured on contamination-free long-horizon engineering tasks, is 15 percentage points. That is a real jump — bigger than what SWE-bench Verified could show at the top end, because that leaderboard is already nearly saturated.

The Kimi K3 story — why we keep watching it

For us, the most interesting row is the fifth one. Kimi K3, Moonshot AI's flagship coding model, lands at 69% ± 5% for an average of $4.65 (€4.28) per task. It is statistically tied with Claude Fable 5 (70% ± 4%) — and it costs less than a quarter as much per task. Fable 5 used 119k output tokens and 88 agent steps to reach its 70%; Kimi K3 used 81k output tokens and 98 steps to reach 69%. Kimi spends more steps and emits fewer tokens per step — a pattern we have seen in our own AI Arena measurements on this rig, where K3 prefers to plan in smaller, cheaper chunks rather than generate huge context-burning outputs. This is exactly the kind of behaviour that survives cost-sensitive production deployment.

Kimi K3 is available to European developers via Moonshot's platform.moonshot.ai API and through aggregators such as OpenRouter, with European data residency available through Moonshot's recently announced EU endpoints. The European AI Act's GPAI obligations, which begin applying on August 2, 2026, do not change the API product itself, but they do shift disclosure responsibilities onto Moonshot as a provider — something European buyers should check before signing enterprise contracts.

Gemini 3.6 Flash: the budget pick that actually works

Google's Gemini 3.6 Flash sits 13th with 49% ± 5%, which sounds middling until you read the cost column: $3.53 (€3.25) per task. It is materially cheaper than Opus 5 (€10.89), GPT-5.6 Sol (€7.72), Fable 5 (€19.90) and even Kimi K3 (€4.28). For roughly half the cost of the leading Anthropic model, Gemini 3.6 Flash resolves roughly two-thirds as many tasks. For European engineering teams building internal coding assistants — copilots for code review, bug triage, refactoring — that is a very reasonable trade. Not every line of code in a CI pipeline needs to be written by a 74% frontier model.

Where Gemini 3.6 Flash gets interesting is the comparison with its own predecessors. Gemini 3.5 Flash, scored under the older medium effort setting, only reached 37% at $7.34 (€6.75) per task — worse accuracy and nearly twice the cost of its successor. Gemini 3.1 Pro, an older pro-tier model Google still lists, scored just 12% on DeepSWE at $9.48 (€8.72). The trajectory inside Google's Gemini 3.x generation is unambiguous: each step down in cost — and up in version number — actually improves value. Gemini 3.6 Flash is the cheapest non-trivial model on the leaderboard that breaks 40%, let alone 49%.

For European customers, Gemini 3.6 Flash is available through Google AI Studio and Vertex AI, with EU residencies in europe-west1 (Belgium) and europe-west4 (Netherlands). Pricing is identical in USD and converted to EUR on the invoice for European billing accounts. There is no separate European SKU.

What the test tells us — and what it does not

DeepSWE is a long-horizon engineering benchmark, not a single-turn code completion test. The tasks are open-ended: a model is given a feature request, a bug report, or a refactoring brief, and it has to navigate a real codebase, make edits across files, and produce a diff that passes a behavioural verifier. The average successful run on the leaderboard took about 88 agent steps and emitted between 60k and 280k output tokens. That is roughly an order of magnitude more work than SWE-bench Lite asks, and it is what separates models that _agentic loop well from models that just write one good function.

Three things stand out in the data:

  1. The frontier is still moving. Opus 4.8 to Opus 5 is a 15-point jump. That is not saturation. The same benchmark that has SWE-bench Verified tightly clustered at the top still has room to grow — which means buyers should not over-anchor on any one leaderboard snapshot.
  2. Cost matters more than ever. The cost spread between the top five models is roughly 4.6× — from $4.65 (Kimi K3) to $21.63 (Fable 5). On a 1,000-task backlog, that is the difference between €4,280 and €19,900. For European engineering teams scaling internal coding agents across hundreds of developers, this is the metric that ends up on the CFO's desk.
  3. Open scaffolds level the field. Every model ran on mini-SWE-agent, the same ~100-line open-source scaffold. Wrapper engineering is no longer a moat for labs — Datacurve made sure of that. What you see on the leaderboard is the model, not the agent frame.

What DeepSWE does not tell us: how a model performs inside your own agentic stack, on your own codebase, with your own context-retrieval layer. DeepSWE runs on public open-source repositories; production code is messier. Treat the leaderboard as a prior on which models are worth a paid evaluation, not as a verdict.

From our perspective, the DeepSWE launch confirms what we have been seeing informally in the AI Arena rig: Opus 5 is the current ceiling, Kimi K3 is the model we actually run for cost-sensitive coding work, and Gemini 3.6 Flash is the budget option that no longer embarrasses itself. The 3–4× cost spread between those three is the most actionable number on the whole leaderboard — and it is the one most European teams will want to recompute against their own task distribution before signing a contract.

Is DeepSWE free to run on my own models?

Yes. The benchmark is open: 113 tasks, hand-written verifiers, and the mini-SWE-agent scaffold are all available on GitHub. You pay only for the API tokens your model consumes, which on the leaderboard ranged from about €2.20 to €24 per task depending on the model.

Which model should a European team actually pick for a coding copilot?

If cost is no object and you want the highest accuracy on long-horizon work, Claude Opus 5 (74%, €10.89/task) is the safest pick. If you need a balanced cost/accuracy trade, Kimi K3 (69%, €4.28/task) is the best value in the top five. For high-volume, low-stakes work like triage comments or first-pass refactors, Gemini 3.6 Flash (49%, €3.25/task) is the cheapest model that does not embarrass itself.

Why are some models marked [max] and others [high]?

Those are effort levels set by each vendor. Anthropic, OpenAI and Moonshot submitted their models at max; Google submitted Gemini 3.6 Flash at high; xAI and Meta submitted at high or xhigh. DeepSWE does not normalise across effort levels — it runs models at the effort their vendor recommends for production coding work. So a cross-effort comparison reflects what customers actually buy, not what each model could do at maximum effort.

X

Don't miss out!

Subscribe for the latest news and updates.