Image
Claude Opus 5 Tops DeepSWE, but Kimi K3 and Gemini 3.6 Flash Are the Real Stories
The first independent benchmark that frontier labs cannot train for is here, and Claude Opus 5 sits on top of it. On DeepSWE — a long-horizon software engineering benchmark released July 25, 2026 by Datacurve — Anthropic's flagship scores 74% pass@1, three points ahead of GPT-5.6 Sol and five points ahead of Claude Fable 5.