I've been testing both Opus 4.1 and Gemini 2.5 on 100k+ token codebases. Anthropic still wins on instruction following and surgical edits, but Gemini handles repository-wide refactors more consistently. Benchmarks like SWE-bench don't tell the whole story — real reliability dips when the context window fills up with unrelated files. Pricing is also a factor: Opus is 3x the cost, and for routine tasks, Sonnet 4.1 or Flash 2.5 might be the sweet spot. What's your experience? Are you staying with one provider or mixing models per task? I'm also curious if anyone has hit the actual 200k token limit on Gemini without degradation, because my results start to drift after around 150k.
Claude Opus 4.1 vs Gemini 2.5 for long context coding
Replies
No comments yet — be the first to share your thoughts.