Are massive context windows actually usable? Vendors like Google and OpenAI push beyond 1M tokens, but independent evals show accuracy drops sharply after 100k. My team tested Claude 3.5 Sonnet, GPT-4o, and Gemini 1.5 Pro on long-document reasoning. All three degraded on retrieval tasks as context grew, though Claude held up best. Meanwhile, per-token pricing scales linearly, so you pay a premium for degraded output. Is RAG still the
Context windows: real-world vs benchmarks
Replies
No comments yet — be the first to share your thoughts.