The spec sheets look great, but who is really hitting 1M tokens with decent recall? I’ve been testing long-context models for document analysis and the quality degrades fast after maybe 50k. Attention becomes diffuse, and the inference cost gets silly. Are these massive context sizes actually solving a real problem, or is it a spec arms race? For practical workflows, retrieval plus a smaller model still seems to beat stuffing everything into one window. I’d like to hear how others are measuring this. Are there benchmarks that simulate genuinely long dependencies, or is everybody just testing with "needle in a haystack" tasks that don't match real workloads? Curious what people's experience is from production systems.
Are 1M context windows actually useful or just marketing?
Replies
No comments yet — be the first to share your thoughts.