Every new release brags about context length, but in practice recall seems to fall off well before the advertised limit. Needle-in-a-haystack scores look great, yet real documents with similar phrasing, tables, and cross-references get muddled around 100k. I have gone back to chunking even on models claiming 1M. Curious whether others see the same,
Is 1M-token context actually usable?
Replies
No comments yet — be the first to share your thoughts.