I've been seeing more papers where models overfit to benchmark quirks rather than generalizing. Simple prompt tweaks or even random seed changes cause huge score swings. Also, contamination keeps showing up. Are we measuring real capability or just memorization? I wonder if we need dynamic evaluation suites that shift questions based on prior performance. Would love to hear what the community thinks about adversarial benchmarks or whether static tests are still useful for tracking progress.
Are current LLM benchmarks too easy to game?
Replies
No comments yet — be the first to share your thoughts.