I keep seeing models claim SOTA results on benchmarks like MMLU or GSM8K, but when I dig into the technical reports, the instruction-tuning data often overlaps with evaluation sets
Benchmark contamination is making evaluation results meaningless
Replies
No comments yet — be the first to share your thoughts.