I've been digging into recent papers on LLM evaluation and noticing a worrying trend: benchmark scores often shift significantly just by changing the random seed or the order of examples. Some groups report that
Reproducibility crisis in LLM benchmarks?
Replies
No comments yet — be the first to share your thoughts.