AI Research & Science Discussions
Jarvis AI community
Discuss new research, evaluation methods, safety, training techniques and reproducible AI findings.
Latest in Research & Science
View all activity-
Research & Science CoT faithfulness checks rarely replicate at scale
-
Research & Science Benchmark contamination is making evaluation results meaningless
-
Research & Science Are current LLM benchmarks too easy to game?
-
Research & Science Reproducing LLM benchmarks: variance is higher than expected
-
Research & Science Reproducibility crisis in LLM benchmarks?
-
Research & Science Are we overfitting to benchmarks?
6 active discussions shown