In an analysis of 24 state-of-the-art language-model benchmarks, Reuel et al. (2024) found that only 4 provided scripts for replicating results, and no more than 10 performed multiple evaluations or reported statistical significance.
Sources
- Can We Trust AI Benchmarks? An Interdisciplinary Review ... (arxiv.org)
- https://proceedings.neurips.cc/paper_files/paper/2024/file/26889e8359e7ef8a7f5d77457364ca55-Paper-Datasets_and_Benchmarks_Track.pdf (proceedings.neurips.cc)
- https://www.researchgate.net/publication/386014536_BetterBench_Assessing_AI_Benchmarks_Uncovering_Issues_and_Establishing_Best_Practices (researchgate.net)