Technology2024globalhigh confidence

In an analysis of 24 state-of-the-art language-model benchmarks, Reuel et al. (2024) found that only 4 provided scripts for replicating results, and no more than 10 performed multiple evaluations or reported statistical significance.

Sources