Koch et al. (2021) found that more than 70% of benchmark datasets used in prominent computer-vision papers had been reused from other domains.
Notes on verification
Directly confirmed by the peer-reviewed NeurIPS 2021 paper (71.9% figure) and corroborated by an independent author interview describing the same qualitative finding.
Sources
- Can We Trust AI Benchmarks? An Interdisciplinary Review ... (arxiv.org)
- https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/file/3b8a614226a953a8cd9526fca6fe9ba5-Paper-round2.pdf (datasets-benchmarks-proceedings.neurips.cc)
- https://aihub.org/2022/02/17/the-life-of-a-dataset-in-machine-learning-research-interview-with-bernard-koch/ (aihub.org)