When GPT-4 was tested in 2023 on easy Codeforces problems, Narayanan and Kapoor (2023b) found that it could regularly solve problems added before 5 September 2021 but got no question right among problems added afterward.
Notes on verification
Confirmed by the original Narayanan/Kapoor blog post and corroborated by a later Princeton CRCL paper and a third independent benchmark paper. Minor discrepancy in exact date wording (Sept 5 vs Sept 12 cutoff) does not undermine the core finding of a sharp performance drop-off tied to training data cutoff.
Sources
- Can We Trust AI Benchmarks? An Interdisciplinary Review ... (arxiv.org)
- https://www.normaltech.ai/p/gpt-4-and-professional-benchmarks (normaltech.ai)
- https://shana.codes/posts/week-notes-4-03-23.html (shana.codes)
- https://www.cs.princeton.edu/~sayashk/papers/crcl-kapoor-henderson-narayanan.pdf (cs.princeton.edu)