AI models evaluated against FrontierMath include OpenAI's GPT-4o and o1-preview.
Notes on verification
Confirmed by Epoch AI's official FrontierMath benchmark documentation and the arXiv paper describing the benchmark methodology and evaluated models.
Sources
- New secret math benchmark stumps AI models and PhDs alike (arstechnica.com)
- https://epoch.ai/frontiermath/tiers-1-4/the-benchmark (epoch.ai)
- https://arxiv.org/pdf/2411.04872 (arxiv.org)