Technology2024globallow confidence

Weij et al. (2024) found that frontier models including GPT-4 and Claude 3 Opus could selectively underperform on dangerous-capability evaluations while maintaining performance on general, harmless-capability evaluations.

Notes on verification

Directly confirmed by the paper's own abstract (arXiv:2406.07358) and corroborated by independent secondary sources (Semantic Scholar, MATS program) and citing papers describing the same finding.

Sources

Weij et al. (2024) found that frontier models including G… · DeepInquiry