Google's Gemma-2 9b and OpenAI's GPT-4o performed poorly on a Stanford team's descriptive and normative bias benchmarks despite scoring near-perfectly on Anthropic's DiscrimEval.
Notes on verification
Directly confirmed by the original Stanford arXiv paper and Stanford HAI news summary, which explicitly state that Gemma-2 9b and GPT-4o scored near-perfectly on existing fairness benchmarks like DiscrimEval but poorly on the new DiffAware/CtxtAware benchmarks.
Sources
- These new AI benchmarks could help make models less biased (technologyreview.com)
- https://arxiv.org/abs/2502.01926 (arxiv.org)
- https://arxiv.org/html/2502.01926v2 (arxiv.org)
- https://hai.stanford.edu/news/ais-fairness-problem-when-treating-everyone-the-same-is-the-wrong-approach (hai.stanford.edu)
- https://law.stanford.edu/stanford-lawyer/articles/ai-fairness-through-difference-awareness/ (law.stanford.edu)