OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch

OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3. In some cases, the numbers on the updated versions showed Astra performing better, while numbers for models from OpenAI’s arch rival Anthropic got worse.

OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch

TL;DR

  • OpenAI modified evaluation benchmarks for its GPT-6 Astra model after its initial announcement, with some scores improving for Astra and others worsening for rivals like Anthropic.
  • The changes were made during a troubled rollout of OpenAI's blog post, which experienced technical difficulties and delays.
  • Specific metrics, such as Astra's hallucination rate and performance on cybersecurity and math evaluations, were altered, sometimes significantly, in subsequent versions of the blog post.
  • Experts and researchers expressed concerns about 'benchmaxxing,' a practice of re-running evaluations under optimized conditions to maximize scores, and the lack of detailed information in system cards regarding evaluation methodologies.
  • The practice of altering benchmark scores, even after initial publication, raises questions about the reliability and transparency of AI performance metrics in a highly competitive industry.
  • OpenAI stated that evaluations inherently have 'noise' and that changes were made to ensure numbers represent their best estimate of model performance for meaningful comparisons.