OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch
OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3. In some cases, the numbers on the updated versions showed Astra performing better, while numbers for models from OpenAI’s arch rival Anthropic got worse.

TL;DR
- OpenAI modified evaluation benchmarks for its GPT-6 Astra model after its initial announcement, with some scores improving for Astra and others worsening for rivals like Anthropic.
- The changes were made during a troubled rollout of OpenAI's blog post, which experienced technical difficulties and delays.
- Specific metrics, such as Astra's hallucination rate and performance on cybersecurity and math evaluations, were altered, sometimes significantly, in subsequent versions of the blog post.
- Experts and researchers expressed concerns about 'benchmaxxing,' a practice of re-running evaluations under optimized conditions to maximize scores, and the lack of detailed information in system cards regarding evaluation methodologies.
- The practice of altering benchmark scores, even after initial publication, raises questions about the reliability and transparency of AI performance metrics in a highly competitive industry.
- OpenAI stated that evaluations inherently have 'noise' and that changes were made to ensure numbers represent their best estimate of model performance for meaningful comparisons.