- Fortune discovered that OpenAI repeatedly changed GPT-6 Astra’s evaluation metrics after the introductory post was published on September 3.
- The release process faced glitches when links remained inactive for nearly two hours before displaying normally.
- OpenAI stated that the article was initially posted and then withdrawn for reasons unrelated to benchmarks, then reposted with several modified metrics.
- GPT-6 Astra’s internal hallucination rate was initially 4.2%, then dropped to 2.0%, and was later adjusted back to 4.2%.
- The hallucination index for GPT-5.6 Sol also changed from 12.2% to 9.4% before returning to the original level.
- Astra’s ARC-AGI-3 score in press documents was 98.6%, but was raised to 99.9% in the public version.
- Astra’s Terminal-Bench 4.0 score increased slightly from 57.7% to 57.9% after the article was updated.
- Some scores for Anthropic models also changed, such as Claude Fable 5.1’s FrontierMath, which once dropped from 87.8% to 78.0% before being readjusted.
- OpenAI explained that benchmark results depend on many factors such as checkpoints, harnesses, reasoning levels, and evaluation configurations, making adjustments before or after release normal.
- The company stated its goal is to reflect the best estimate of performance so that users can compare models more accurately.
- Stanford University researchers noted that the “benchmaxxing” phenomenon—running benchmarks multiple times to achieve the highest result—is a reality in the AI industry and called for more transparency in evaluation methods.
- Some experts proposed that AI companies should clearly disclose all changes in testing conditions when updating benchmarks so that the research community and customers understand the exact meaning of the figures.
📌 AI benchmarks are increasingly becoming a vital competitive factor but also spark significant controversy regarding transparency. OpenAI’s continuous adjustment of GPT-6 Astra’s metrics after announcement raises questions about how model performance is measured and presented. While the company maintains this is a normal technical update process, many experts argue that the AI industry needs more public standards for testing conditions and reasons for benchmark changes so that users, researchers, and investors can evaluate them fairly.
