OpenAI Adjusts GPT-6 Astra Evaluation Data, Raising Transparency Concerns

icon MarsBit
Share
AI summary iconSummary
OpenAI has reportedly modified GPT-6 Astra’s evaluation data, raising concerns about transparency. Astra’s hallucination rate fell from 4.2% to 2%, while competitors such as Anthropic and GPT-5.6 Sol experienced declines in math scores. OpenAI states the changes reflect accurate performance, but Stanford researchers warn of “benchmaxxing.” As altcoins to watch gain momentum amid market shifts, reliable data remains critical. Inflation data trends also underscore the need for clear, reproducible benchmarks in both AI and crypto sectors.

Breaking News from Beating AI: Since OpenAI released GPT-6 Astra on September 3, multiple model evaluation benchmarks have been continuously adjusted, with some revisions improving Astra’s performance while others caused declines in competing models, sparking external concerns about AI “benchmark stuffing” and evaluation transparency. For instance, Astra’s hallucination rate was previously lowered from 4.2% to 2%, while GPT-5.6 Sol dropped from 12.2% to 9.4%; both later reverted to 4.2% and 12.2%, respectively. In math evaluations, Anthropic’s Fable 5.1 score fell from 87.8% to 78% before rebounding to 83%, while GPT-5.6 Sol dropped from 83% to 80.5% before returning to 83%. Additionally, Astra’s score on the ARC-AGI-3 benchmark rose from an initial draft of 98.6% to 99.99% on the final page, and its programming evaluation score increased slightly from 57.7% to 57.9%. OpenAI stated that evaluation results can be influenced by model versions, tool configurations, reasoning levels, and test runs, and that these adjustments were made to ensure data more accurately reflects the models’ peak performance. However, Stanford University researchers suggest that frequently rerunning evaluations may involve so-called “Benchmaxxing”—adjusting test conditions to maximize benchmark scores. Industry experts note that as competition among AI models intensifies, benchmark data has become a critical metric for assessing model capabilities and capturing market share, with growing attention being paid to improving the transparency and reproducibility of benchmark tests.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.