OpenAI Adjusts GPT-6 Astra Benchmark Data Amid Scrutiny Over Performance Metrics
OpenAI has faced scrutiny after repeatedly modifying performance metrics for its newly released GPT-6 Astra model. Reports indicate that the internal hallucination rate was adjusted from 4.2% to 2% before reverting to 4.2%, while scores on the ARC-AGI-3 benchmark rose from 98.6% to 99.99%. OpenAI stated that these fluctuations are due to variations in model checkpoints, tool configurations, and testing methods. The incident highlights the industry-wide practice of 'Benchmaxxing,' where companies optimize testing conditions to inflate benchmark results.
Summaries are written by AI from the original article. Not investment advice.