OpenAI briefly halved GPT-6 Astra's hallucination rate in post-launch benchmark edit
The company also boosted its own math scores and cut Anthropic's, raising questions about 'benchmaxxing' in AI evaluation.
By Turki Al-Mutairi·· 3 min

Loading the Newsroom
Curating from trusted global sources…
1 briefing · “ai evaluation”
The company also boosted its own math scores and cut Anthropic's, raising questions about 'benchmaxxing' in AI evaluation.