OpenAI has unveiled its latest flagship model, GPT-6 Astra, in a briefing where company president Greg Brockman declared, “Welcome to the AGI era.” However, the first independent benchmarking results from Artificial Analysis paint a more sobering picture. According to the firm, which measures more than 500 AI models and over 1,000 endpoints, GPT-6 Astra performs on par with its predecessor, GPT-5.6 Sol, and trails models from Anthropic and Meta.
In the Artificial Analysis Intelligence Index, GPT-6 Astra scores 61 points, the same as GPT-5.6 Sol. Anthropic’s Claude Fable 5.1 leads the index with 66 points, and Meta’s Muse Spark 1.3 (max) also outperforms Astra. Astra did better in the Artificial Analysis Coding Agent Index, where it scored 67 points in the Codex harness, on par with Claude Opus 5 (70 points) and Fable 5 in Claude Code, and on par with Muse Spark 1.3 in Muse Code.
The benchmarking firm noted that while GPT-6 Astra outperformed in coding tasks, it faced significant competition in other areas. The Intelligence Index, which evaluates mathematics, science, programming, long document reasoning, and factual knowledge, showed that Astra saved 10 percent in output tokens compared to its predecessor, but the price increase offset this benefit. At maximum effort, Astra costs 75 percent more per task than GPT-5.6 Sol.
Cost was another factor highlighted by the benchmarking firm. OpenAI has raised prices 2.5 times, from $4 to $10 per million input tokens and from $20 to $50 per million output tokens. However, GPT-6 Astra showed improved token efficiency, needing roughly one third of the tokens used by GPT-5.6 Sol (max) and about one fifth of those used by Claude Opus 5 (xhigh). At max effort, Astra costs about the same per task as GPT-5.6 Sol, while scoring two index points higher. In the Coding Agent Index, Astra’s cost per task is less than half that of Claude Fable 5.
The analysis also revealed significant progress in AA-Omniscience, the firm’s knowledge and hallucination benchmark, where the hallucination rate dropped from 92 to 51 percent, and accuracy increased by four points. In agentic knowledge work, Astra improved in AA-Briefcase, gaining around 80 Elo points, and in Analytical Quality Elo, but lost ground in Presentation Quality Elo. In GDPval-AA v2, a benchmark covering economically valuable tasks, the model lost approximately 80 Elo points.
For developers and companies, the benefits of GPT-6 Astra are mixed. In agentic coding, Astra offers strong value for money, but the gap to Anthropic’s models persists in raw model intelligence as API costs increase. OpenAI developed the model in its largest training run to date, using over 100,000 GPUs at the Stargate data center in Texas, and classified it as “critical” under its own cybersecurity framework.
While the AGI claim is compelling, benchmarking results suggest that GPT-6 Astra’s performance is still competitive but not yet leading. Whether this marks the beginning of the AGI era remains to be seen, and the benchmark tables from Artificial Analysis do not provide definitive confirmation.
Source: https://www.trendingtopics.eu/gpt-6-astra-trails-top-models-from-anthropic-and-meta-in-benchmarks/
Thinking about building an AI product?
Get in Touch