GPT-6 Astra: What Analysis and Benchmarks Say About the Model

Independent benchmarks show GPT-6 Astra excels in agentic coding at a fraction of the cost, but its higher price offsets efficiency gains in general intelligence tasks.

GPT-6 Astra: What Analysis and Benchmarks Say About the Model

GPT-6 Astra, released on September 3, 2026, delivers meaningful gains in agentic coding tasks. It matches the strongest models in the market at less than half their cost. On general intelligence, however, its lower token usage is offset by a higher price.

Why this matters for Tess users

Tess brings multiple AI models into one workspace. Choosing the right model for each task directly affects both output quality and cost. GPT-6 Astra is OpenAI's new reasoning model and the successor to GPT-5.6 Sol. Independent results from Artificial Analysis help show where it is a strong choice — and where switching may not yet make financial sense.

Pricing: substantially higher

GPT-6 Astra is priced at 2.5x GPT-5.6 Sol's rates: from $4 / $20 to $10 / $50 per million input / output tokens. It keeps the 90% discount for cache reads and the 25% premium for cache writes.

Quick takeaway: there are two distinct stories. In the Coding Agent Index, Astra matches leading models at less than half the cost. In the Intelligence Index, it is more token-efficient than its predecessor, but the price increase erases that advantage.

Highlights from the Coding Agent Index

➤ Competes with the market leaders

In Codex, GPT-6 Astra scores 67 in the index — roughly equal to Claude Opus 5 and Fable 5 in Claude Code, as well as Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the index with 70.

➤ 70% more token-efficient than GPT-5.6 Sol

Astra uses one third of the tokens consumed by GPT-5.6 Sol (max effort) in the Codex harness, and one fifth of the tokens used by Claude Opus 5 (xhigh). Its effort settings occupy the token-efficiency Pareto frontier.

➤ Leads the coding cost-efficiency frontier

At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring two points higher. Per task, it costs less than half as much as Claude Fable 5 for the same score.

Artificial Analysis — Coding Agent Index
Source: Artificial Analysis — Coding Agent Index

The coding cost improvements are primarily driven by token efficiency: at max effort, Astra reduces token use by around 3x compared with GPT-5.6 Sol (max).

Artificial Analysis — coding token efficiency
Source: Artificial Analysis — token efficiency in the Coding Agent Index

Artificial Analysis — Coding Agent Index vs. cost per task
Source: Artificial Analysis — Coding Agent Index vs. cost per task

Highlights from the Intelligence Index

➤ Ties GPT-5.6 Sol on general intelligence

GPT-6 Astra matches GPT-5.6 Sol at 61 in the index. That is five points below Claude Fable 5.1 (max with fallback) and also behind Meta's Muse Spark 1.3 (max).

➤ Around 10% fewer output tokens, but at a higher price

At max effort, Astra uses about 10% fewer output tokens than GPT-5.6 Sol and defines a new efficiency frontier for intelligence versus output tokens per task. But its 2.5x price increase makes it 75% more expensive per task than its predecessor.

➤ Cuts hallucinations roughly in half

In AA-Omniscience, Artificial Analysis's knowledge and hallucination benchmark, the hallucination rate drops from 92% to 51% at max effort. Unlike some models, this gain does not come at the expense of accuracy: Astra's accuracy also rises by four points.

➤ Gains roughly 80 Elo points in AA-Briefcase

GPT-6 Astra improves by around 80 points in AA-Briefcase, a long-horizon knowledge-work evaluation that tests multi-week projects with linked tasks and thousands of source files. It improves rubric scores and Analytical Quality Elo, while Presentation Quality Elo declines; GPT-5.6 Sol (max) remains the leader on presentation quality.

➤ Mixed progress in other evaluations

The model gains six points on Humanity's Last Exam, but loses around 80 Elo points in GDPval-AA v2, which measures economically valuable tasks across 44 occupations. Artificial Analysis also found 2–3 point regressions in τ³-Banking, SciCode and AA-LCR.

Artificial Analysis — Intelligence Index vs. output tokens per task
Source: Artificial Analysis — Intelligence Index vs. output tokens per task

Artificial Analysis — Intelligence Index vs. cost per task
Source: Artificial Analysis — Intelligence Index vs. cost per task

Artificial Analysis — AA-Omniscience
Source: Artificial Analysis — AA-Omniscience (knowledge and hallucination)

Artificial Analysis — AA-Briefcase and GDPval-AA v2
Source: Artificial Analysis — AA-Briefcase and GDPval-AA v2

Artificial Analysis — Intelligence Index evaluation breakdown
Source: Artificial Analysis — Intelligence Index v4.1.1 evaluation breakdown

In practice

  • For coding and agentic tasks: GPT-6 Astra is a strong option, delivering performance at the level of market-leading models at a fraction of the cost of the most expensive alternative.
  • For general reasoning and intelligence: lower token use does not offset the price increase, so compare alternatives before migrating.
  • For fact-sensitive tasks: the reduction in hallucinations is one of Astra's most important gains, but outputs still require appropriate review.
  • For long and complex projects: analytical quality improves materially, while presentation quality may require additional editing.

The best way to choose a model is to test it against your own workflow. Tess makes that comparison possible by giving your team access to multiple AI models in one workspace.


Original data, charts and methodology: Artificial Analysis — “Benchmarking GPT-6 Astra”, published September 3, 2026.

Build with TESS

Turn ideas from this article into working AI workflows.

Create agents, automations, and knowledge-powered workflows in one platform built for teams.