GPT-6.1 Sol: One Point Behind Astra, a Quarter of the Bill

GPT-6.1 Sol scores 52, one point behind Astra, at $0.72 an index task against Astra's $3.26. On coding, xhigh beats max and Astra's Codex score.

GPT-6.1 Sol: One Point Behind Astra, a Quarter of the Bill

GPT-6.1 Sol came out on September 29, 2026, seven days after GPT-6 Sol. It is not a new top of the Intelligence Index. Artificial Analysis puts it at 52, one point behind GPT-6 Astra at 53, four points ahead of GPT-6 Sol at 48, and five ahead of GPT-5.6 Sol at 47. Opus 5.5 still leads that chart at 58. Sonnet 5.5 is at 56.

The bill is the part that moved. At max effort, one Intelligence Index task costs $0.72, against $3.26 for Astra, $1.05 for GPT-6 Sol, and $1.99 for GPT-5.6 Sol. Less than a quarter of Astra, and 31% under the Sol it replaces. The list price did not change. Input is still $2 per million tokens and output is still $10. The new cut is the cache read: $0.10, a 95% discount, half of GPT-6 Sol's $0.20.

Intelligence Index and cost per task, GPT-6.1 Sol at 52
Artificial Analysis. GPT-6.1 Sol at max is 52, one point behind Astra. Every effort level sits on the cost frontier. Astra and Opus sit further right.

The setting is the product

On the API the id is gpt-6.1-sol. Context is 1.05 million tokens. It is in the API today, and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu. It is not in Chat yet. An Ultrafast tier, up to 8x generation speed in Codex, is planned for the coming days. That speed is a separate product from the price below.

Artificial Analysis ran five efforts. Medium already matches last week's Sol at max.

Effort Intelligence Index Cost per index task
Low 42 $0.13
Medium 48 $0.21
High 50 $0.32
Xhigh 51 $0.39
Max 52 $0.72

Medium at 48 is GPT-6 Sol's max score, at $0.21 against $1.05. The last point, from xhigh at 51 to max at 52, almost doubles the task price, from $0.39 to $0.72. All five points sit on the cost frontier: for a given index score, Artificial Analysis does not see a cheaper model. GPT-6 Sol's old points fall off that line.

The model spends 10% to 30% more output tokens than GPT-6 Sol to get there. Low and medium land on the token-efficiency frontier. Max does not. The tokens are cheap enough that max still wins on dollars, and verbose enough that it loses on tokens. The full index run produced 67 million output tokens, under a median of 82 million. That is the run, not the per-task count.

Index versus output tokens per task
Artificial Analysis. Low and medium sit on the token frontier. Max uses enough extra tokens to fall off it.

Coding: buy xhigh, not the headline

On the Coding Agent Index, in Codex, GPT-6.1 Sol at max scores 60. That is 3 points over GPT-6 Sol at max, 57, and 2 points under Astra at 62. Xhigh scores 63. That is 1 point over Astra, 3 points over its own max, and 6 points over GPT-6 Sol at max. Artificial Analysis puts that xhigh point at less than 15% of Astra's cost per coding task.

The bars above 63 are a different harness. Sonnet 5.5 in Claude Code is at 68. Opus 5.5 in Claude Code is at 66. Fable 5.1 in Claude Code is at 62, tied with Astra's Codex score. Grok 4.7 on Grok Build is at 56. A 63 in Codex is not a pass of a 68 in Claude Code. It is a pass of Astra on the harness OpenAI's own agent uses, at a fraction of the bill.

Coding Agent Index and cost per task
Artificial Analysis. In Codex, xhigh is 63 and max is 60. Astra's Codex max is 62, further right on the cost chart.

Knowledge work recovered. It did not take the lead.

Last week's Sol got cheaper and thinner. This one spends the week differently. On AA-Briefcase, Artificial Analysis measures about 80 Elo over GPT-6 Sol, from a higher rubric score and higher analytical quality. Presentation quality slips. The chart puts the new score in Astra's band, and leaves Opus 5.5 at 1822 and Sonnet 5.5 at 1811 at the top of the same figure. Eighty Elo is a real recovery. It is not the knowledge-work lead.

On AA-Omniscience the mechanism is also different from last week. Accuracy rises 8 points. The hallucination rate falls from 60% to 54%. Last week's drop was mostly Sol attempting fewer questions. This chart's accuracy is the share of questions answered correctly out of all questions, whether or not the model tries. Both the hit rate and the error rate moved the right way. Opus and Astra still sit above it on that index.

AA-Briefcase score, rubric, and presentation
Artificial Analysis. Briefcase rises about 80 Elo. Presentation quality is the series that slips.

AA-Omniscience index, accuracy, and hallucination rate
Artificial Analysis. Accuracy rises 8 points. Hallucination falls from 60% to 54%.

What OpenAI's own charts add, and where max is the wrong quote

OpenAI's launch charts are not files you can embed. The claims below are theirs. BenchLM's rows are the max-effort cells from those same charts, and BenchLM says it keeps max even when a lower effort scores higher. That split is the useful part.

On DeepSWE v1.1, OpenAI says GPT-6.1 Sol matches Astra at about a fifth of the cost, and beats GPT-6 Sol's best score by 6.4 points, at a lower effort and a lower cost than that best score. BenchLM's max row is 71.9%. The match with Astra is not that max cell.

On OSWorld 2.0's offline set, partial reward, OpenAI says max effort beats GPT-6 Sol by 7 points, comes within 2.1 points of Astra, and costs about a seventh as much per task. BenchLM lists the Sol figure as 71.4% partial.

On AutomationBench 1.0.6, the lead OpenAI quotes is medium effort: 2.2 points over Opus 5.5, at about a third of the cost, and 4.8 points over GPT-6 Sol at the same setting. BenchLM's max row is 36.1%. OpenAI also notes that the Fable point on this chart leaves out the cost of fallbacks, which fired on about 40% of Fable's tasks.

On Terminal-Bench Science 0.1, OpenAI says max effort more than doubles GPT-6 Sol, at less than half the cost. The average task is $5.47, against $23.21 for Opus 5.5 and $23.80 for Astra. Astra still has the highest score they tested, 68.1%. BenchLM lists Sol at 57.0%. Use Astra for the hardest scientific terminal work. The $5.47 is the reason to use Sol on the rest.

On GDP.pdf, OpenAI says Sol scores above Opus 5.5 with fallbacks at less than half the cost per task, and approaches Astra at about a fifth of the cost. Artificial Analysis, on its own GDP.pdf, measures a 6 point gain over GPT-6 Sol.

OpenAI's factuality check is a third measurement, not Omniscience. At low effort, the share of answers with a factual error on hard, user-flagged chats falls from 11.4% to 7.7%. Across the efforts they tested, Sol stays within 1.9 points of Astra, at less than a fifth of the cost. Those prompts were chosen because an earlier model had already been wrong. They are not a typical chat.

In practice

- Move GPT-6 Sol traffic to gpt-6.1-sol. Medium on this index is 48 at $0.21, which is last week's max score at about a fifth of last week's max bill. The list price is still $2 and $10. The new discount is the $0.10 cache read.

- For coding agents, use xhigh. In Codex that is 63, one point over Astra's 62, at under 15% of Astra's task cost. Max is 60. The last index point, 51 to 52, is the expensive one, $0.39 to $0.72.

- Leave Astra in place for the hardest scientific terminal work. OpenAI's own chart still has Astra at 68.1% and prices Sol's run at $5.47 against Astra's $23.80. Leave Opus 5.5 in place when you need the index at 58 or the top of Briefcase.

- Treat the DeepSWE "matches Astra" line and the 71.9% max row as two quotes. The match is a lower effort. The 71.9% is BenchLM's max cell.

- Treat the AutomationBench lead the same way. The 2.2 points over Opus 5.5 are medium against medium. The max row BenchLM lists is 36.1%.

- Quote Terminal-Bench Science as OpenAI's comparison: Sol more than doubles GPT-6 Sol, Astra still leads at 68.1%, and the task prices are $5.47, $23.21, and $23.80.

- Tools stay on the Responses API. Chat Completions on this model does not take tool calls. The model is in Work and Codex, not in Chat yet.

Sources: Introducing GPT-6.1 Sol, Artificial Analysis, the effort comparison, and BenchLM. The five charts, the index of 52, and the $0.72 task are Artificial Analysis. The DeepSWE, OSWorld, AutomationBench, science-terminal, and factuality lines are OpenAI's. BenchLM's 71.9%, 71.4%, 36.1%, and 57.0% are the max-effort cells from those charts.

Keep reading

Explore more product news and best practices for teams building with Plataforma Tess pre_prod.

Build with TESS

Turn ideas from this article into working AI workflows.

Create agents, automations, and knowledge-powered workflows in one platform built for teams.