GPT-6 Sol and Luna: Half the Bill, a Thinner Deliverable

GPT-6 Sol and Luna match GPT-5.6 on the Intelligence Index at about half the price. Knowledge-work scores fell, and Sol answers less often.

GPT-6 Sol and GPT-6 Luna, released on September 22, 2026, are not a new top of the Intelligence Index. They are the same neighborhood as GPT-5.6, at about half the API price. That is the part worth switching. The part worth not switching is the deliverable: on knowledge work, both models got shorter, and they drop pieces the rubric asked for.

Sol is the coding and agent model. On the OpenAI API the id is gpt-6-sol, at $2 per million input tokens and $10 per million output. Luna is the volume model, gpt-6-luna, at $0.10 and $0.50. OpenAI's line is that they carry Astra's gains into something you can run all day. Artificial Analysis's line is sharper. The index barely moved. The bill did. A class of professional work got worse.

Why this matters

A workspace that already runs Astra, Opus 5.5, and GPT-5.6 Sol does not need two new names. It needs to know which traffic moves, and which job should stay where the deliverable was already good.

Sol is built for complex coding and agentic workflows. Its context window is 1,050,000 tokens, max output is 128,000, and the knowledge cutoff is April 20, 2026. Default effort is medium. The scores below are max. Luna's published role is focused, high-volume work: a summary, an extraction, a short answer. Its knowledge cutoff is May 18, 2026.

Both are in the API today. ChatGPT Work and Codex get them for Plus, Pro, Business, Enterprise, and Edu. Free and Go can try Luna in the desktop app. The ChatGPT app and site roll out through the day.

The bill

The rate card is half of GPT-5.6, which OpenAI describes as 50% under that generation's promotional price. Cache reads stay at 10% of input. Cache writes stay at 1.25 times input.

Per 1M tokens GPT-6 Sol GPT-5.6 Sol GPT-6 Luna GPT-5.6 Luna
Input $2 $4 $0.10 $0.20
Output $10 $20 $0.50 $1.20
Cache read $0.20 $0.40 $0.01 $0.02
Cache write $2.50 $5 $0.125 $0.25

The task bill falls by about the same amount, and not because the models got terser. On the Intelligence Index, Sol max uses 31,000 output tokens per task against 29,000 for GPT-5.6 Sol max. Luna uses 51,000 against 41,000. Slightly more tokens, much less money. Artificial Analysis puts the index run at $1.06 per task for Sol max, against $1.99, about half. Luna max is $0.07 against $0.18, about 60% less.

That $1.06 is the intelligence-index task. It is not the coding-agent task, which Sol max runs at $2.99, also about half of GPT-5.6 Sol max. Keep the two invoices apart.

Two switches give the discount back. Fast mode is 2x the rate. And any request whose input passes 272,000 tokens is repriced on the whole prompt: input and cache at 2x, output at 1.5x. A long agent context is not the price in the table. Batch and Flex are half of standard, on top of the new list.

Intelligence Index and cost per task for GPT-6 Sol and Luna
Source: Artificial Analysis. Sol max scores 48, one point over GPT-5.6 Sol max. Luna max stays at 37. The cost curve is what moved.

Level on the index, behind the models you already know

On Artificial Analysis's Intelligence Index, Sol max scores 48. GPT-5.6 Sol max scores 47. Luna max scores 37, the same as GPT-5.6 Luna max. "Level" is the right word. A one-point step is not a new tier.

The same chart still has Claude Opus 5.5 at 58, and GPT-6 Astra and Claude Fable 5.1 at 53. Sol did not take their jobs. It took the stretch of the cost frontier where a team was already paying Sol prices. GPT-6 Sol's effort line sits left of GPT-5.6 Sol. Luna does the same under it. OpenAI captured that frontier by cutting price, not by passing Astra.

Coding moved a little. The invoice moved a lot.

In OpenAI's Codex harness, the Coding Agent Index puts Sol max at 57, up 2 from GPT-5.6 Sol max at 55. Luna max scores 41, down 2 from 43. The Sol gain is Terminal-Bench 4.0 at 43% against 37%, and SWE-Atlas-QnA at 58% against 54%. Luna's drop is SWE-Atlas-QnA at 44% against 49%, and DeepSWE v1.1 at 64% against 66%.

At $2.99 per coding task, Sol max costs about half of its predecessor and sits on the coding cost frontier. Fable 5.1 and Astra, in their own harnesses, are still at 62. Opus 5 is at 60. Grok 4.7 on Grok Build is at 56, one point under Sol, on a different harness and a different bill. Do not read 57 as a pass of those models.

There is a second Terminal-Bench number, and it is not this one. Inside the Intelligence Index, Sol's Terminal-Bench 4.0 is 44% against 40%, and Luna's is 13% against 12%. The 43% versus 37% figure is the Coding Agent Index, in Codex. Label them or they will get quoted as one score.

Coding Agent Index and cost per task, GPT-6 Sol at 57
Source: Artificial Analysis. Codex harness. Sol max is 57 and on the cost frontier. Luna max is 41, two points under GPT-5.6 Luna.

The hallucination drop is mostly a refusal

Both models invent less on AA-Omniscience, at max effort. Sol's hallucination rate goes from 92% to 60%. Luna's goes from 93% to 77%. The index, which rewards a correct answer and punishes a wrong one, moves Sol from 22 to 27 and Luna from −10 to 1.

Read the mechanism before you treat that as a smarter model. Sol attempts 83% of questions, against 99% for GPT-5.6 Sol max. Wrong answers fall by about a quarter. Accuracy also falls, from 59% to 54%. Luna's accuracy stays about flat, 44% against 43%, while it answers fewer questions. Sol got safer by declining. It did not get more often right.

AA-Omniscience index, accuracy, and hallucination rate
Source: Artificial Analysis. Sol's hallucination rate falls from 92% to 60%. Accuracy falls with it, from 59% to 54%.

The deliverable got thinner

The regression is in knowledge work, and Artificial Analysis checked it by hand. On GDPval-AA v2.1, real tasks across 44 occupations, Sol max drops about 100 Elo. The chart puts GPT-5.6 Sol max at 1588 and GPT-6 Sol max at 1487. Luna max drops about 75 Elo, to 1367. Astra is at 1542 on the same board. Opus 5.5 is at 1846. A price cut does not close that gap.

On AA-Briefcase, Sol is level and Luna drops about 45 Elo. The misses are presentation and missing rubric items, not a random wrong fact. Shorter output, required pieces left out. For a deck, a brief, or a multi-week file set, that is the failure mode. The model finishes. The deliverable is thinner than the one GPT-5.6 handed you.

AutomationBench-AA did go up: Sol 62% against 60%, Luna 53% against 50%. That is a workflow pass rate, not a document. It does not cancel the GDPval drop.

GDPval-AA v2.1, GPT-6 Sol at 1487 and Luna at 1367
Source: Artificial Analysis. Sol falls about 100 Elo from GPT-5.6 Sol. Luna falls about 75. Opus 5.5 remains at 1846.

In practice

- Move GPT-5.6 Sol and Luna traffic onto gpt-6-sol and gpt-6-luna when the job is the same shape. The Intelligence Index is flat. The rate card is about half, and the models spend slightly more tokens, not fewer.

- Do not retire Astra or Opus 5.5 because Sol max scores 48. On this index Astra is 53 and Opus 5.5 is 58. Sol won the cost frontier under them.

- Keep briefs, decks, and long knowledge-work files on the model that was already winning GDPval. Sol dropped about 100 Elo and Luna about 75, on shorter deliverables that skip rubric items.

- Treat Sol's hallucination drop as a refusal policy. The rate went from 92% to 60% because it attempts 83% of questions instead of 99%. Accuracy fell from 59% to 54%. Use it when silence is cheaper than a wrong answer.

- Keep the two Terminal-Bench figures apart. 43% against 37% is Sol on the Coding Agent Index, in Codex. 44% against 40% is Sol on the Intelligence Index eval. Luna's coding score went down, 41 against 43.

- Assume the half-price table stops at 272,000 input tokens. Past that, the whole request is 2x on input and cache and 1.5x on output. Fast mode is 2x and spends the discount.

- Sol's charts are max effort. The API default is medium. On Sol, Chat Completions will function-call only with reasoning effort set to none. Tools belong on the Responses API.

Sources: Introducing GPT-6 Sol and Luna, GPT-6 Sol and GPT-6 Luna on the OpenAI API, and Artificial Analysis. The charts and the Elo drops are Artificial Analysis's. The rate card is OpenAI's.

Sigue leyendo

Explora más novedades y mejores prácticas para equipos que construyen con Plataforma Tess pre_prod.

Construye con TESS

Convierte ideas de este artículo en flujos de IA funcionales.

Crea agentes, automatizaciones y flujos con conocimiento en una plataforma hecha para equipos.