OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, filling in the two tiers below the GPT-6 Astra flagship it shipped earlier in the month. Both launched at half the price of their GPT-5.6 versions, with Sol at $2 per million input tokens and Luna at $0.10.
That makes this release bigger than a routine model refresh. Every app built on GPT-5.6 Sol or Luna now has a successor at half the per-token cost, and the GPT-6 family finally runs from a frontier tier down to a model cheap enough for high-volume extraction and classification.
The tier names hide how much the lineup overlaps. On independent benchmarks, Sol at its highest effort setting only narrowly outscores Astra at its lowest, yet Astra costs less per task and starts answering in seconds where Sol takes close to two minutes. With three tiers and five or six effort levels each, the real choice is a tier and an effort setting for every step of an agent.
So the question with GPT-6 Sol vs Luna vs Astra is which one to use for each step in 2026.
What are GPT-6 Sol and Luna?
GPT-6 Sol and GPT-6 Luna are the middle and entry tiers of the GPT-6 family. Both carry names from the GPT-5.6 lineup, but they're new models with new pricing.
OpenAI says the two models were trained with methods similar to Astra's, and they keep Astra's 1.05M-token context window and 128K-token output limit.
GPT-6 Sol
GPT-6 Sol is the tier OpenAI positions for complex coding and agentic workflows. It costs $2 per million input tokens and $10 per million output tokens, and its reasoning effort runs from none to max, with medium as the default.
The name is where it gets confusing. In GPT-5.6, Sol was the frontier tier, with Terra in the middle and Luna at the bottom. GPT-6 puts Astra above Sol, so the same name now sits one rung lower, and gpt-6-sol is a separate model from gpt-5.6-sol.
GPT-6 Luna
GPT-6 Luna is the GPT-6 model OpenAI builds for focused, high-volume tasks, where each call is short and the output is easy to check.
At $0.10 input and $0.50 output per million tokens, Luna costs one twentieth of Sol and one hundredth of Astra on both lines. It supports the same none-to-max effort range as Sol.
Where GPT-6 Astra fits
GPT-6 Astra sits above both as the model OpenAI positions for the hardest end-to-end work. At $10/$50 per million tokens it costs five times Sol's price, and its effort range starts at low, with no none setting.
GPT-6 Sol vs Luna vs Astra: overview
All three GPT-6 models accept text and images and return text, with the same context window and output limit. The differences are in price and effort range.
| GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| API model ID | gpt-6-astra | gpt-6-sol | gpt-6-luna |
| Input / output per 1M | $10 / $50 | $2 / $10 | $0.10 / $0.50 |
| Cached input per 1M | $1.00 | $0.20 | $0.01 |
| Context window | 1.05M tokens | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 128K tokens | 128K tokens |
| Reasoning effort | low to max | none to max | none to max |
| Knowledge cutoff | Apr 30, 2026 | Apr 20, 2026 | May 18, 2026 |
| Best for | Hardest agent steps | General production work | High-volume, checkable steps |
Because the spec sheets barely differ, choosing between the three comes down to what each tier delivers per dollar at a given effort level.
GPT-6 Sol vs Luna vs Astra: pricing
Output tokens cost five times the input rate on all three GPT-6 tiers, and reasoning tokens bill as output. As a result, the effort level you pick often moves the bill more than the tier's headline input price.
Per-token pricing
Astra costs five times Sol on every line of the rate card, and Sol costs twenty times Luna. Compared with GPT-5.6, Sol and Luna come in at exactly half of the previous promotional rates of $4/$20 and $0.20/$1.20.
The service-tier multipliers are the same across the family. Batch and Flex are half of Standard, Fast mode is double, and regional processing adds 10% where it's available. For Sol and Luna, EU data residency only works with Standard processing, so an EU workload can't pair residency with the Batch or Flex discount.
Caching and long-context pricing
Cached input bills at 10% of the uncached rate on all three models, and cache writes bill at 1.25x. GPT-6 also keeps the cache when you change reasoning effort or toggle tools mid-conversation.
That second behavior matters for routing by step. Raising Sol from medium to high for a hard follow-up keeps the cached prefix, while moving the same conversation from Sol to Astra starts the cache over, since a cached prefix belongs to the model that processed it.
Long prompts have their own threshold. Once input passes 272K tokens, the whole request bills at 2x the input and cache rates and 1.5x the output rate, so a 300K-token Sol request pays $1.20 for input alone where a 270K-token request pays $0.54. If most of that context repeats between calls, prompt caching is the first place to recover it.
Cost per task
Per-token prices don't account for how many tokens each model spends. On the Artificial Analysis Intelligence Index at max effort, Astra uses about 27K output tokens per task against Sol's 31K.
As a result, the per-task gap is narrower than the rate card suggests. Astra at max costs $3.26 per task against Sol's $1.06, about three times as much where the per-token ratio is five, while Luna at max costs $0.07.
Effort level stretches the range further. Sol costs $0.13 per task at low and $1.06 at max, an eightfold spread inside one model.
Route GPT-6 Sol, Luna, and Astra through one endpoint
Switch GPT-6 tiers per request, see what every call costs by model, and get 20% off Astra, Sol, and Luna during the Respan launch deal.
GPT-6 Sol vs Luna vs Astra: performance
At max effort, Sol and Luna score close to their GPT-5.6 predecessors on independent benchmarks, so most of what changed in this release is cost. Astra still scores highest at every effort level the three tiers share.
Intelligence and coding
On the same Intelligence Index, a composite of 10 evaluations, Astra scores 53 at max effort, Sol 48, and Luna 37. Artificial Analysis's launch analysis found Sol's and Luna's scores level with GPT-5.6 overall, with gains on some evaluations and regressions on others.
Coding tells a similar story. On the Coding Agent Index, run in OpenAI's Codex harness, Sol at max scores 57, up two points from GPT-5.6 Sol, while Luna scores 41, down two from its predecessor.
Knowledge work
Knowledge work is where Sol and Luna gave ground. On GDPval-AA, a benchmark adapted from OpenAI's dataset of occupational tasks, Sol dropped about 100 Elo and Luna about 75 against their GPT-5.6 versions. The same analysis traced the drop to weaker presentation and deliverables that left out required elements.
Hallucination moved the other way. Sol's hallucination rate at max effort fell from 92% to 60%, mostly because it declines to answer more often, which also lowered its accuracy from 59% to 54%. Luna's rate fell from 93% to 77%.
On AA-Briefcase, a long-horizon knowledge work evaluation, Sol held level with its predecessor while Luna dropped about 45 Elo. For a step whose output someone reads and acts on, such as a migration plan or a written report, that's a reason to score the new models against your own rubric before moving the step over.
Speed and latency
At max effort, Sol generates 126 output tokens per second against Astra's 53. Speed figures for Luna aren't available yet.
Effort level drives time to first token far more than tier does. Astra returns its first token in about 2.8 seconds at low and 6.2 at medium, then jumps to 79 seconds at high and 361 at max, while Sol goes from 1.4 seconds at low to 8.1 at high and 107 at max.
For any step a user waits on, low or medium effort on either tier keeps the delay in single-digit seconds.
Performance by effort level
The same benchmarks score each model at every effort level it supports, including GPT-6 Luna, and the tiers overlap in two places.
| Effort | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna |
|---|---|---|---|
| max | 53 · $3.26 | 48 · $1.06 | 37 · $0.07 |
| xhigh | 52 · $2.31 | 44 · $0.53 | 34 · $0.04 |
| high | 51 · $1.73 | 43 · $0.37 | 32 · $0.03 |
| medium | 50 · $1.54 | 40 · $0.25 | 29 · $0.02 |
| low | 46 · $0.82 | 34 · $0.13 | 21 · $0.0045 |
| none | n/a | 28 · $0.33 | 18 · $0.01 |
Intelligence Index score · cost per task
At the top, Sol at max outscores Astra at low by two points, but costs about 30% more per task and takes 107 seconds to its first token instead of about 3. At the bottom, Luna at max (37) beats Sol at low (34) for roughly half the cost per task.
The none setting also costs more than it looks. Sol with no reasoning scores 28 and costs $0.33 per task, more than Sol at low, because it spends more output tokens per task, and Luna shows the same pattern at $0.01 against $0.0045. That matters if you call Sol or Luna through Chat Completions, which only supports function calling at none.
Which is best: Sol, Luna, or Astra?
No single GPT-6 model is the right choice for a whole agent. Astra scores highest at every effort level it shares with Sol, while Sol and Luna cover far more of the cost range, so the useful question is which tier and effort to use for each kind of step.
Pick a tier and an effort level
Across the benchmark data, a handful of configurations cover most agent steps:
- GPT-6 Astra at medium or high - For steps where a wrong answer is expensive, like planning a multi-file change or recovering after a tool call fails. Astra at medium scores 50 for $1.54 per task, above anything Sol reaches.
- GPT-6 Astra at low - For a hard step that also needs a fast first token. It scores 46 and starts answering in about 3 seconds.
- GPT-6 Sol at medium or high - The default for general agent work, tool calling, and code edits, scoring 40 to 43 for $0.25 to $0.37 per task.
- GPT-6 Luna at high or max - For extraction, classification, and routing, where the output has a right answer you can check. Luna at max outscores Sol at low for about half the cost.
Sol at max is the hardest setting to justify. Astra at low comes within two points of it for less per task and answers in seconds, so a step that needs more than Sol at xhigh usually belongs on Astra.
Route each agent step to a different tier
Every call in an agent run can set its own model and effort, so the tier decision can happen per step instead of once per application. A coding agent might triage an incoming issue with Luna at high, plan the change with Astra at medium, make the edits with Sol at medium, and write the pull request summary with Luna again.
Where those boundaries go depends partly on caching. Switching models restarts the prompt cache, so the split works best at natural handoffs, such as passing a finished plan to the step that executes it. Tool calling adds one more constraint, since tool calls on Sol and Luna at any effort above none need the Responses API.
A fixed split like this is simpler to reason about than an LLM router that picks a model per request from its content. Each step's assignment is also a single setting you can score on your own traffic, with a rubric-based evaluator such as LLM-as-a-judge.
Route GPT-6 Sol, Luna, and Astra with Respan

Routing agent steps across GPT-6 tiers pays off when switching models takes one change and every call's cost is visible afterward. Respan lets you route, observe, and evaluate every LLM call, and GPT-6 Astra, Sol, and Luna are 20% off on Respan during the launch deal.
Respan's LLM gateway puts GPT-6 Sol, GPT-6 Luna, and Astra behind one OpenAI-compatible endpoint with 1,000+ models, so moving a step to a different tier means changing one word in the request.
- Fallback when a provider fails - Define a fallback chain once, at the org level or per request, and traffic moves to the next model in the chain when a provider errors or rate-limits.
- Cost for every step - Each request is logged as a span with latency and cost attached, and spend breaks down by model, request, and end customer, with hard limits that stop spend at a ceiling.
- Traces across the whole run - Every LLM call, tool run, and agent turn nests in one trace, so a bad output leads back to the step and model that produced it.
- Evals before and after a switch - Build datasets from production logs, run them through each model with your evaluators, then deploy the same evaluators on live traffic so a regression after moving a step to Luna surfaces in real time.
Find the right GPT-6 tier for every step
Route each step to GPT-6 Sol, Luna, or Astra through one endpoint, see what every call costs by model, and get 20% off all three during the Respan launch deal.
FAQ
Is GPT-6 Sol the same as GPT-5.6 Sol?
No. gpt-6-sol is a new model that sits below Astra in the GPT-6 family, while gpt-5.6-sol was the top tier of GPT-5.6. GPT-6 Sol costs $2/$10 per million tokens against GPT-5.6 Sol's $4/$20 promotional rate, and at max effort it scores one point higher on independent benchmarks.
Is there a GPT-6 Terra?
OpenAI hasn't released a GPT-6 Terra. The GPT-6 family is Astra, Sol, and Luna, and GPT-5.6 Terra is still listed in OpenAI's model catalog as the option that balances intelligence and cost.
Which ChatGPT plans include GPT-6 Sol and Luna?
GPT-6 Sol and Luna are in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can use Luna in the desktop app, and neither model is in Chat yet.
Does GPT-6 Sol support tool calling in Chat Completions?
Only with reasoning effort set to none. For tool calls at low effort or above, OpenAI requires the Responses API, and the same applies to GPT-6 Luna. Astra's model page doesn't list the same restriction.
Can I use GPT-6 Sol, Luna, and Astra in the same app?
Yes. Through the Respan gateway, all three sit behind one endpoint, so each call can name its own GPT-6 model and effort level, and every call's cost and trace land in one place. OpenAI's API also accepts any of the three on each request, since they share the same endpoints.



