

Gemini 4 Argon matches GPT-6 Astra on intelligence and costs less per token, but it uses more than twice the output tokens per task, which decides the cheaper model once the launch discount ends. Opus 5.5 scores higher than both and costs the most per task.

GPT-6.1 Sol vs Luna vs Astra comes down to tier and effort level. Compare pricing and benchmarks, and see which model to use for each step of your agent.
Dylan Cable · 2 days ago
Compare OpenCode vs Claude Code on models, pricing, features, and performance, and see how to trace both coding agents on your own codebase.
Dylan Cable · 4 days ago
Grok 4.7 is the newest Grok model for coding and knowledge work. Compare Grok 4.7 vs Claude Opus 5.5 on benchmarks, pricing, speed, and cost per task.
Dylan Cable · September 24, 2026OpenAI released GPT-6 Astra on September 3, 2026, and Google answered with Gemini 4 Argon on September 30. Anthropic shipped Opus 5.5 in between, which means all three major labs put a new top-tier model on the market within a single month.
Frontier intelligence has become the main thing these labs compete on. Agents now run long, multi-step tasks, and a model that reasons slightly better at each step makes fewer errors that compound across the whole run, so the top of the capability range is where a new release changes what an agent can be trusted to do.
At that level, though, the scores have started to converge. Argon and Astra tie on independent intelligence testing, and Opus 5.5 leads both by a few points, so the differences that decide which model you run show up in cost per task and in how often each model answers wrong.
So Gemini 4 Argon vs GPT-6 Astra vs Opus 5.5 comes down to pricing and benchmarks, measured per task instead of per token.
Gemini 4 Argon ties GPT-6 Astra at 53 on the Artificial Analysis Intelligence Index, and Opus 5.5 scores 58. Argon hallucinates far less often than Astra, but it also answers correctly less often and spends more than twice Astra's output tokens on an average task.
On price, Argon undercuts Astra per task only while its launch discount lasts. At list price it lands between Astra and Opus 5.5, and it isn't broadly available through the API yet.
| Gemini 4 Argon | GPT-6 Astra | Opus 5.5 | |
|---|---|---|---|
| List price (per 1M) | $4 / $20 | $10 / $50 | $4 / $20 |
| Launch price (per 1M) | $2 / $10 | N/A | N/A |
| Intelligence Index | 53 (high) | 53 (max) | 58 (max) |
| Cost per task | $3.98 ($1.99 launch) | $3.26 | $5.98 |
| Output tokens per task | 62k | 27k | Not published |
| Hallucination rate | 15% | 51% | Not published |
| Terminal Bench 4 | 57% | 59% | 60% |
| Context window | 1M | 1.05M | 1M |
| API access | Limited rollout | Available | Available |
Gemini 4 Argon is Google's new frontier model. It accepts text, image, video, and speech input, returns text, and has a 1M-token context window.
The headline spec change is output length. In its Gemini 4 Argon announcement, Google raised the output limit to 1M tokens, up from 64K, so a single response can now carry a full-file rewrite or a long report without being split across calls.
Google announced Gemini 4 Argon on September 30, 2026. The rollout starts with trusted cyber defenders in Google's Fairwind Program, and paid API customers and Google AI Ultra subscribers come after that.
Google hasn't put a date on the second stage. Until it does, Argon is a model to evaluate and plan around rather than one you can deploy today.
Independent testing from Artificial Analysis puts Gemini 4 Argon at 53 on its Intelligence Index at high effort, 23 points above Gemini 3.1 Pro Preview. That's a large jump in one generation, and it reopens the Gemini vs ChatGPT question at the top tier, where Google is now level with OpenAI's frontier model.
On agentic work, Argon scores 78% on AutomationBench-AA, seven points ahead of Claude Sonnet 5.5 at 71%.
Argon launched at an introductory price of $2 per million input tokens and $10 per million output tokens. After the introductory period, the price rises to $4/$20, and Google hasn't said when that switch happens.
Cached input tokens are 95% off the input price, which works out to $0.10 per million at the launch price and $0.20 at list. Output is where the bill can grow, though, since the 1M output limit means a single maxed-out response would cost $20 at list.
Per-token prices only cover part of the cost with reasoning models, because each model decides how many tokens to spend before it answers. Cost per task captures both, by measuring what each model costs to complete an average Intelligence Index task.
Gemini 4 Argon
At $4/$20 list, Argon matches Opus 5.5 per token and costs well under Astra. Its token count changes that math: Argon averages 62k output tokens per task, so it costs $3.98 per task at list, or $1.99 during the launch period.
GPT-6 Astra
Astra's list price is $10/$50, two and a half times Argon's per-token rate, with cached input at $1.00 per million. Because it averages 27k output tokens per task, Astra finishes at $3.26 per task, 72 cents under Argon at list.
Opus 5.5
Opus 5.5 is also $4/$20, with cache reads at $0.20 per million. At max effort it costs $5.98 per task, $2.00 more than Argon at list and $2.72 more than Astra.
If your workload runs close to these averages, Argon comes out cheaper than Astra only while the launch price lasts. Workloads with short, fixed-length outputs may narrow that gap, since Argon's lower per-token price counts for more when every model writes less.
Route GPT-6 Astra and Opus 5.5 through one API
Compare frontier models side by side on your own requests, with cost and latency on every call. Try Respan for free.
Gemini 4 Argon
Argon scores 53 on the Intelligence Index at high effort, one point ahead of GPT-6.1 Sol at 52.
GPT-6 Astra
Astra reaches the same 53, measured at max effort.
Opus 5.5
Opus 5.5 scores 58 at max effort, five points above both.
A five-point gap on a composite index won't show up on every request. For the hardest reasoning steps in an agent, though, it's the one independent measure where these three separate, and it points to Opus 5.5.
Terminal Bench 4 tests agents on tasks completed in a command-line environment, which makes it a useful check for coding and DevOps agents.
Gemini 4 Argon
Argon scores 57%, two points behind Astra and a 53-point improvement over Gemini 3.1 Pro Preview.
GPT-6 Astra
Astra scores 59%.
Opus 5.5
Opus 5.5 scores 60%, one point ahead of Astra.
The three frontier models sit within three points of each other here, and Claude Sonnet 5.5 scores 64%, above all of them. Because results move this much from one agentic benchmark to the next, testing on your own workload will likely tell you more than any single score.
Artificial Analysis's AA-Omniscience benchmark tracks two numbers: how often a model answers correctly, and how often it gives a wrong answer instead of declining.
Gemini 4 Argon
Argon has a 15% hallucination rate and 50% accuracy. Lower accuracy alongside a low hallucination rate suggests Argon declines to answer more often instead of guessing.
GPT-6 Astra
Astra answers correctly more often, at 63%, but its hallucination rate is 51%, so it guesses far more often when it doesn't know.
Opus 5.5
Artificial Analysis hasn't published AA-Omniscience results for Opus 5.5.
Which profile fits depends on what a wrong answer costs you. A step that feeds an automated action or a customer-facing reply may be better off with a model that declines, while a step that needs an answer every time may suit Astra.
Gemini 4 Argon
Argon takes text, image, video, and speech input, with a 1M-token context window and a 1M-token output limit.
GPT-6 Astra
Astra has a 1.05M-token context window. At 27k output tokens per task, it writes less than half as much as Argon on the same work, which keeps output costs down on steps where you pay for every token the model writes.
Opus 5.5
Opus 5.5 has a 1M-token context window and accepts text and image input. It also offers a fast mode priced at $8/$40 per million tokens, for steps where response time matters more than cost.
None of the three wins across every axis, so the useful question is which model fits each step of your workload:
Running more than one of these in production means sending each step to a different model, and handling that split is the job of an LLM router. Whichever mix you pick, base it on cost per task measured on your own traffic rather than per-token prices.

Respan is an AI router that sends every model call through one API, with observability and evals built in. Route, observe, and evaluate every LLM call, so model choices come from measured cost and quality instead of a benchmark average.
GPT-6 Astra and Opus 5.5 are already on the gateway, and when a new frontier model is added, moving traffic to it is a one-word change.
Together, those give you a cost per task and a quality score for each model, measured on the requests your users actually send.
Route and compare frontier models in one place
Send every model call through one API, including new models as they reach the gateway, and compare them side by side on the same request. Try Respan for free.
Gemini 4 Argon and GPT-6 Astra tie at 53 on the Artificial Analysis Intelligence Index. Argon has a much lower hallucination rate (15% vs 51%), while Astra answers correctly more often (63% vs 50%) and uses fewer output tokens per task. Which one fits better depends on whether a step needs caution or coverage, and on whether you're paying Argon's launch price or its list price.
Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens during its introductory period, then $4/$20 after that. Cached input is 95% off. On an average Intelligence Index task, that comes to $1.99 at the launch price and $3.98 at list.
Not broadly yet. Google is rolling Gemini 4 Argon out to cyber defenders in its Fairwind Program first, with paid API customers and Google AI Ultra subscribers next, and no date for that stage. GPT-6 Astra and Opus 5.5 are available through the Respan gateway today.
Gemini 4 Argon has a 1M-token context window and a 1M-token output limit, up from 64K in the previous version.
Respan's Playground runs the same input through several models side by side, and Experiments run a model through your evaluators across a dataset built from production requests. Setting up that dataset now and scoring GPT-6 Astra and Opus 5.5 against it gives you a baseline that's ready to extend to Argon later.