Google just put a price tag on Gemini 4 Argon, its newest frontier model — and the numbers deserve a closer look than the announcement headlines suggest. An introductory rate of $2 per million input tokens and $10 per million output tokens sounds aggressive for a frontier-tier model. It is. But it is also an introductory rate: once the promo window closes, Argon reverts to $4 / $20 per million tokens. And once you put that number next to what capable models actually cost on an open market today, the picture gets more interesting.
We pulled live per-token prices from the Qubax price index (374 models, updated continuously) and compared them against the OpenRouter benchmark price that every listing on our index carries. Here is what Gemini 4 Argon's launch price means for anyone shipping on an API right now.
What Google announced
Per Google's official announcement, Gemini 4 Argon is built for "deep reasoning across complex, long-horizon workflows" — software engineering, legal and financial analysis, and cybersecurity defense. Three facts matter for developers:
- Introductory price: $2 per 1M input tokens, $10 per 1M output tokens, with cached input tokens at a 95% discount.
- Post-intro price: $4 per 1M input and $20 per 1M output — double the intro rate on input, double on output.
- Output limit: 1 million tokens, up from 64,000, so a single trajectory can generate hundreds of thousands of tokens.
That last point quietly changes the cost math. If your agent routinely generates 200K output tokens per run, the output rate dominates your bill, and the difference between $10 and $20 per 1M output tokens is the difference between $2.00 and $4.00 per run — per call.
Rollout is phased: Argon is initially available to trusted cyber defenders through Google's Fairwind Program, with broader paid-API access to follow. So most developers will be pricing it before they can use it — which is exactly the right time to compare it against what is already on the market.
Argon vs the current frontier price table
Here is the same capability tier, priced, from our live index. The Qubax column is the sell price on our open market where compute providers compete on price; the OpenRouter column is the benchmark price for the same model. (Prices as of October 2, 2026, source: qubax.ai/price-index.)
| Model | Qubax in $/1M | Qubax out $/1M | OpenRouter in $/1M | OpenRouter out $/1M |
|---|---|---|---|---|
| GPT-5.6 Sol | $0.30 | $1.50 | $1.00 | $5.00 |
| Gemini 3.1 Pro | $2.70 | $16.20 | $1.80 | $10.80 |
| Claude Opus 5.5 | $3.02 | $12.08 | $4.00 | $20.00 |
| Claude Opus 5 | $6.75 | $27.01 | $5.00 | $25.00 |
| GLM 5.3 | $0.087 | $0.27 | $0.12 | $0.49 |
| Gemini 3.8 Flash | $0.34 | $1.35 | $0.38 | $1.88 |
| GPT-5.6 Luna | $0.045 | $0.27 | $0.10 | $0.60 |
Against this table, Argon's $2/$10 intro rate lands in familiar territory: cheaper than Claude Opus-tier pricing at list, roughly in line with Gemini 3.1 Pro, well above GPT-5.6 Sol. The post-intro $4/$20 rate, however, would make it one of the most expensive mainstream frontier models on the market — at a moment when the rest of the price table is moving the other way.
The output-token trap
Frontier comparisons usually fixate on input price, but long-horizon agents invert that. A model with a 1M-token output limit is designed for workloads where output dwarfs input — multi-hour coding trajectories, full legal document generation, chained agentic runs. On those workloads:
- Argon at intro rates: a 500K-output run costs $5.00.
- Argon at post-intro rates: the same run costs $10.00.
- GLM 5.3 at Qubax pricing: the same run costs about $0.14.
That is a 70x spread between the most and least expensive way to get the same class of long-form work done. The same pattern showed up when we compared GPT-6 Sol against Claude Opus 5 on agentic coding: output-heavy agent workloads punish list-price thinking fastest. The gap between Argon's intro and post-intro rates alone (2x) is bigger than most quarter-over-quarter price movements elsewhere in the industry.
What open-market pricing is doing to launch prices
There is a second-order effect worth naming. Launch prices like Argon's are set against a benchmark — and that benchmark is itself under pressure. On our index, the discount to the OpenRouter benchmark price across the models above ranges from modest (Claude Opus 5.5 at ~24% off on input) to extreme (Kimi K2 Thinking at ~98% off, Qwen3 32B at ~98% off). Compute providers competing on an open market are pricing capable open-weight models at fractions of list:
| Model | Qubax in $/1M | Qubax out $/1M | OpenRouter in $/1M | Effective discount (in) |
|---|---|---|---|---|
| Kimi K2 Thinking | $0.009 | $0.038 | $0.60 | ~98% |
| GLM 4.6 | $0.0065 | $0.026 | $0.43 | ~98% |
| Qwen3 32B | $0.0012 | $0.0042 | $0.08 | ~98% |
| DeepSeek V4 Pro | $0.094 | $0.188 | $0.177 | below benchmark* |
| MiniMax M3 | $0.0045 | $0.018 | $0.23 | ~98% |
*DeepSeek V4 Pro's Qubax sell price reflects wholesale competition among providers; its OpenRouter benchmark reflects a different listing tier. (Prices as of October 2, 2026, source: qubax.ai/price-index.)
The takeaway for pricing strategy: when open-weight models with near-frontier benchmark scores trade at hundredths of a dollar per million tokens, a frontier lab's "aggressive" launch price is aggressive only relative to its own history — not relative to the market a developer can actually buy into today. This is the same dynamic we tracked when OpenAI and Anthropic's list prices collided with cheaper Chinese rivals earlier this year: the price war didn't end, it moved downstream.
Practical guidance if you're evaluating Argon
- Price the post-intro rate, not the intro rate. Google's own announcement says $4/$20 applies after the introductory period. Any migration or architecture decision should be priced at $4/$20.
- Model your output tokens first. Argon's 1M output limit targets exactly the workloads where output rate dominates. Estimate tokens-out per run before comparing on input price.
- Keep a routing fallback. For the 80% of requests that don't need frontier-level long-horizon reasoning, a $0.02–$0.30 tier model (GLM 5.3 Flash, GPT-5.6 Luna, Gemini 3.7 Flash class) produces a blended cost that frontier-only pricing cannot match. You can check per-model pricing on the model directory, and the quickstart covers swapping a model ID in one line.
- Watch the benchmark column, not just the list price. The same model can trade at very different prices depending on where and how you buy it. Our live price index tracks the Qubax sell price against the OpenRouter benchmark for all 374 models, updated continuously.
If you want to see this pricing live rather than in a table, you can create a free Qubax account and query any model in the index through a single OpenAI-compatible endpoint — the price you see is the price you pay, per token, with no subscription layer.
FAQ
What is Gemini 4 Argon's API price?
Gemini 4 Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, with cached input at 95% off. After the introductory period, pricing rises to $4 per million input and $20 per million output.
How does Gemini 4 Argon pricing compare to Claude and GPT?
At intro rates, Argon ($2/$10) is cheaper than Claude Opus 5 ($5/$25 on the OpenRouter benchmark) and comparable to GPT-5.6-class frontier pricing. At its post-intro rate of $4/$20, it moves toward the top of the frontier price table.
When can developers access Gemini 4 Argon?
Google is rolling Argon out in phases, starting with trusted cyber defenders through its Fairwind Program, followed by paid API customers and Google AI Ultra subscribers. No public release date has been given yet.
What is the cheapest frontier-class model per million tokens right now?
Among models with near-frontier benchmark scores, open-weight options like GLM 5.3 and Kimi K2 Thinking trade well under $0.10 per million input tokens on Qubax's open market — roughly 40x to 100x below Argon's intro price, depending on the workload.
Does Qubax offer Gemini models?
Yes. Qubax carries the Gemini line including Gemini 3.8 Flash and Gemini 3.1 Pro, alongside models from OpenAI, Anthropic, DeepSeek, Moonshot, Zhipu and others, all priced per token on a single OpenAI-compatible API.