Comparison

Cheapest GPT API: Luna vs OSS 120B vs 4o Mini

The cheapest GPT API is not the model with the most traffic. Live prices for six GPT models versus OpenRouter, and what 1,000 typical requests actually cost.

Qubax AI12 min read
Cheapest GPT API: Luna vs OSS 120B vs 4o Mini — illustration
In this article 6 sections

The cheapest GPT API is not the model your traffic already uses. On the morning of October 5, 2026, GPT-6 Luna had 275,483 requests in the prior 24 hours, and its price sits only 6.4% under the OpenRouter benchmark. The models that win on price — gpt-oss-20b, GPT 5 Nano, and OpenAI GPT OSS 120B — cost a fraction of Luna on the same token mix, and two of them barely show up in traffic. If you searched for a cheapest GPT API number, the rest of this page is that number, with the workload math written out so you can recompute it.

Prices below are the latest row on the price index as of October 5, 2026. They are what you pay per million tokens. Six models are in the table. GPT-5.6 Luna is left out on purpose: its latest input price is $0.1587 per million, above the OpenRouter benchmark of $0.1000, so it fails a "cheaper than OpenRouter" test. Flagship Sol and Astra SKUs are a different budget and are not part of a cheapest-GPT question.

Six GPT prices, one morning

ModelModel idContextRequests, 24hInput $/MOutput $/MOpenRouter in $/MOpenRouter out $/MInput vs OpenRouter
gpt-oss-20bopenai-gpt-oss-20b128,000148$0.0003$0.0013$0.0180$0.090098.3% less
GPT 5 Nanogpt-5-nano400,000303$0.0038$0.0300$0.0250$0.200084.8% less
OpenAI GPT OSS 120Bopenai-gpt-oss-120b128,00018,013$0.0062$0.0246$0.0700$0.300091.1% less
GPT-4o Minigpt-4o-mini128,0002,683$0.0069$0.0274$0.1500$0.600095.4% less
GPT-6 Lunagpt-6-luna1,050,000275,483$0.0468$0.2339$0.0500$0.25006.4% less
GPT-5.4 Minigpt-5.4-mini400,0002,586$0.0562$0.3375$0.3750$2.250085.0% less

Prices as of October 5, 2026, source: qubax.ai/price-index. Figures are rounded to 4 decimal places from the latest price row. Unrounded, GPT-4o Mini is $0.006853 / $0.027413 and GPT-6 Luna is $0.046778 / $0.233888. OpenAI GPT OSS 120B's latest row was computed on October 1, 2026; the other five rows were computed on October 5. Request counts are the trailing 24 hours as of that morning, not a quality score. OpenRouter price means the benchmark stored for that model, not a live scrape of one provider on openrouter.ai.

Two things in that table are easy to miss.

First, "mini" is not a price rank. GPT-5.4 Mini is the most expensive model in the set, on both input and output. GPT-6 Luna is cheaper than GPT-5.4 Mini here, has a longer context window (1,050,000 vs 400,000), and had about 106 times the observed traffic (275,483 vs 2,586). If the reason you reached for a mini was cost, Luna beats this particular mini on the price index, and the open-weight rows beat both.

Second, percent off OpenRouter is not dollars saved. GPT-4o Mini is 95.4% under a $0.15 OpenRouter input price. gpt-oss-20b is 98.3% under a benchmark that is already $0.0180, so the dollar gap is smaller. The next section prices an actual request.

What 1,000 requests actually cost

List prices hide the mix. A typical app call is input-heavy: a system prompt, a few turns of history, and a short answer. The worked example is 1,000 requests of 4,000 input tokens and 800 output tokens. Cost = requests × (input tokens / 1,000,000 × input price + output tokens / 1,000,000 × output price). Using the 4-decimal prices in the table:

ModelCost, 1,000 requestsOpenRouter, same mixYou save vs OpenRouterCost, 1,000,000 requests
gpt-oss-20b$0.0022$0.144098.4%$2.24
GPT 5 Nano$0.0392$0.260084.9%$39.20
OpenAI GPT OSS 120B$0.0445$0.520091.4%$44.48
GPT-4o Mini$0.0495$1.080095.4%$49.52
GPT-6 Luna$0.3743$0.40006.4%$374.32
GPT-5.4 Mini$0.4948$3.300085.0%$494.80

On this mix, GPT-6 Luna costs 8.4 times OpenAI GPT OSS 120B ($0.3743 / $0.0445). GPT-4o Mini costs 1.1 times GPT OSS 120B ($0.0495 / $0.0445) — close enough that price is not why you would pick one over the other. GPT 5 Nano is slightly cheaper than GPT OSS 120B here ($0.0392 vs $0.0445) because the mix is input-heavy and Nano's input rate ($0.0038) is below OSS 120B's ($0.0062).

That ranking flips when the answer is as long as the prompt. At 1,000 input and 1,000 output tokens, 1,000 requests cost $0.0308 on GPT OSS 120B and $0.0338 on GPT 5 Nano. The crossover is mechanical. Set the two rates equal:

0.0062 × input + 0.0246 × output = 0.0038 × input + 0.0300 × output

which simplifies to input = 2.25 × output. Nano is cheaper only when input tokens exceed about 2.25 times output tokens. Below that ratio, OSS 120B's lower output rate wins. gpt-oss-20b is cheaper than both on either mix, because it is cheaper on input and on output.

An agent-shaped call — 20,000 input tokens and 1,500 output tokens — does not change the order. One thousand of those calls cost $0.0080 on gpt-oss-20b, $0.1210 on GPT 5 Nano, $0.1609 on GPT OSS 120B, $0.1791 on GPT-4o Mini, $1.2869 on GPT-6 Luna, and $1.6303 on GPT-5.4 Mini. Luna is still about 8 times OSS 120B ($1.2869 / $0.1609). The 6.4% gap versus OpenRouter on Luna is $0.0881 per thousand of those agent calls ($1.2869 vs $1.3750). That is real money at millions of calls, and it is not a reason to leave an open-weight model that already passes your eval.

Pull the same rows yourself on the price index before you lock a budget to a single morning's snapshot. When the eval passes, create a Qubax API key and send the model id from the table. The host and the request shape do not change when you swap ids.

OpenRouter is not the OpenAI list

The "you save" column is versus OpenRouter. It is not versus the price on OpenAI's own pricing page, and those two public numbers disagree for several of these models.

OpenAI's published list, checked October 5, 2026, shows short-context rates per million tokens of $0.10 input and $0.50 output for gpt-6-luna, $0.15 and $0.60 for gpt-4o-mini, $0.75 and $4.50 for gpt-5.4-mini, and $0.05 and $0.40 for gpt-5-nano (OpenAI API pricing). The same page lists a separate long-context rate for gpt-6-luna ($0.20 / $0.75) and a batch rate ($0.05 / $0.25 on short context). The price index stores one input rate and one output rate per model. This comparison does not assume a long-context surcharge on our side.

For GPT-4o Mini, the OpenRouter benchmark ($0.15 / $0.60) matches both OpenAI's published list and the price on the OpenRouter GPT-4o Mini page. The 95.4% gap is a real gap against both public lists.

For GPT-6 Luna, the OpenRouter benchmark ($0.05 / $0.25) matches OpenAI's batch rate, not the higher published list rate of $0.10 / $0.50. Against OpenRouter, Luna is 6.4% cheaper. Against that published list, the same $0.0468 / $0.2339 rate is 53.2% cheaper. Quote the comparison you actually mean. A write-up that says "53% off" while the OpenRouter column says 6.4% is mixing two benchmarks. This page uses OpenRouter for the savings column, and cites the OpenAI list separately so the two are not collapsed.

GPT OSS 120B is an open-weight model. The weights and model card are on Hugging Face. It does not appear as a line on OpenAI's published token table the way gpt-4o-mini does. Its OpenRouter benchmark on the price index is $0.0700 / $0.3000. Provider prices on OpenRouter move, so that benchmark is the comparison figure, not a claim that every OpenRouter provider quotes $0.07. The same caution applies to gpt-oss-20b, whose benchmark is $0.0180 / $0.0900.

The same method, applied to Claude, is the cheapest Claude API comparison: one morning, latest price row, OpenRouter as the benchmark name.

Which id to send

Pick on constraints, then on price. Price alone will always return gpt-oss-20b, and that is the correct answer to the literal question. It is a weak answer to "what should production call on Monday" if you have not run it.

Use gpt-oss-20b when the prompt fits in 128,000 tokens and you have already evaluated it. At $2.24 per million requests of the 4,000/800 mix, it is the price floor. 148 requests in 24 hours means you are early on that id, not that the discount is fake.

Use OpenAI GPT OSS 120B when you want the cheap model that already has volume. 18,013 requests in 24 hours, 128,000 context, $44.48 per million requests on the worked mix, 91.4% under its OpenRouter benchmark. It is the default in this set if 128,000 tokens is enough and you do not have an eval that specifically fails the open-weight model. It also wins against GPT 5 Nano whenever the answer is long relative to the prompt, which is the common case for summaries, drafts, and tool-call explanations.

Use GPT-4o Mini when you already depend on that model id and a migration is the expensive part. It is 11% more than GPT OSS 120B on the worked mix, and 95.4% under the $0.15 / $0.60 OpenRouter list. Against GPT OSS 120B at $0.0062 / $0.0246, that is a familiarity choice. Context is the same 128,000 tokens.

Use GPT-6 Luna when the job needs the 1,050,000-token window, or when your eval says the open-weight models fail and Luna passes. Do not use it because it is the popular GPT. Popular and cheap have come apart. 275,483 requests say a lot of callers have not made that distinction. At $374.32 per million of the worked requests, versus $44.48 on GPT OSS 120B, the gap is $329.84 per million requests. That is the cost of skipping the eval.

Skip GPT-5.4 Mini if the goal is cost. It is 32% more than GPT-6 Luna on the worked mix ($0.4948 / $0.3743), with a shorter context window and far less traffic. Its OpenRouter gap (85.0%) looks attractive next to Luna's 6.4% and still loses to Luna, Nano, GPT-4o Mini, and both OSS models on dollars per request. A large percent off a high benchmark is not a cheap API.

Skip GPT-5.6 Luna for a cost migration. Input at $0.1587 per million is above the $0.1000 OpenRouter benchmark, and output at $0.6346 is above the $0.6000 benchmark. A model that costs more than the benchmark is not a candidate in a cheapest-GPT shortlist, however many requests it served overnight.

Swap the model id, not the client

The request is an OpenAI chat completion. Setup, keys, and the error shape are in the quickstart. The only field this comparison asks you to change is model.

python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.qubax.ai/v1",
    api_key="YOUR_API_KEY",
)

# Price-floor id. Swap to openai-gpt-oss-120b or gpt-6-luna after the eval.
resp = client.chat.completions.create(
    model="openai-gpt-oss-20b",
    messages=[
        {"role": "user", "content": "List three risks in this refund policy, one line each."}
    ],
)
print(resp.choices[0].message.content)

Run that prompt, or your real one, against openai-gpt-oss-20b, openai-gpt-oss-120b, and gpt-6-luna. Keep the cheapest id that still passes. If 20b fails and 120b passes, the worked-mix penalty versus the floor is $42.24 per million requests ($44.48 minus $2.24), not the $329.84 jump to Luna. If both open-weight ids fail, Luna is the expensive answer you have evidence for, and the 6.4% OpenRouter gap is a small extra, not the decision.

Price is the easy half. Run the same prompts against gpt-oss-20b, OpenAI GPT OSS 120B, and GPT-6 Luna, then keep the cheapest pass. Start on Qubax and check the first invoice against the table above. If the invoice does not match, the token mix is different from 4,000/800, not the rate card.

FAQ

What is the cheapest GPT API right now?

gpt-oss-20b, model id openai-gpt-oss-20b, at $0.0003 per million input tokens and $0.0013 per million output tokens as of October 5, 2026. One thousand requests of 4,000 input and 800 output tokens cost $0.0022. It had 148 requests in the prior 24 hours, so treat it as a price floor with thin observed traffic, not as proof that a production route is already busy.

Is GPT-6 Luna the cheapest GPT model?

No. It is the GPT model in this set with the most traffic (275,483 requests in 24 hours) and the only one with a 1,050,000-token context window. On the 4,000/800 mix it costs $0.3743 per thousand requests, about 8.4 times OpenAI GPT OSS 120B, and only 6.4% under its OpenRouter benchmark of $0.05 input and $0.25 output per million tokens.

How do these prices compare with OpenAI's published list?

They are a different comparison from the OpenRouter column. GPT-4o Mini matches OpenAI's published list and the OpenRouter page at $0.15 / $0.60, so the 95.4% gap holds against both. GPT-6 Luna's published short-context rate is $0.10 / $0.50; the OpenRouter benchmark is $0.05 / $0.25, in line with OpenAI's batch rate. Against the published list, Luna's $0.0468 / $0.2339 is about 53% less. Against OpenRouter, it is 6.4% less. Say which one you mean.

When is GPT 5 Nano cheaper than GPT OSS 120B?

When input tokens are more than about 2.25 times output tokens. Nano's input rate ($0.0038) is below OSS 120B's ($0.0062), and its output rate ($0.0300) is above OSS 120B's ($0.0246). The 4,000/800 example is input-heavy, so Nano wins that one ($0.0392 vs $0.0445 per thousand requests). A 1,000/1,000 mix flips it ($0.0338 vs $0.0308). gpt-oss-20b is cheaper than both on either mix.

Which model id should the request send?

The slug: openai-gpt-oss-20b, gpt-5-nano, openai-gpt-oss-120b, gpt-4o-mini, gpt-6-luna, or gpt-5.4-mini. The host is https://api.qubax.ai/v1. Changing models is a string change, not a client change. Context limits still apply: the two OSS models and GPT-4o Mini are 128,000 tokens, GPT 5 Nano and GPT-5.4 Mini are 400,000, and GPT-6 Luna is 1,050,000.

Try GPT-5.6 on Qubax

OpenAI's latest. Up to 99% below OpenRouter.

View pricing

Article tags

#gpt#pricing#openai#comparison

Share this article

Qubax AI

Qubax AI

AI models up to 99% below OpenRouter · Pay with crypto

Reading about GPT-5.6 and GPT-5?

Access them — plus 400+ other models — through one API. OpenAI's latest. Up to 99% below OpenRouter.

Related articles

All articles →