Back to blog
News·5 min read·973 words

DeepSeek Just Crossed $1 Billion in Annual Revenue — and It's Still the Cheapest Frontier Model You Can Call

DeepSeek hit $1B in annual recurring revenue — proof the open-weights, efficiency-first model works commercially. What it means for your inference bill.

DeepSeek Just Crossed $1 Billion in Annual Revenue — and It's Still the Cheapest Frontier Model You Can Call — illustration

DeepSeek Just Crossed $1 Billion in Annual Revenue — and It's Still the Cheapest Frontier Model You Can Call

DeepSeek, the Chinese AI lab that shook the industry in early 2025 with its impossibly cheap training runs, has quietly hit a milestone nobody was watching for: $1 billion in annual recurring revenue. That makes it one of the only open-weights labs on the planet with a real, self-sustaining business — and it changes the math for every developer still paying premium prices for inference.

What actually happened

Per TLDR AI's September 25 briefing, DeepSeek crossed $1B ARR, a surge driven by API usage and the V4 model family. The timing is notable: the news landed the same week OpenAI introduced a premium "Pro Max" tier for ChatGPT and Meta launched real-time avatar "digital humans" for Muse. The industry is splitting in two directions at once — proprietary labs pushing premium tiers up, and open-weights labs pushing efficiency down.

DeepSeek is firmly in the second camp, and the revenue number proves the efficiency thesis works commercially, not just on leaderboards.

Why $1B ARR from an open-weights lab matters

Open weights and revenue sound contradictory. If you give the weights away, how do you make a billion dollars? Three ways, and DeepSeek does all of them:

  • Hosted API at scale. Cheap per-token prices attract enormous volume. When your marginal cost per token is a fraction of a frontier closed model, volume does the heavy lifting.
  • Efficiency as a feature. The V4 family — especially V4 Flash — is built for high-volume, low-cost inference. Developers running millions of requests a day care about cost per completed task more than a two-point benchmark delta.
  • Serving derivatives and enterprise deployments. Companies that want the model but not the ops burden pay for managed access.

The contrast with the pricing war above it is stark. GPT-6 Sol landed at $2 per million input tokens and Claude Opus 5.5 at $4/$20 — big cuts by closed-lab standards. DeepSeek V4 Flash-style pricing sits orders of magnitude lower.

What it costs on the open market right now

Here's where it gets interesting for builders. Pricing differs wildly depending on where you buy inference — and this is the part most developers never check.

On Qubax, where compute providers compete on an open market and prices track actual wholesale cost, DeepSeek V4 Pro runs about $0.08 per million input tokens and $0.16 per million output — roughly 79% below the standard retail benchmark of ~$0.38/$0.76. DeepSeek V4.1 Flash is even cheaper: around $0.016/$0.063, versus $0.35/$0.29 at typical retail. Same weights, radically different invoice.

The lesson isn't "DeepSeek is cheap." It's that the price of the same model can vary 5–20x depending on who's serving it — and most developers are still paying the top of that range by default.

The broader picture: the efficiency race is real

DeepSeek's milestone lands in a week full of signals that cheap-and-good is winning:

  • OpenRouter just crossed 10 million developers, and its CEO described routing between models as "critical infrastructure" — developers increasingly treat models as interchangeable commodities and shop on price.
  • GLM-5.3's open weights are out on Hugging Face, claiming open-source state of the art on Terminal Bench 3.0 — the open-weights frontier keeps closing on closed labs.
  • MiniMax M3-class models now cost fractions of a cent per million tokens while scoring respectably on agentic benchmarks.

When the second-cheapest option is nearly as good as the most expensive one, revenue flows to whoever can serve tokens at the lowest cost. That's DeepSeek's entire playbook.

What this means for developers

  1. Audit your inference bill this week. If you're calling a flagship model for tasks like summarization, extraction, classification, or simple chat, you're almost certainly overpaying. A V4-Flash-class model handles most of that work at 1–5% of the cost.
  2. Route, don't hardcode. Point easy tasks at a cheap model, hard ones at a flagship. Even a simple two-tier router typically cuts bills 60–80%.
  3. Buy on the open market. Marketplaces that pass wholesale pricing through — rather than fixed retail markups — are where the DeepSeek price gap actually reaches your invoice. Check the live rates on Qubax's model index before you commit to a single provider.

Will DeepSeek keep growing?

The open question is whether revenue follows efficiency or exclusivity. Closed labs are betting on premium tiers and frontier-only capabilities (GPT-6 Astra's "Critical" cyber rating, for example). DeepSeek is betting that 95% of real-world workloads don't need the frontier — just good-enough intelligence at a price that scales. A billion dollars of ARR suggests the second bet is paying off faster than anyone expected.

Want to see what DeepSeek actually costs with real open-market pricing? Browse live rates for DeepSeek, GLM, Kimi, and every frontier model on Qubax AI — what you pay tracks what inference actually costs, not what a price list says it should.

FAQ

How much revenue does DeepSeek make?

DeepSeek crossed $1 billion in annual recurring revenue as of late September 2026, per TLDR AI — a first for an open-weights-first lab at this scale.

Is DeepSeek still cheaper than GPT or Claude?

By a wide margin. DeepSeek's V4 family costs a small fraction of flagship GPT-6 or Claude pricing, and on open marketplaces like Qubax the gap widens further because prices track wholesale cost.

Can I use DeepSeek models commercially?

Yes. DeepSeek's recent releases ship under permissive licenses allowing commercial use, and they're available via API from DeepSeek directly or through marketplaces like Qubax.

What's the best DeepSeek model for production?

For high-volume, cost-sensitive workloads, DeepSeek V4 Flash or V4.1 Flash are the workhorses. For harder reasoning and agentic tasks, V4 Pro offers stronger quality while staying far cheaper than frontier closed models.

🤖

Try Claude Opus 5 on Qubax

Anthropic's most powerful model. Up to 49% off.

View pricing

Article tags

#deepseek#open weights#ai industry#ai pricing#news
Share:Post on XTelegramLinkedInYHacker NewsReddit
Qubax AI

Qubax AI

AI Models at up to 99% off · Pay with crypto

Reading about Claude Opus 5 and DeepSeek? Access them — plus 400+ other models — through one API. Anthropic's most powerful model. Up to 49% off.

Related articles