Back to blog
News·7 min read·1272 words

Etched's Valuation Doubles to $21 Billion in a Month — Jane Street Leads Mega Round for AI Inference Hardware

Etched raised $700M at a $21B valuation led by Jane Street, doubling its worth in a month. Its custom prefill/decode inference silicon promises faster, cheaper AI — and downward pressure on API prices.

Etched's Valuation Doubles to $21 Billion in a Month — Jane Street Leads Mega Round for AI Inference Hardware — illustration

The AI hardware arms race just got a new front-runner. Etched, the inference-chip startup designing custom silicon to run frontier AI models faster and cheaper, announced on Tuesday that it has raised $700 million at a $21 billion valuation — doubling its worth in barely a month.

The round was led by Jane Street, the legendary quantitative trading firm, after it installed Etched's first shipped AI cluster system in its own data center and ran real workloads on it. Other investors include Kleiner Perkins, Sequoia Capital, Andreessen Horowitz, Peter Thiel, and Tiger Global.

Even by AI standards, the valuation step-up is jaw-dropping. Etched was valued at $5 billion in December 2025. It raised a $300 million Series C at a $10.3 billion valuation in July 2026. Now investors have nearly doubled that figure again — an $11 billion jump in about 30 days.

Why Investors Are So Excited

Etched delivers its AI technology as full systems it calls "frontier inference clusters" — what Nvidia would call "AI factories." But unlike generic GPU racks, Etched has designed two custom components from scratch to accelerate inference, the compute-heavy process that happens after a user submits a prompt.

"Inference is built in two stages," explained co-founder and COO Robert Wachen: prefill and decode.

  • Prefill phase — mathematically and compute-intensive. The system must understand the prompt, including all context. Etched built a prefill chip that operates at low voltage, allowing it to pack in more transistors without the typical heat problems of high-end AI chips. More transistors means more tokens processed per second.
  • Decode phase — memory-intensive. The system generates output tokens — the actual answer the user sees. Etched created a new type of memory and interconnect it calls cluster-scale memory, which lets many chips connect together and share a single memory pool at very high speed and low latency.

The result, Etched promises, is higher throughput and lower per-token cost — the two metrics that matter most to anyone running production AI workloads.

The Perception Problem Etched Is Still Fighting

Etched is still battling the perception from its early days that it "etches" a particular model into its chips — meaning each chip is custom-designed to run one frontier model and nothing else. That was the original intention. It is no longer the case.

Etched's systems can run any frontier model. The company has pivoted from a single-model ASIC strategy to a programmable inference platform, and the Jane Street deployment — running real quant workloads on a rack in a real data center — is the proof point investors needed.

Jane Street said in the blog post announcing the round: "We tested the chip and are pleased with the early results. Etched's unique approach to inference delivers the precision we will need to support our most demanding workloads. We're excited to now have our own rack running in our datacenter."

For a firm whose entire business depends on microseconds and exactness, that endorsement carries weight.

What This Means for the AI Inference Market

The Etched raise signals that investors are no longer content to bet only on the model layer — the GPT-5s, Claude Opuses, and Gemini Pros of the world. The infrastructure that runs those models is becoming a market in its own right, and it is growing faster than the models themselves.

Consider the economics. A frontier model might cost $2–15 per million output tokens at retail. If you are an API provider serving billions of tokens per day, even a 20% reduction in inference cost flows straight to the bottom line. Etched's pitch — purpose-built silicon that slashes the cost of both prefill and decode — targets exactly that gap.

This is also why inference cost has become the dominant concern for developers building AI applications. The question is no longer "which model is smartest?" but "which model gives me the intelligence I need at a price I can sustain?"

The Bigger Picture: Inference Is the New Frontier

For most of the AI boom, training was the bottleneck. Companies raced to build bigger clusters to train bigger models. That is still happening — but as models have matured, the cost center has shifted to inference. Running a 2-trillion-parameter model for millions of users, 24/7, is where the real compute bill lives.

Etched is not alone in targeting this. Nvidia's Blackwell and Rubin architectures, Cerebras's wafer-scale chips, Groq's LPU, and SambaNova's RDU all attack inference from different angles. What makes Etched notable is the speed of its commercial validation — a real quant fund buying a real rack and leading a real round at a $21 billion valuation, all within months of shipping.

What Developers Should Watch

If you are building on AI APIs today, the Etched story is more than a funding headline. It is a signal that inference costs are on a downward trajectory — and that the platforms that pass those savings to developers will win.

  • Watch per-token pricing trends. As custom inference silicon comes online, expect downward pressure on API pricing across all providers. Models that cost $15/M output tokens today may cost a fraction of that within a year.
  • Watch latency. Etched's cluster-scale memory is designed to reduce decode latency. If it delivers, expect sub-second response times on frontier models to become the norm, not the exception.
  • Watch the provider landscape. Cheaper inference benefits API aggregators and gateways that can route across multiple models and pass savings through. Platforms like Qubax AI that offer access to dozens of frontier models at competitive rates stand to gain as the underlying compute gets cheaper.

The Takeaway

Etched's $21 billion valuation is a bet that inference — not training — is where the next trillion dollars of AI infrastructure spending will flow. Whether the company can deliver on its speed and cost promises at scale remains to be seen, but Jane Street putting its own data center where its money is suggests the early results are real.

For developers, the message is simpler: the cost of intelligence is falling, and the tools to access it are getting better. The best time to build on AI APIs was a year ago. The second-best time is now.


Want to compare inference costs across frontier models with real, live pricing? Check out [Qubax AI's model catalog](https://qubax.ai/models) — 370+ models, transparent pricing, one API.

FAQ

What does Etched do?

Etched designs custom silicon and full inference cluster systems optimized for running frontier AI models. Its chips target the two phases of inference — prefill (compute-heavy) and decode (memory-heavy) — with purpose-built hardware for each.

Why is Jane Street leading the round?

Jane Street installed Etched's first shipped AI cluster in its own data center, tested it on real quantitative workloads, and was satisfied enough to lead the $700 million round at a $21 billion valuation.

How fast did Etched's valuation grow?

Etched went from a $5 billion valuation in December 2025 to $10.3 billion in July 2026 to $21 billion in August 2026 — roughly quadrupling in eight months.

What is the difference between training and inference?

Training is the process of building a model by feeding it data. Inference is the process of using that trained model to generate outputs. As models have grown larger, inference has become the dominant cost for AI providers.

How does this affect developers using AI APIs?

Cheaper inference hardware puts downward pressure on per-token API pricing. Developers can expect lower costs and lower latency from frontier models as companies like Etched, Nvidia, Groq, and Cerebras compete on inference efficiency.

Article tags

#etched#ai-hardware#inference#jane-street#funding
Share:Post on XTelegramLinkedInYHacker NewsReddit
Qubax AI

Qubax AI

AI Models at up to 99% off · Pay with crypto

Access GPT, Claude, Gemini, GLM & 340+ models through one OpenAI-compatible API. Up to 99% off. Pay with 200+ cryptocurrencies. No credit card needed.

Related articles