I asked three cheap, popular models the same simple question: "Write a 150-word explanation of how HTTP caching headers work, for a junior developer."
Each one answered in about 150 words. Each one billed me for 800 to 2,500 output tokens.
The difference is reasoning tokens: "thinking" the model does before it answers. You pay for every one of them at the output rate, they never appear in the response, and most dashboards don't break them out. For a task like this one, they are pure waste.
Here is how to measure them in your own app, and the one parameter that removed them for me.
The measurement
Every OpenAI-compatible API reports hidden tokens in usage.completion_tokens_details.reasoning_tokens. This script calls each model three times, first with default settings and then with reasoning_effort turned down, and prints the medians.
"""hidden_tokens.py: how much of your output bill is hidden thinking?"""
import os, statistics, time
from openai import OpenAI
client = OpenAI(base_url="https://api.qubax.ai/v1",
api_key=os.environ["QUBAX_API_KEY"])
PROMPT = ("Write a 150-word explanation of how HTTP caching "
"headers work, for a junior developer.")
# model -> (output $ per 1M tokens, lowest reasoning_effort it accepts)
MODELS = {
"glm-5.3-flash": (0.0426, "low"),
"deepseek-v4.1-flash": (0.058824, "none"),
"claude-haiku-5.5": (0.340909, "low"),
}
def call(model, **extra):
t0 = time.perf_counter()
r = client.chat.completions.create(
model=model, max_tokens=4000, temperature=0,
messages=[{"role": "user", "content": PROMPT}], **extra)
d = r.usage.completion_tokens_details
hidden = (getattr(d, "reasoning_tokens", 0) or 0) if d else 0
words = len((r.choices[0].message.content or "").split())
return {"out": r.usage.completion_tokens, "hidden": hidden,
"words": words, "s": time.perf_counter() - t0}
for model, (price, effort) in MODELS.items():
for label, extra in [("default", {}),
(f"effort={effort}", {"reasoning_effort": effort})]:
rows = [call(model, **extra) for _ in range(3)]
med = lambda k: statistics.median(r[k] for r in rows)
per_1k = med("out") * price / 1e6 * 1000
print(f"{model:<22}{label:<14}{med('out'):>6.0f} out "
f"{med('hidden'):>6.0f} hidden {med('words'):>4.0f} words "
f"{med('s'):>5.1f}s ${per_1k:.4f}/1k calls")Change base_url and the key to point at any OpenAI-compatible endpoint. The model names above are the ones I tested.
The results (Oct 11, 2026, median of 3 runs)
| Model | Setting | Billed output tokens | Hidden | Words shown | Time | Cost per 1k calls |
|---|---|---|---|---|---|---|
| GLM 5.3 Flash | default | 2,483 | 2,256 | 152 | 39.6s | $0.1058 |
| GLM 5.3 Flash | reasoning_effort="low" | 232 | 0 | 159 | 8.4s | $0.0099 |
| DeepSeek V4.1 Flash | default | 1,525 | 1,293 | 150 | 12.8s | $0.0897 |
| DeepSeek V4.1 Flash | reasoning_effort="none" | 200 | 0 | 132 | 3.6s | $0.0118 |
| Claude Haiku 5.5 | default | 1,652 | 1,354 | 150 | 7.4s | $0.5632 |
| Claude Haiku 5.5 | reasoning_effort="low" | 289 | 0 | 129 | 2.7s | $0.0985 |
The same answer, 83–91% cheaper and 3–5× faster, from one parameter.
Hidden-token counts vary a lot between runs, so I ran the whole script again. On the second run, defaults were 74–90% hidden, and the setting still cut cost by 69–91%. The exact numbers move; the direction never did.
At default settings, 82–91% of what I paid for on that run was text I never saw. Latency follows the bill: GLM 5.3 Flash went from 40 seconds to 8 because it stopped writing a hidden essay before the real one.
Two things that surprised me
1. Values aren't portable. DeepSeek accepts "none" and goes to zero hidden tokens. On GLM, "low" already gave zero. On the same models, "minimal" sometimes kept a few hundred hidden tokens. Measure each model; don't assume.
2. Some models ignore it. In a separate run, GPT-5.5 kept 400–600 hidden tokens whatever value I sent. If a model doesn't respond to the knob, the fix is choosing a different model for that job, not tuning.
When to keep thinking on
Hidden reasoning isn't always waste. It earns its cost on multi-step math, tricky refactors, planning agents, and anything where a wrong answer is expensive. It's waste on:
- classification and routing ("which queue does this ticket go to?")
- extraction to JSON
- summaries, rewrites, translations
- short chat replies and UI copy
A simple rule that works: set `reasoning_effort` per call site, not per app. Your "summarize this email" helper and your "plan this database migration" agent shouldn't share a setting.
FAST = {"reasoning_effort": "low"} # extraction, labels, summaries
DEEP = {} # planning, math, hard code
def summarize(text):
return client.chat.completions.create(
model="claude-haiku-5.5", **FAST,
messages=[{"role": "user", "content": f"Summarize in 2 lines:\n{text}"}],
).choices[0].message.contentCheck your own traffic
Log completion_tokens_details.reasoning_tokens next to completion_tokens for a day. If hidden tokens are more than half your output on endpoints that do simple work, you've found the cheapest optimization you'll make this year.
I ran these tests through Qubax, an OpenAI-compatible API with 400+ models (often well below OpenRouter prices) where you can switch models by changing one string, so the script runs unchanged on every model. Prices in the table are the per-token rates on Oct 11, 2026; live rates are on the price index.
FAQ
What are reasoning tokens?
They're tokens a model generates while "thinking" before it writes the visible answer. They're billed at the output-token price but never returned in the response text. You can see the count in usage.completion_tokens_details.reasoning_tokens.
Does lowering reasoning_effort make answers worse?
For simple tasks like summaries, extraction and labels, it made no visible difference in my tests. For multi-step math, planning or hard code, keep reasoning on and measure before switching it off.
Which reasoning_effort value should I use?
It depends on the model. In these tests, DeepSeek V4.1 Flash went to zero hidden tokens with "none", and GLM 5.3 Flash and Claude Haiku 5.5 with "low". Test each model you use.
References
- OpenAI API reference,
reasoning_effortandcompletion_tokens_details: https://platform.openai.com/docs/api-reference/chat/create - Anthropic extended thinking docs: https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking
- DeepSeek API docs: https://api-docs.deepseek.com/
What's the worst hidden-token ratio you've found in production? I'd love to see numbers from other stacks.