Tutorial

Your LLM bill is 80% hidden thinking tokens. One parameter fixes it.

Three cheap models billed 800–2,500 output tokens for a 150-word answer. Most of it was hidden reasoning. Here's how to measure it and the one parameter that cut cost 69–91%.

Qubax AI5 min read
Your LLM bill is 80% hidden thinking tokens. One parameter fixes it. — illustration
In this article 6 sections

I asked three cheap, popular models the same simple question: "Write a 150-word explanation of how HTTP caching headers work, for a junior developer."

Each one answered in about 150 words. Each one billed me for 800 to 2,500 output tokens.

The difference is reasoning tokens: "thinking" the model does before it answers. You pay for every one of them at the output rate, they never appear in the response, and most dashboards don't break them out. For a task like this one, they are pure waste.

Here is how to measure them in your own app, and the one parameter that removed them for me.

The measurement

Every OpenAI-compatible API reports hidden tokens in usage.completion_tokens_details.reasoning_tokens. This script calls each model three times, first with default settings and then with reasoning_effort turned down, and prints the medians.

python
"""hidden_tokens.py: how much of your output bill is hidden thinking?"""
import os, statistics, time
from openai import OpenAI

client = OpenAI(base_url="https://api.qubax.ai/v1",
                api_key=os.environ["QUBAX_API_KEY"])
PROMPT = ("Write a 150-word explanation of how HTTP caching "
          "headers work, for a junior developer.")
# model -> (output $ per 1M tokens, lowest reasoning_effort it accepts)
MODELS = {
    "glm-5.3-flash": (0.0426, "low"),
    "deepseek-v4.1-flash": (0.058824, "none"),
    "claude-haiku-5.5": (0.340909, "low"),
}

def call(model, **extra):
    t0 = time.perf_counter()
    r = client.chat.completions.create(
        model=model, max_tokens=4000, temperature=0,
        messages=[{"role": "user", "content": PROMPT}], **extra)
    d = r.usage.completion_tokens_details
    hidden = (getattr(d, "reasoning_tokens", 0) or 0) if d else 0
    words = len((r.choices[0].message.content or "").split())
    return {"out": r.usage.completion_tokens, "hidden": hidden,
            "words": words, "s": time.perf_counter() - t0}

for model, (price, effort) in MODELS.items():
    for label, extra in [("default", {}),
                         (f"effort={effort}", {"reasoning_effort": effort})]:
        rows = [call(model, **extra) for _ in range(3)]
        med = lambda k: statistics.median(r[k] for r in rows)
        per_1k = med("out") * price / 1e6 * 1000
        print(f"{model:<22}{label:<14}{med('out'):>6.0f} out "
              f"{med('hidden'):>6.0f} hidden {med('words'):>4.0f} words "
              f"{med('s'):>5.1f}s  ${per_1k:.4f}/1k calls")

Change base_url and the key to point at any OpenAI-compatible endpoint. The model names above are the ones I tested.

The results (Oct 11, 2026, median of 3 runs)

ModelSettingBilled output tokensHiddenWords shownTimeCost per 1k calls
GLM 5.3 Flashdefault2,4832,25615239.6s$0.1058
GLM 5.3 Flashreasoning_effort="low"23201598.4s$0.0099
DeepSeek V4.1 Flashdefault1,5251,29315012.8s$0.0897
DeepSeek V4.1 Flashreasoning_effort="none"20001323.6s$0.0118
Claude Haiku 5.5default1,6521,3541507.4s$0.5632
Claude Haiku 5.5reasoning_effort="low"28901292.7s$0.0985

The same answer, 83–91% cheaper and 3–5× faster, from one parameter.

Hidden-token counts vary a lot between runs, so I ran the whole script again. On the second run, defaults were 74–90% hidden, and the setting still cut cost by 69–91%. The exact numbers move; the direction never did.

At default settings, 82–91% of what I paid for on that run was text I never saw. Latency follows the bill: GLM 5.3 Flash went from 40 seconds to 8 because it stopped writing a hidden essay before the real one.

Two things that surprised me

1. Values aren't portable. DeepSeek accepts "none" and goes to zero hidden tokens. On GLM, "low" already gave zero. On the same models, "minimal" sometimes kept a few hundred hidden tokens. Measure each model; don't assume.

2. Some models ignore it. In a separate run, GPT-5.5 kept 400–600 hidden tokens whatever value I sent. If a model doesn't respond to the knob, the fix is choosing a different model for that job, not tuning.

When to keep thinking on

Hidden reasoning isn't always waste. It earns its cost on multi-step math, tricky refactors, planning agents, and anything where a wrong answer is expensive. It's waste on:

  • classification and routing ("which queue does this ticket go to?")
  • extraction to JSON
  • summaries, rewrites, translations
  • short chat replies and UI copy

A simple rule that works: set `reasoning_effort` per call site, not per app. Your "summarize this email" helper and your "plan this database migration" agent shouldn't share a setting.

python
FAST = {"reasoning_effort": "low"}   # extraction, labels, summaries
DEEP = {}                            # planning, math, hard code

def summarize(text):
    return client.chat.completions.create(
        model="claude-haiku-5.5", **FAST,
        messages=[{"role": "user", "content": f"Summarize in 2 lines:\n{text}"}],
    ).choices[0].message.content

Check your own traffic

Log completion_tokens_details.reasoning_tokens next to completion_tokens for a day. If hidden tokens are more than half your output on endpoints that do simple work, you've found the cheapest optimization you'll make this year.

I ran these tests through Qubax, an OpenAI-compatible API with 400+ models (often well below OpenRouter prices) where you can switch models by changing one string, so the script runs unchanged on every model. Prices in the table are the per-token rates on Oct 11, 2026; live rates are on the price index.

FAQ

What are reasoning tokens?

They're tokens a model generates while "thinking" before it writes the visible answer. They're billed at the output-token price but never returned in the response text. You can see the count in usage.completion_tokens_details.reasoning_tokens.

Does lowering reasoning_effort make answers worse?

For simple tasks like summaries, extraction and labels, it made no visible difference in my tests. For multi-step math, planning or hard code, keep reasoning on and measure before switching it off.

Which reasoning_effort value should I use?

It depends on the model. In these tests, DeepSeek V4.1 Flash went to zero hidden tokens with "none", and GLM 5.3 Flash and Claude Haiku 5.5 with "low". Test each model you use.

References

  • OpenAI API reference, reasoning_effort and completion_tokens_details: https://platform.openai.com/docs/api-reference/chat/create
  • Anthropic extended thinking docs: https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking
  • DeepSeek API docs: https://api-docs.deepseek.com/

What's the worst hidden-token ratio you've found in production? I'd love to see numbers from other stacks.

Try Claude on Qubax

Anthropic models on Qubax. Up to 74% off.

View pricing

Article tags

#reasoning tokens#llm cost#python#reasoning_effort

Share this article

Qubax AI

Qubax AI

AI models up to 99% below OpenRouter · Pay with crypto

Reading about Claude and GLM 5?

Access them — plus 400+ other models — through one API. Anthropic models on Qubax. Up to 74% off.

Related articles

All articles →