Feeding an entire book, contract, or codebase to an AI model used to mean choosing between expensive long-context models or building a retrieval system. That trade-off is gone: several frontier-class models now accept 1,000,000+ tokens in a single request, and the cheapest of them cost less than a coffee per billion tokens processed.
In this tutorial you'll build a complete long-document Q&A pipeline in Python using the OpenAI-compatible API — no vector database, no embeddings, no chunking library. Just map-reduce over a single context window, with cost tracking at every step.
All prices below are live API prices as of October 7, 2026, source: [qubax.ai/price-index](https://qubax.ai/price-index).
What a million tokens actually costs
First, the data. Here are the cheapest models that accept ≥1M-token contexts right now, with their price per million tokens and what the same tokens cost on OpenRouter retail:
| Model | Context | Input $/M | Output $/M | OpenRouter input $/M |
|---|---|---|---|---|
| MiniMax M2 | 1,000,000 | $0.0038 | $0.0151 | $0.26 |
| DeepSeek V4.1 Flash (0731) | 1,048,576 | $0.0156 | $0.2169 | $0.18 |
| GLM 5.2 | 1,048,576 | $0.0197 | $0.0788 | $0.09 |
| Grok 4.20 Beta | 2,000,000 | $0.0626 | $0.2505 | $1.25 |
| Claude Opus 4.6 | 1,000,000 | $3.5577 | $14.2308 | $5.00 |
Two things worth noticing:
- The spread is enormous. MiniMax M2 and Claude Opus 4.6 both accept 1M tokens, but Opus costs 936× more per input token. For document Q&A — where you're mostly paying for input — the cheap end wins on pure economics.
- Real traffic confirms it. On Qubax's own API in the last 7 days, DeepSeek V4.1 Flash alone processed 8.3 billion input tokens across 45,000+ requests — by far the most of any model. Developers are already pushing long-context workloads to the budget tier.
For this tutorial we'll use DeepSeek V4.1 Flash: 1M context, near-bottom pricing, and the highest real-world long-context traffic of any model right now.
The architecture: map-reduce, not RAG
Long-document Q&A has two standard approaches:
- RAG: chunk the document, embed chunks, retrieve the top-k relevant ones, answer from those.
- Long-context map-reduce: put the whole document in the context window and let the model read everything.
RAG is still right when your document exceeds the context window or when latency matters more than recall. But when the document fits (and 1M tokens ≈ 750,000 words ≈ five normal novels), single-context beats RAG on quality — the model sees cross-references between chapter 2 and chapter 47 that any retrieval system would miss.
The map-reduce pattern we'll build:
- Map: split the document into a few large segments (we use ~300K tokens each, well inside the limit) and extract everything relevant to the question from each segment.
- Reduce: combine the per-segment findings into one final answer with citations.
This is more reliable than a single mega-prompt for two reasons: attention degrades over very long inputs (models recall the beginning and end better than the middle — the "lost in the middle" problem documented by Stanford researchers), and per-segment extraction keeps each request's output small and cheap.
Step 1: Setup and cost-aware client
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["QUBAX_API_KEY"],
base_url="https://api.qubax.ai/v1",
)
MODEL = "deepseek-v4-flash-0731"
# Live prices (Oct 7, 2026): $0.0156/M input, $0.2169/M output
PRICE_IN = 0.0156 / 1_000_000
PRICE_OUT = 0.2169 / 1_000_000
def cost_of(usage) -> float:
return usage.prompt_tokens * PRICE_IN + usage.completion_tokens * PRICE_OUTAny OpenAI SDK works (see the official OpenAI Python SDK docs — only the base_url changes). Token counting uses tiktoken, OpenAI's o200k_base tokenizer, which is a close match for how modern models segment text; exact billing always comes from the API's usage field.
Step 2: Token-aware segmentation
Don't split by characters — split by tokens, with overlap so no fact is cut in half:
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
def segment(text: str, max_tokens: int = 300_000, overlap: int = 2_000) -> list[str]:
"""Split text into overlapping token windows."""
ids = enc.encode(text)
segments = []
step = max_tokens - overlap
for start in range(0, len(ids), step):
window = ids[start : start + max_tokens]
segments.append(enc.decode(window))
if start + max_tokens >= len(ids):
break
return segmentsA 500-page document is roughly 375K tokens, so two segments of 300K with 2K overlap covers it. A 1M-token corpus becomes four segments.
Step 3: Map — extract per segment
def extract_from_segment(question: str, segment: str, index: int) -> str:
resp = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "system", "content":
"You are a research assistant. Extract EVERY passage, fact, figure, "
"and quote from the text that is relevant to the user's question. "
"If nothing is relevant, reply exactly: NO_RELEVANT_CONTENT. "
"Quote passages verbatim where possible."},
{"role": "user", "content":
f"Question: {question}\n\nDocument segment {index}:\n\n{segment}"},
],
temperature=0,
)
return resp.choices[0].message.contentStep 4: Reduce — synthesize the final answer
def answer_question(question: str, document: str) -> dict:
segs = segment(document)
findings = []
total_cost = 0.0
for i, seg in enumerate(segs):
out = extract_from_segment(question, seg, i)
u = out # each call returns text; usage captured below
if out.strip() != "NO_RELEVANT_CONTENT":
findings.append(f"[Segment {i+1}]\n{out}")
if not findings:
return {"answer": "The document does not address this question.", "cost": total_cost}
combined = "\n\n---\n\n".join(findings)
reduce_resp = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "system", "content":
"Synthesize a single, complete answer to the question from the "
"extracted findings below. Cite segments as [S1], [S2]. "
"Resolve contradictions explicitly."},
{"role": "user", "content": f"Question: {question}\n\nFindings:\n{combined}"},
],
temperature=0,
)
# track cost across all calls
total_cost = len(enc.encode(document)) * PRICE_IN * 1.05 # map inputs + 5% overlap
total_cost += reduce_resp.usage.prompt_tokens * PRICE_IN
total_cost += reduce_resp.usage.completion_tokens * PRICE_OUT
return {
"answer": reduce_resp.choices[0].message.content,
"segments_searched": len(segs),
"estimated_cost_usd": round(total_cost, 6),
}Step 5: Run it
document = open("quarterly_report.pdf.txt").read() # 500 pages ≈ 375K tokens
result = answer_question(
"What are all the contingent liabilities mentioned, and in which sections?",
document,
)
print(result["answer"])
print(f"Searched {result['segments_searched']} segments · cost ≈ ${result['estimated_cost_usd']}")For a 375K-token document on DeepSeek V4.1 Flash, the map pass costs about $0.006. The whole query — map plus reduce — lands under one cent. The same query against Claude Opus 4.6's 1M context would run about $1.40 in input alone, and against MiniMax M2 about $0.0015 — all three fit the document; the budget tier just makes experimentation free.
Practical notes from running this in production
- Watch for the middle-blur. Even with 1M context, models recall segment edges best. If a fact you know is in the document doesn't surface, try splitting into more, smaller segments.
- Cache the segmentation.
tiktokenencoding of a large document takes a few seconds. Cache segments by document hash so repeat queries skip straight to the map step. - Parallelize the map. Segments are independent — run
extract_from_segmentcalls concurrently (e.g. withasyncio+AsyncOpenAI). Four segments in parallel cuts wall-clock time ~4×. - Log real usage, not estimates. The API returns exact
usageper call; sum it rather than estimating from character counts. Our estimate above is deliberately conservative (105% of raw tokens).
When to still use RAG
Single-context map-reduce wins when the document fits and answer quality matters. Switch to RAG when:
- The corpus exceeds ~1M tokens (multiple books, an entire codebase history)
- You need sub-second responses on a static corpus (retrieval + short prompt is faster than re-reading 375K tokens per query — even at $0.006/query, latency is ~30–60s)
- Many users ask many questions against the same corpus (cost scales with every query × full document; embeddings amortize)
A hybrid works well too: RAG for the fast path, and fall back to full-context map-reduce when retrieval confidence is low.
FAQ
How many pages is 1 million tokens?
Roughly 750,000 words, or about 1,500–2,500 PDF pages depending on density. DeepSeek V4.1 Flash accepts 1,048,576 tokens and Grok 4.20 Beta accepts 2,000,000.
Do I still need a vector database for long documents?
Not if the document fits in the context window. Vector search adds infrastructure and can miss cross-references. Use embeddings only when the corpus exceeds the context limit or you need low-latency repeated queries.
Which is the cheapest 1M-context model?
As of October 7, 2026, MiniMax M2 at $0.0038 per million input tokens and $0.0151 per million output tokens — roughly 936× cheaper than Claude Opus 4.6's 1M context on input. DeepSeek V4.1 Flash and GLM 5.2 are close behind with 1,048,576-token windows.
Why not just send the whole document in one request?
You can — and for documents under ~200K tokens it's often the simplest approach. Map-reduce over segments improves recall on very long inputs (models pay less attention to the middle of huge contexts) and keeps each request's output focused, which reduces cost and hallucination.
Does the OpenAI Python SDK work with these models?
Yes. All models in the table above are reachable through the standard OpenAI SDK — you only change the base_url and API key. The tutorial's code runs unchanged.
Check live prices for any model at the [Qubax price index](https://qubax.ai/price-index) — updated continuously from real API traffic. Or browse all 399+ models to find the right context window and price for your workload.