Tutorial

How to Answer Questions About 500-Page Documents With a 1M-Token Model (Python Map-Reduce Tutorial)

Build a complete long-document Q&A pipeline in Python with map-reduce over 1M-token contexts. Live price data: MiniMax M2 at $0.0038/M vs Claude Opus 4.6 at $3.56/M for the same context window.

Qubax AI8 min read
How to Answer Questions About 500-Page Documents With a 1M-Token Model (Python Map-Reduce Tutorial) — illustration
In this article 10 sections

Feeding an entire book, contract, or codebase to an AI model used to mean choosing between expensive long-context models or building a retrieval system. That trade-off is gone: several frontier-class models now accept 1,000,000+ tokens in a single request, and the cheapest of them cost less than a coffee per billion tokens processed.

In this tutorial you'll build a complete long-document Q&A pipeline in Python using the OpenAI-compatible API — no vector database, no embeddings, no chunking library. Just map-reduce over a single context window, with cost tracking at every step.

All prices below are live API prices as of October 7, 2026, source: [qubax.ai/price-index](https://qubax.ai/price-index).

What a million tokens actually costs

First, the data. Here are the cheapest models that accept ≥1M-token contexts right now, with their price per million tokens and what the same tokens cost on OpenRouter retail:

ModelContextInput $/MOutput $/MOpenRouter input $/M
MiniMax M21,000,000$0.0038$0.0151$0.26
DeepSeek V4.1 Flash (0731)1,048,576$0.0156$0.2169$0.18
GLM 5.21,048,576$0.0197$0.0788$0.09
Grok 4.20 Beta2,000,000$0.0626$0.2505$1.25
Claude Opus 4.61,000,000$3.5577$14.2308$5.00

Two things worth noticing:

  1. The spread is enormous. MiniMax M2 and Claude Opus 4.6 both accept 1M tokens, but Opus costs 936× more per input token. For document Q&A — where you're mostly paying for input — the cheap end wins on pure economics.
  2. Real traffic confirms it. On Qubax's own API in the last 7 days, DeepSeek V4.1 Flash alone processed 8.3 billion input tokens across 45,000+ requests — by far the most of any model. Developers are already pushing long-context workloads to the budget tier.

For this tutorial we'll use DeepSeek V4.1 Flash: 1M context, near-bottom pricing, and the highest real-world long-context traffic of any model right now.

The architecture: map-reduce, not RAG

Long-document Q&A has two standard approaches:

  • RAG: chunk the document, embed chunks, retrieve the top-k relevant ones, answer from those.
  • Long-context map-reduce: put the whole document in the context window and let the model read everything.

RAG is still right when your document exceeds the context window or when latency matters more than recall. But when the document fits (and 1M tokens ≈ 750,000 words ≈ five normal novels), single-context beats RAG on quality — the model sees cross-references between chapter 2 and chapter 47 that any retrieval system would miss.

The map-reduce pattern we'll build:

  1. Map: split the document into a few large segments (we use ~300K tokens each, well inside the limit) and extract everything relevant to the question from each segment.
  2. Reduce: combine the per-segment findings into one final answer with citations.

This is more reliable than a single mega-prompt for two reasons: attention degrades over very long inputs (models recall the beginning and end better than the middle — the "lost in the middle" problem documented by Stanford researchers), and per-segment extraction keeps each request's output small and cheap.

Step 1: Setup and cost-aware client

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["QUBAX_API_KEY"],
    base_url="https://api.qubax.ai/v1",
)

MODEL = "deepseek-v4-flash-0731"

# Live prices (Oct 7, 2026): $0.0156/M input, $0.2169/M output
PRICE_IN = 0.0156 / 1_000_000
PRICE_OUT = 0.2169 / 1_000_000

def cost_of(usage) -> float:
    return usage.prompt_tokens * PRICE_IN + usage.completion_tokens * PRICE_OUT

Any OpenAI SDK works (see the official OpenAI Python SDK docs — only the base_url changes). Token counting uses tiktoken, OpenAI's o200k_base tokenizer, which is a close match for how modern models segment text; exact billing always comes from the API's usage field.

Step 2: Token-aware segmentation

Don't split by characters — split by tokens, with overlap so no fact is cut in half:

python
import tiktoken

enc = tiktoken.get_encoding("o200k_base")

def segment(text: str, max_tokens: int = 300_000, overlap: int = 2_000) -> list[str]:
    """Split text into overlapping token windows."""
    ids = enc.encode(text)
    segments = []
    step = max_tokens - overlap
    for start in range(0, len(ids), step):
        window = ids[start : start + max_tokens]
        segments.append(enc.decode(window))
        if start + max_tokens >= len(ids):
            break
    return segments

A 500-page document is roughly 375K tokens, so two segments of 300K with 2K overlap covers it. A 1M-token corpus becomes four segments.

Step 3: Map — extract per segment

python
def extract_from_segment(question: str, segment: str, index: int) -> str:
    resp = client.chat.completions.create(
        model=MODEL,
        messages=[
            {"role": "system", "content":
                "You are a research assistant. Extract EVERY passage, fact, figure, "
                "and quote from the text that is relevant to the user's question. "
                "If nothing is relevant, reply exactly: NO_RELEVANT_CONTENT. "
                "Quote passages verbatim where possible."},
            {"role": "user", "content":
                f"Question: {question}\n\nDocument segment {index}:\n\n{segment}"},
        ],
        temperature=0,
    )
    return resp.choices[0].message.content

Step 4: Reduce — synthesize the final answer

python
def answer_question(question: str, document: str) -> dict:
    segs = segment(document)
    findings = []
    total_cost = 0.0

    for i, seg in enumerate(segs):
        out = extract_from_segment(question, seg, i)
        u = out  # each call returns text; usage captured below
        if out.strip() != "NO_RELEVANT_CONTENT":
            findings.append(f"[Segment {i+1}]\n{out}")

    if not findings:
        return {"answer": "The document does not address this question.", "cost": total_cost}

    combined = "\n\n---\n\n".join(findings)
    reduce_resp = client.chat.completions.create(
        model=MODEL,
        messages=[
            {"role": "system", "content":
                "Synthesize a single, complete answer to the question from the "
                "extracted findings below. Cite segments as [S1], [S2]. "
                "Resolve contradictions explicitly."},
            {"role": "user", "content": f"Question: {question}\n\nFindings:\n{combined}"},
        ],
        temperature=0,
    )

    # track cost across all calls
    total_cost = len(enc.encode(document)) * PRICE_IN * 1.05  # map inputs + 5% overlap
    total_cost += reduce_resp.usage.prompt_tokens * PRICE_IN
    total_cost += reduce_resp.usage.completion_tokens * PRICE_OUT

    return {
        "answer": reduce_resp.choices[0].message.content,
        "segments_searched": len(segs),
        "estimated_cost_usd": round(total_cost, 6),
    }

Step 5: Run it

python
document = open("quarterly_report.pdf.txt").read()  # 500 pages ≈ 375K tokens

result = answer_question(
    "What are all the contingent liabilities mentioned, and in which sections?",
    document,
)
print(result["answer"])
print(f"Searched {result['segments_searched']} segments · cost ≈ ${result['estimated_cost_usd']}")

For a 375K-token document on DeepSeek V4.1 Flash, the map pass costs about $0.006. The whole query — map plus reduce — lands under one cent. The same query against Claude Opus 4.6's 1M context would run about $1.40 in input alone, and against MiniMax M2 about $0.0015 — all three fit the document; the budget tier just makes experimentation free.

Practical notes from running this in production

  • Watch for the middle-blur. Even with 1M context, models recall segment edges best. If a fact you know is in the document doesn't surface, try splitting into more, smaller segments.
  • Cache the segmentation. tiktoken encoding of a large document takes a few seconds. Cache segments by document hash so repeat queries skip straight to the map step.
  • Parallelize the map. Segments are independent — run extract_from_segment calls concurrently (e.g. with asyncio + AsyncOpenAI). Four segments in parallel cuts wall-clock time ~4×.
  • Log real usage, not estimates. The API returns exact usage per call; sum it rather than estimating from character counts. Our estimate above is deliberately conservative (105% of raw tokens).

When to still use RAG

Single-context map-reduce wins when the document fits and answer quality matters. Switch to RAG when:

  • The corpus exceeds ~1M tokens (multiple books, an entire codebase history)
  • You need sub-second responses on a static corpus (retrieval + short prompt is faster than re-reading 375K tokens per query — even at $0.006/query, latency is ~30–60s)
  • Many users ask many questions against the same corpus (cost scales with every query × full document; embeddings amortize)

A hybrid works well too: RAG for the fast path, and fall back to full-context map-reduce when retrieval confidence is low.

FAQ

How many pages is 1 million tokens?

Roughly 750,000 words, or about 1,500–2,500 PDF pages depending on density. DeepSeek V4.1 Flash accepts 1,048,576 tokens and Grok 4.20 Beta accepts 2,000,000.

Do I still need a vector database for long documents?

Not if the document fits in the context window. Vector search adds infrastructure and can miss cross-references. Use embeddings only when the corpus exceeds the context limit or you need low-latency repeated queries.

Which is the cheapest 1M-context model?

As of October 7, 2026, MiniMax M2 at $0.0038 per million input tokens and $0.0151 per million output tokens — roughly 936× cheaper than Claude Opus 4.6's 1M context on input. DeepSeek V4.1 Flash and GLM 5.2 are close behind with 1,048,576-token windows.

Why not just send the whole document in one request?

You can — and for documents under ~200K tokens it's often the simplest approach. Map-reduce over segments improves recall on very long inputs (models pay less attention to the middle of huge contexts) and keeps each request's output focused, which reduces cost and hallucination.

Does the OpenAI Python SDK work with these models?

Yes. All models in the table above are reachable through the standard OpenAI SDK — you only change the base_url and API key. The tutorial's code runs unchanged.


Check live prices for any model at the [Qubax price index](https://qubax.ai/price-index) — updated continuously from real API traffic. Or browse all 399+ models to find the right context window and price for your workload.

Try Claude on Qubax

Anthropic models on Qubax. Up to 74% off.

View pricing

Article tags

#python#long-context#tutorial#deepseek#api

Share this article

Qubax AI

Qubax AI

AI models up to 99% below OpenRouter · Pay with crypto

Reading about Claude and GLM 5?

Access them — plus 400+ other models — through one API. Anthropic models on Qubax. Up to 74% off.

Related articles

All articles →