Compare Models

Side-by-side price, context, modalities, status, and measured latency — then copy the ID and build.

Loading live data…

How to compare AI models

Pick up to four models to see them side by side: price per 1 million input and output tokens on Qubax and on OpenRouter, context window, supported inputs (text, images, audio), and measured latency. Latency is time to first token (TTFT) at the median (p50) and the slow tail (p95), measured on real Qubax traffic.

Price alone is rarely the right way to choose. For chat and agents, output price matters more than input price because responses are long. For retrieval and long documents, input price and context window dominate. For interactive apps, p95 latency is what users feel. Benchmarks from LMArena, Epoch AI and LiveBench are linked from each model page.

Every model in the table is available through the same OpenAI-compatible endpoint, https://api.qubax.ai/v1. Switching models is a one-line change to the model ID, so you can A/B test two models on your own prompts before committing.

Frequently asked questions

Which AI model is the cheapest?
It depends on the job. Small open-weight models (Llama, Qwen, DeepSeek Flash) cost a few cents per million tokens; frontier models (GPT-5.x, Claude Opus, Gemini Pro) cost dollars. Use this table to compare the models that meet your quality bar, then pick the cheapest of those.
What is a context window?
The maximum number of tokens (roughly 0.75 words each) a model can read in one request, including your prompt, documents and chat history. Larger windows let you send whole files or long conversations without trimming.
What does TTFT p50 / p95 mean?
Time to first token: how long until the model starts streaming its answer. p50 is the typical request; p95 means 95% of requests were faster than this. Lower is better for chat and voice apps.
Are these the same models as on OpenAI, Anthropic or OpenRouter?
Yes. Qubax routes to the same official models through an open market of compute providers, which is why prices are usually lower. The API format is identical to OpenAI's, so existing SDKs and tools work unchanged.