One dimension is not a complete decision
Cost, token volume, capacity, and evidence coverage answer different questions. Compare them with a representative prompt, actual output length, retry behavior, and your quality bar.
Find a model for a taskCompare documented context capacity and maximum output limits across tracked models. A larger window may help a workload, but it does not prove better answers or lower cost.
Context and output limits come from the model records and their specification sources. Verify endpoint, tier, and effective limits before production use.
| Rank | Model | Provider | Context window | Max output | Specs checked |
|---|---|---|---|---|---|
| 1 | Gemini 1.5 Pro | 2,097,152 | 8,192 | 2026-07-17 | |
| 2 | GPT-5.6 Luna | OpenAI | 1,050,000 | 128,000 | 2026-07-13 |
| 3 | GPT-5.6 Sol | OpenAI | 1,050,000 | 128,000 | 2026-07-13 |
| 4 | GPT-5.6 Terra | OpenAI | 1,050,000 | 128,000 | 2026-07-13 |
| 5 | Gemini 1.5 Flash | 1,048,576 | 8,192 | 2026-07-17 | |
| 6 | Gemini 2.5 Flash | 1,048,576 | 65,536 | 2026-07-13 | |
| 7 | Gemini 2.5 Flash-Lite | 1,048,576 | 65,536 | 2026-07-13 | |
| 8 | Gemini 2.5 Pro | 1,048,576 | 65,536 | 2026-07-13 | |
| 9 | Gemini 3 Flash Preview | 1,048,576 | 65,536 | 2026-07-14 | |
| 10 | Gemini 3.1 Flash-Lite | 1,048,576 | 65,536 | 2026-07-14 | |
| 11 | Gemini 3.1 Pro Preview | 1,048,576 | 65,536 | 2026-07-14 | |
| 12 | Gemini 3.5 Flash | 1,048,576 | 65,536 | 2026-07-14 | |
| 13 | Gemini 3.5 Flash-Lite | 1,048,576 | 65,536 | 2026-07-21 | |
| 14 | Gemini 3.6 Flash | 1,048,576 | 65,536 | 2026-07-21 | |
| 15 | Kimi K3 | Moonshot AI | 1,048,576 | Requires verification | 2026-07-17 |
| 16 | Claude Fable 5 | Anthropic | 1,000,000 | 128,000 | 2026-07-14 |
| 17 | Claude Mythos 5 | Anthropic | 1,000,000 | 128,000 | 2026-07-14 |
| 18 | Claude Opus 4.8 | Anthropic | 1,000,000 | 128,000 | 2026-07-13 |
| 19 | Claude Opus 5 | Anthropic | 1,000,000 | 128,000 | 2026-07-28 |
| 20 | Claude Sonnet 4.6 | Anthropic | 1,000,000 | 128,000 | 2026-07-13 |
| 21 | Claude Sonnet 5 | Anthropic | 1,000,000 | 128,000 | 2026-07-13 |
| 22 | DeepSeek V4 Flash non-thinking | DeepSeek | 1,000,000 | 384,000 | 2026-07-13 |
| 23 | DeepSeek V4 Pro | DeepSeek | 1,000,000 | 384,000 | 2026-07-13 |
| 24 | Grok 4.5 | xAI | 500,000 | Requires verification | 2026-07-17 |
| 25 | gpt-5.4-mini | OpenAI | 400,000 | 128,000 | 2026-07-13 |
| 26 | gpt-5.4-nano | OpenAI | 400,000 | 128,000 | 2026-07-13 |
| 27 | Kimi K2.6 | Moonshot AI | 262,144 | Requires verification | 2026-07-17 |
| 28 | Kimi K2.7 Code | Moonshot AI | 262,144 | Requires verification | 2026-07-17 |
| 29 | Kimi K2.7 Code High-Speed | Moonshot AI | 262,144 | Requires verification | 2026-07-17 |
| 30 | Command A | Cohere | 256,000 | 8,000 | 2026-07-17 |
| 31 | Mistral Medium 3.5 | Mistral | 256,000 | Requires verification | 2026-07-17 |
| 32 | Claude Haiku 4.5 | Anthropic | 200,000 | 64,000 | 2026-07-13 |
| 33 | Gemini 3.5 Live Translate Preview | 131,072 | 65,536 | 2026-07-14 | |
| 34 | GPT-OSS 120B on Groq | Groq | 131,072 | 65,536 | 2026-07-15 |
| 35 | GPT-OSS 20B on Groq | Groq | 131,072 | 65,536 | 2026-07-15 |
| 36 | Llama 4 Scout on Groq | Groq | 131,072 | 8,192 | 2026-07-15 |
| 37 | Qwen3-32B on Groq | Groq | 131,072 | 40,960 | 2026-07-15 |
| 38 | Command R 08-2024 | Cohere | 128,000 | 4,000 | 2026-07-17 |
| 39 | Command R+ 08-2024 | Cohere | 128,000 | 4,000 | 2026-07-17 |
| 40 | GPT-4o | OpenAI | 128,000 | 16,384 | 2026-07-13 |
| 41 | GPT-4o mini | OpenAI | 128,000 | 16,384 | 2026-07-13 |
A context window is a documented capacity limit, not a guarantee that a task will use the full window well. Check the context window calculator before choosing a model.
Cost, token volume, capacity, and evidence coverage answer different questions. Compare them with a representative prompt, actual output length, retry behavior, and your quality bar.
Find a model for a taskNo. Each board measures one defined dimension. Token count, price, context capacity, and evidence coverage do not substitute for testing your own completed tasks.