One dimension is not a complete decision
Cost, token volume, capacity, and evidence coverage answer different questions. Compare them with a representative prompt, actual output length, retry behavior, and your quality bar.
Find a model for a taskSee what kind of evidence is documented for tracked models before relying on a comparison. Official facts, user reports, and controlled tests remain separate.
More sources do not mean a model is better. Community reports are not controlled benchmarks, and provider documentation is not independent performance evidence.
| Model | Official facts | Community reports | Controlled StackLens test | Last reviewed |
|---|---|---|---|---|
| Claude Fable 5 | 2 | 5 | Not run | 2026-07-13 |
| Claude Sonnet 5 | 2 | 5 | Not run | 2026-07-13 |
| Gemini 3.5 Flash | 2 | 5 | Not run | 2026-07-13 |
| GPT-5.6 Sol | 1 | 5 | Not run | 2026-07-13 |
| GPT-5.6 Terra | 1 | 3 | Not run | 2026-07-13 |
| GPT-5.6 Luna | 1 | 3 | Not run | 2026-07-13 |
Read the methodology for source boundaries. A community report is a lead for testing, not a controlled performance result.
Cost, token volume, capacity, and evidence coverage answer different questions. Compare them with a representative prompt, actual output length, retry behavior, and your quality bar.
Find a model for a taskNo. Each board measures one defined dimension. Token count, price, context capacity, and evidence coverage do not substitute for testing your own completed tasks.