Transparent decision data

AI model rankings you can inspect

Compare measured token efficiency, a fixed API cost scenario, and the evidence behind model decisions. Every board shows its inputs, sources, and limits.

Measured tokenizer data

Token efficiency board

Lower measured token count for the same text means lower token volume in this sample set. It does not predict model quality.

RankTokenizerSamplesTotal tokensTokens / 100 charsBasis
1OpenAI Current311427.2o200k_base
2DeepSeek V4 Pro311727.9Pinned tokenizer JSON
3Llama 3311828.2Pinned tokenizer JSON
4OpenAI Legacy312128.9cl100k_base
5Qwen 3.5312930.8Pinned tokenizer JSON
6Mistral Medium 3.5313231.5Pinned tokenizer JSON
7Gemma 3313532.2Pinned SentencePiece model

Checked 2026-07-19. Plain text only; no chat templates, roles, tools, images, or API wrappers. The exact source samples and per-sample results are available in the tokenizer comparison.

Source-tracked cost scenario

API cost board

1M input + 300K output tokens per month. Cost only; no equal-quality assumption.

RankModelInput / 1MOutput / 1MWorkload costChecked
1GPT-OSS 20B on GroqGroq$0.07$0.30$0.162026-07-15
2Llama 4 Scout on GroqGroq$0.11$0.34$0.212026-07-15
3Gemini 2.5 Flash-LiteGoogle$0.10$0.40$0.222026-08-05
4DeepSeek V4 Flash non-thinkingDeepSeek$0.14$0.28$0.222026-08-05
5Command R 08-2024Cohere$0.15$0.60$0.332026-07-17
6GPT-4o miniOpenAI$0.15$0.60$0.332026-08-05
7GPT-OSS 120B on GroqGroq$0.15$0.60$0.332026-07-15
8Qwen3-32B on GroqGroq$0.29$0.59$0.472026-07-15
9GPT-5.6 LunaOpenAI$0.20$1.20$0.562026-08-05
10gpt-5.4-nanoOpenAI$0.20$1.25$0.582026-08-05
11DeepSeek V4 ProDeepSeek$0.43$0.87$0.702026-08-05
12Gemini 3.1 Flash-LiteGoogle$0.25$1.50$0.702026-08-05
13Gemini 3.5 Live Translate PreviewGoogle$0.25$1.50$0.702026-08-05
14Gemini 2.5 FlashGoogle$0.30$2.50$1.052026-08-05
15Gemini 3.5 Flash-LiteGoogle$0.30$2.50$1.052026-08-05
16Gemini 3 Flash PreviewGoogle$0.50$3.00$1.402026-08-05
17gpt-5.4-miniOpenAI$0.75$4.50$2.102026-08-05
18Kimi K2.6Moonshot AI$0.95$4.00$2.152026-07-17
19Kimi K2.7 CodeMoonshot AI$0.95$4.00$2.152026-07-17
20Claude Haiku 4.5Anthropic$1.00$5.00$2.502026-08-05
21Gemini 3.6 FlashGoogle$1.50$7.50$3.752026-08-05
22Mistral Medium 3.5Mistral$1.50$7.50$3.752026-07-17
23Grok 4.5xAI$2.00$6.00$3.802026-07-17
24Gemini 3.5 FlashGoogle$1.50$9.00$4.202026-08-05
25Gemini 2.5 ProGoogle$1.25$10.00$4.252026-08-05
26Kimi K2.7 Code High-SpeedMoonshot AI$1.90$8.00$4.302026-07-17
27Claude Sonnet 5Anthropic$2.00$10.00$5.002026-08-05
28Command ACohere$2.50$10.00$5.502026-07-17
29Command R+ 08-2024Cohere$2.50$10.00$5.502026-07-17
30GPT-4oOpenAI$2.50$10.00$5.502026-08-05

This is a billing comparison for one reference workload, not a quality or “best model” ranking. Change the assumptions in the API cost calculator before making a routing decision.

Separate evidence types

Quality evidence coverage

These rows show how much evidence is documented for each model. More sources do not prove better quality.

ModelOfficial factsCommunity reportsControlled StackLens testLast reviewed
Claude Fable 52 sources5 sourcesNot run2026-07-13
Claude Sonnet 52 sources5 sourcesNot run2026-07-13
Gemini 3.5 Flash2 sources5 sourcesNot run2026-07-13
GPT-5.6 Sol1 source5 sourcesNot run2026-07-13
GPT-5.6 Terra1 source3 sourcesNot run2026-07-13
GPT-5.6 Luna1 source3 sourcesNot run2026-07-13

Official documentation, user reports, and controlled tests are not interchangeable. Read the methodology before using any signal in a production decision.

StackLens index

No fabricated composite score

StackLens does not currently publish a weighted quality score. Public leaderboards use different prompts, dates, models, and scoring rules, so combining them into one number would create false precision.

Find a model for a task
FAQ

Ranking questions

Is the cheapest model the best model?

No. The cost board compares token charges for one fixed workload and does not measure quality, latency, retries, or completed-task success.

How are token rankings calculated?

The tokenizer board sums measured token counts across the published English, Chinese, JSON, and code samples. The exact text, tokenizer basis, source, and check date remain visible.

Why is there no overall model quality ranking?

The available public sources do not use one shared controlled task set. StackLens keeps official facts, community reports, and independent tests separate rather than inventing a combined quality score.