Measured tokenizer data

Token efficiency rankings by tokenizer

See which tokenizer produces fewer tokens for the same published English, Chinese, JSON, and code samples. This is a token-volume comparison, not a model quality ranking.

Measured tokenizer data

Measured token counts for the same text

The board sums the same source samples across each tokenizer. Different chat wrappers, system prompts, tools, and model-specific processing can change production counts.

RankTokenizerSamplesTotal tokensTokens / 100 charsBasis
1OpenAI Current3 / 311427.2o200k_base
2DeepSeek V4 Pro3 / 311727.9Pinned tokenizer JSON
3Llama 33 / 311828.2Pinned tokenizer JSON
4OpenAI Legacy3 / 312128.9cl100k_base
5Qwen 3.53 / 312930.8Pinned tokenizer JSON
6Mistral Medium 3.53 / 313231.5Pinned tokenizer JSON
7Gemma 33 / 313532.2Pinned SentencePiece model

Checked 2026-07-19. Plain text only; no chat templates, roles, tools, images, or API wrappers. Open the reproducible tokenizer comparison to inspect the source text and each result.

Use the result carefully

One dimension is not a complete decision

Cost, token volume, capacity, and evidence coverage answer different questions. Compare them with a representative prompt, actual output length, retry behavior, and your quality bar.

Find a model for a task
FAQ

Ranking methodology questions

Does a higher ranking mean a better model?

No. Each board measures one defined dimension. Token count, price, context capacity, and evidence coverage do not substitute for testing your own completed tasks.