LLM API

Claude API Pricing and Token Costs

Creating an Anthropic API key does not itself add a fee. Claude API usage is billed by model and can include input, cache writes, cache hits, output tokens, and execution-mode pricing.

Decision summary: Choose Anthropic when long-context quality and careful instruction following justify the cost.
Pricing checked 2026-08-26high confidence

Does a Claude API key cost money?

The API key itself does not have a separate purchase price. Charges come from the Claude API usage attached to the account, including applicable token, caching, and execution-mode rates.

Claude API pricing per token

Anthropic publishes text rates per 1 million tokens. The table below also converts the tracked default input, cache-hit, and output rates to 1,000 tokens for smaller workload estimates.

Cheapest Claude API model

Claude Haiku 4.5 has the lowest tracked default input and output rates in this table: $1.00 input and $5.00 output per 1M tokens. Test task quality before routing production traffic.

Pricing overview

Usage-based API pricing with separate input, output, cache write, cache hit, batch, and fast-mode rates where available.

Source-tracked model data

Anthropic API model prices in one table

Example cost uses 1,000,000 input tokens and 300,000 output tokens at each model's default tracked rate, with no cache discount.

ModelInput priceCached inputOutput priceCache writeExample cost
Claude Opus 5
Stable
$5.00 / 1M
$0.005 / 1K
Not tracked$25.00 / 1M
$0.025 / 1K
Not tracked$12.50
Claude Fable 5
Stable
$10.00 / 1M
$0.01 / 1K
$1.00 / 1M
$0.001 / 1K
$50.00 / 1M
$0.05 / 1K
5m: $12.50 / 1M
$0.0125 / 1K
1h: $20.00 / 1M
$0.02 / 1K
$25.00
Claude Mythos 5
Limited availability
$10.00 / 1M
$0.01 / 1K
$1.00 / 1M
$0.001 / 1K
$50.00 / 1M
$0.05 / 1K
5m: $12.50 / 1M
$0.0125 / 1K
1h: $20.00 / 1M
$0.02 / 1K
$25.00
Claude Sonnet 5
Stable
$2.00 / 1M
$0.002 / 1K
$0.20 / 1M
$0.0002 / 1K
$10.00 / 1M
$0.01 / 1K
5m: $2.50 / 1M
$0.0025 / 1K
1h: $4.00 / 1M
$0.004 / 1K
$5.00
Claude Opus 4.8
Stable
$5.00 / 1M
$0.005 / 1K
$0.50 / 1M
$0.0005 / 1K
$25.00 / 1M
$0.025 / 1K
5m: $6.25 / 1M
$0.00625 / 1K
1h: $10.00 / 1M
$0.01 / 1K
$12.50
Claude Sonnet 4.6
Stable
$3.00 / 1M
$0.003 / 1K
$0.30 / 1M
$0.0003 / 1K
$15.00 / 1M
$0.015 / 1K
5m: $3.75 / 1M
$0.00375 / 1K
1h: $6.00 / 1M
$0.006 / 1K
$7.50
Claude Haiku 4.5
Stable
$1.00 / 1M
$0.001 / 1K
$0.10 / 1M
$0.0001 / 1K
$5.00 / 1M
$0.005 / 1K
5m: $1.25 / 1M
$0.00125 / 1K
1h: $2.00 / 1M
$0.002 / 1K
$2.50

Cost boundary: this example compares token charges only. It does not assume equal output quality, latency, retry rates, or output length across models.

Monthly cost calculator

Estimate monthly Anthropic API cost

Choose a model and enter a typical request. The estimate uses the same source-tracked pricing records as the table above.

Estimated monthly token cost$0.00
Input
$0.00
Output
$0.00
Per 1,000 requests
$0.00

Estimate excludes untracked provider-specific charges.

Compare with other providers
StackLens assessment

How to shortlist a Claude API model

Compare the lowest-cost suitable Claude route first, then include cache behavior, tokenizer changes, temporary pricing, and access limits in the rollout decision.

Establish a lower-cost baseline

Claude Haiku 4.5

Test routine and high-volume cases first, then record which requests need escalation rather than routing every request to a higher-priced model.

Compare Sonnet routes

Claude Sonnet 5 and Claude Sonnet 4.6

Account for Sonnet 5 introductory pricing and the provider-documented tokenizer difference when comparing measured workload cost.

Review premium and limited routes

Claude Opus 4.8, Claude Fable 5, and Claude Mythos 5

Use representative difficult tasks and confirm access boundaries before assigning production traffic to these higher-priced routes.

What affects cost

  • Long context
  • Long generated answers
  • Cache write and cache hit behavior
  • Retries
  • Batch and fast-mode terms

Best for

  • coding-heavy teams
  • document analysis
  • long-context assistants

Not ideal for

  • very high output volume without caps
  • simple extraction tasks that smaller models can handle
  • teams tied to OpenAI-only features
StackLens assessment

How to control monthly Anthropic API cost

Claude budgets are especially sensitive to generated output, repeated long context, cache behavior, and temporary model pricing.

Route routine work to Haiku

Test Claude Haiku 4.5 on routine and high-volume requests before assigning every request to Sonnet or Opus.

Measure cache writes and hits separately

A cache write and a cache hit have different rates, so a useful forecast needs the expected reuse pattern rather than one blanket discount.

Cap generated output

Claude output rates are higher than input rates in the tracked standard profiles, making unnecessary answer length an important budget variable.

Date temporary prices

Use the documented post-introductory Sonnet 5 rates for workloads that continue after 2026-08-31.

Separate batch and fast-mode scenarios

Do not mix standard, batch, and fast-mode rates in one estimate; model each execution path separately.

Count retries and agent loops

Include repeated model calls and regenerated answers when forecasting coding or agent workflows.

Cost example

Claude output volume can outweigh input cost

Example using Claude Haiku 4.5 standard pricing with 1,000,000 input tokens. The only change is output volume.

300,000 output tokens: $1.00 input + $1.50 output = $2.50

2,000,000 output tokens: $1.00 input + $10.00 output = $11.00

Budget review

Anthropic budget mistakes to avoid

  • Applying the cache-hit rate to context that must first be written or is rarely reused.
  • Extending Claude Sonnet 5 introductory pricing beyond its documented end date.
  • Routing routine requests to Sonnet or Opus before testing Haiku on representative tasks.
  • Forecasting one model call when the workflow can retry, use tools, or run an agent loop.
FAQ

Anthropic API pricing questions

How does Claude Sonnet 5 introductory pricing affect a budget?

The tracked introductory rates end after 2026-08-31. Forecast workloads that continue into September with the documented rates effective 2026-09-01 instead of extending the introductory price indefinitely.

When can Claude prompt caching reduce input cost?

Caching can help when stable context is reused, but cache writes and cache hits have different rates. Measure how often the same context is written and reused instead of applying the cache-hit rate to every input token.

How do I convert Anthropic API pricing per million tokens to cost per token?

Divide the tracked price per 1,000,000 tokens by 1,000,000. Calculate input and output separately because their rates differ, then multiply each rate by the tokens used.

How should I compare Anthropic API pricing across models?

Use the same input volume, output volume, cache assumptions, and request count for each model. Cost alone does not establish equal quality, latency, or retry behavior.