LLM API

DeepSeek API pricing

Useful for cost-sensitive teams evaluating DeepSeek with explicit cache-miss defaults and separately tracked cache-hit pricing.

Decision summary: Use DeepSeek in a shortlist when cache-miss pricing, context length, and output limits fit your workload; only model cache-hit savings with explicit assumptions.
Pricing checked 2026-08-16high confidence

Pricing overview

DeepSeek pricing separates cache-hit input, cache-miss input, and output. StackLens uses cache-miss input by default.

Source-tracked model data

DeepSeek API model prices in one table

Example cost uses 1,000,000 input tokens and 300,000 output tokens at each model's default tracked rate, with no cache discount.

ModelInput priceCached inputOutput priceContextExample cost
DeepSeek V4 Flash non-thinking
Stable
$0.14 / 1M$0.0028 / 1M$0.28 / 1M1M tokens$0.224
DeepSeek V4 Pro
Stable
$0.435 / 1M$0.0036 / 1M$0.87 / 1M1M tokens$0.696

Cost boundary: this example compares token charges only. It does not assume equal output quality, latency, retry rates, or output length across models.

What affects cost

  • Cache-miss input volume
  • Explicit cache-hit assumptions
  • Output length
  • Thinking or non-thinking mode
  • Context length and maximum output

Best for

  • cost-sensitive experiments
  • teams testing provider routing
  • large-context workflows after quality testing

Not ideal for

  • buyers assuming cache-hit pricing without evidence
  • teams that cannot monitor output length
  • production routing without quality and latency checks
FAQ

DeepSeek API pricing questions

Does StackLens use DeepSeek cache-hit pricing by default?

No. StackLens uses cache-miss input pricing by default and only applies cache-hit pricing when the user explicitly supplies a cache-hit percentage.

What are the tracked DeepSeek V4 Pro limits?

StackLens stores a 1,000,000 token context length and 384,000 maximum output tokens from the supplied official values.