LLM API

OpenAI API Pricing and Token Costs

Creating an OpenAI API key does not itself add a fee. API usage is billed separately by model and can include input, cached input, cache writes, and output tokens.

Decision summary: Choose OpenAI when ecosystem maturity and multimodal coverage matter; control cost with routing and output limits.
Pricing checked 2026-08-21high confidence

Does an OpenAI API key cost money?

The API key itself does not have a separate purchase price. Charges come from the API usage attached to the account, including applicable model token rates and any separately priced provider features.

OpenAI API pricing per token

OpenAI publishes text rates per 1 million tokens. The table below also converts each tracked rate to 1,000 tokens so small requests are easier to estimate.

Cheapest OpenAI API model

GPT-4o mini has the lowest tracked default input and output rates in this table: $0.15 input and $0.60 output per 1M tokens. Test task quality before routing production traffic.

Pricing overview

Usage-based API pricing by model, input tokens, output tokens, context tier, caching, and modality.

Source-tracked model data

OpenAI API model prices in one table

Example cost uses 1,000,000 input tokens and 300,000 output tokens at each model's default tracked rate, with no cache discount.

ModelInput priceCached inputOutput priceContextExample cost
GPT-5.6 Luna
Stable
$0.20 / 1M
$0.0002 / 1K
$0.02 / 1M
$0.00002 / 1K
$1.20 / 1M
$0.0012 / 1K
1.1M tokens$0.56
GPT-5.6 Sol
Stable
$5.00 / 1M
$0.005 / 1K
$0.50 / 1M
$0.0005 / 1K
$30.00 / 1M
$0.03 / 1K
1.1M tokens$14.00
GPT-5.6 Terra
Stable
$2.00 / 1M
$0.002 / 1K
$0.20 / 1M
$0.0002 / 1K
$12.00 / 1M
$0.012 / 1K
1.1M tokens$5.60
gpt-5.4-mini
Stable
$0.75 / 1M
$0.00075 / 1K
$0.075 / 1M
$0.000075 / 1K
$4.50 / 1M
$0.0045 / 1K
400K tokens$2.10
gpt-5.4-nano
Stable
$0.20 / 1M
$0.0002 / 1K
$0.02 / 1M
$0.00002 / 1K
$1.25 / 1M
$0.00125 / 1K
400K tokens$0.575
GPT-4o
Stable
$2.50 / 1M
$0.0025 / 1K
$1.25 / 1M
$0.00125 / 1K
$10.00 / 1M
$0.01 / 1K
128K tokens$5.50
GPT-4o mini
Stable
$0.15 / 1M
$0.00015 / 1K
$0.075 / 1M
$0.000075 / 1K
$0.60 / 1M
$0.0006 / 1K
128K tokens$0.33

Cost boundary: this example compares token charges only. It does not assume equal output quality, latency, retry rates, or output length across models.

Monthly cost calculator

Estimate monthly OpenAI API cost

Choose a model and enter a typical request. The estimate uses the same source-tracked pricing records as the table above.

Estimated monthly token cost$0.00
Input
$0.00
Output
$0.00
Per 1,000 requests
$0.00

Estimate excludes untracked provider-specific charges.

Compare with other providers
Official price history

OpenAI API recorded model price changes

Verified changes for exact models and pricing profiles available through this provider. New model generations are not treated as price changes.

OpenAI · Standard · Short Context

GPT-5.6 Luna

First recorded 2026-07-31
RateEarlier priceLater recorded priceChangePercentage
Input / 1M$1.00$0.20-$0.80-80.0%
Cached input / 1M$0.10$0.02-$0.08-80.0%
Output / 1M$6.00$1.20-$4.80-80.0%
Cache write / 1M$1.25$0.25-$1.00-80.0%
Checked 2026-07-31Official source ↗

StackLens first observed the lower price on the current official OpenAI pricing source; the provider's original effective date was not independently established.

OpenAI · Standard · Long Context

GPT-5.6 Luna

First recorded 2026-07-31
RateEarlier priceLater recorded priceChangePercentage
Input / 1M$2.00$0.40-$1.60-80.0%
Cached input / 1M$0.20$0.04-$0.16-80.0%
Output / 1M$9.00$1.80-$7.20-80.0%
Cache write / 1M$2.50$0.50-$2.00-80.0%
Checked 2026-07-31Official source ↗

Long-context rates are derived from OpenAI's documented 2x input and 1.5x output multipliers applied to the current base rates.

OpenAI · Standard · Short Context

GPT-5.6 Terra

First recorded 2026-07-31
RateEarlier priceLater recorded priceChangePercentage
Input / 1M$2.50$2.00-$0.50-20.0%
Cached input / 1M$0.25$0.20-$0.05-20.0%
Output / 1M$15.00$12.00-$3.00-20.0%
Cache write / 1M$3.13$2.50-$0.625-20.0%
Checked 2026-07-31Official source ↗

StackLens first observed the lower price on the current official OpenAI pricing source; the provider's original effective date was not independently established.

OpenAI · Standard · Long Context

GPT-5.6 Terra

First recorded 2026-07-31
RateEarlier priceLater recorded priceChangePercentage
Input / 1M$5.00$4.00-$1.00-20.0%
Cached input / 1M$0.50$0.40-$0.10-20.0%
Output / 1M$22.50$18.00-$4.50-20.0%
Cache write / 1M$6.25$5.00-$1.25-20.0%
Checked 2026-07-31Official source ↗

StackLens first observed the lower price on the current official OpenAI pricing source; the provider's original effective date was not independently established.

StackLens assessment

How to shortlist an OpenAI API model

Begin with the lowest-cost tracked route that can pass the workload's acceptance test, then escalate only the cases that need another model.

Start with lower-cost routes

GPT-4o mini and gpt-5.4-nano

Use a representative test set to determine whether the lower tracked token rates are sufficient for routine requests.

Test middle price tiers

gpt-5.4-mini and gpt-5.6-luna

Compare these routes on the lower-cost model's failure cases, including retries and generated output length.

Escalate selected workloads

GPT-4o, gpt-5.6-terra, and gpt-5.6-sol

Reserve higher tracked rates for workloads where measured task outcomes justify the additional cost.

What affects cost

  • Output length
  • Repeated context
  • Long-context tier selection
  • Cache write behavior
  • Retries and agent loops
  • Regional processing premiums

Best for

  • production AI features
  • multimodal apps
  • teams that value ecosystem depth

Not ideal for

  • teams optimizing only for lowest cost
  • workloads without token monitoring
  • buyers that need fixed monthly spend
StackLens assessment

How to reduce monthly OpenAI API cost

A useful OpenAI budget starts with request routing, output limits, repeated context, and retry behavior rather than one blended token estimate.

Route simple tasks to lower-cost models

Route simple tasks to lower-cost tracked models such as gpt-5.4-nano, gpt-5.4-mini, gpt-5.6-luna, or GPT-4o mini where quality is good enough.

Cap output tokens

Output tokens can dominate spend because generated answers are billed separately from input.

Cache repeated context

Repeated system prompts, policy text, and document context can inflate input spend when sent on every request.

Monitor retries and agent loops

Retries, tool calls, and agent loops can multiply cost even when each individual request looks small.

Measure cost per feature

Track cost by product feature or workflow so expensive paths are visible before the monthly bill arrives.

Compare before and after routing

Run the same workload through the original route and a lower-cost route, then compare quality and spend together.

Cost example

Output length can dominate total cost

Example using GPT-4o mini tracked pricing with 1,000,000 input tokens. The only change is output volume.

300,000 output tokens: $0.15 input + $0.18 output = $0.33

2,000,000 output tokens: $0.15 input + $1.20 output = $1.35

Budget review

Common budget mistakes to avoid

  • Choosing a flagship model for every request before measuring quality needs.
  • Letting output length grow without product-level caps.
  • Sending the same long context repeatedly instead of caching or trimming it.
  • Ignoring retries and agent loops in monthly forecasts.
FAQ

OpenAI API pricing questions

Is OpenAI the cheapest LLM API?

Not for every workload. GPT-4o mini and gpt-5.4-nano have low tracked token rates, while larger OpenAI models cost more. Compare the models that pass your task test using the same input volume, output volume, and retry assumptions.

Which OpenAI cost drivers should I measure first?

Measure generated output, repeated context, retries, and agent loops. Then separate routine requests that can use a lower-cost route from the smaller set that needs a higher-priced model.

Does an OpenAI API key have a separate cost?

StackLens does not add a separate API-key fee to its estimates. The calculator models tracked token usage; account, billing, and any non-token charges should be checked against current official OpenAI terms.

How do I estimate OpenAI API cost per request?

Multiply input tokens by the model input rate, multiply output tokens by the output rate, then divide each per-million rate by 1,000,000. Add the two results and account for caching, retries, and other tracked charges.

Is ChatGPT subscription pricing the same as OpenAI API pricing?

No. This page covers API token usage, not ChatGPT subscription pricing. Treat ChatGPT plans and API billing as separate purchasing decisions and verify current account terms before purchase.