Measure a real prompt
Use representative input rather than a word-count shortcut. Output length should reflect the answer your application actually returns.
Count prompt tokensChoose a model, enter the monthly budget and token shape of one request, and see the request volume that budget can support.
requests within the reserved budget
Calculating...
Use representative input rather than a word-count shortcut. Output length should reflect the answer your application actually returns.
Count prompt tokensUse request volume, cache behavior, and workload assumptions when you need a broader estimate than one request.
Open monthly cost calculatorRouting routine traffic can change capacity, but verify quality, latency, retries, and tool-call behavior separately.
Model routing costIt divides a monthly budget by the estimated input and output token cost of one request, then shows how many requests the budget can support under those assumptions.
No. The result covers the input and output token charges entered in the form. Retries, tool calls, caching, storage, batch terms, and other separately billed features need to be added to a fuller workload estimate.
Output rates are often higher than input rates, so long generated answers can use a large share of the budget even when the prompt is short. Test representative output lengths before setting a production limit.