Repeated context is the lever
System instructions, tool definitions, and stable document context are common candidates. New user input remains billed at the standard input rate in this scenario.
Estimate how much repeated prompt context could cost when a provider charges a lower cached-input rate. Compare standard and cached scenarios using tracked model prices.
Calculating...
System instructions, tool definitions, and stable document context are common candidates. New user input remains billed at the standard input rate in this scenario.
The calculator applies your cache-hit percentage only to repeated context. Measure real hit rate and expiration behavior after launch.
Minimum tokens, cache-write charges, storage duration, and API formatting can change the real result. Review the selected model's official source.
Read the methodologyIt compares standard input pricing with cached-input pricing for repeated context, plus output cost, using the selected model's tracked default profile.
No. Cache-write, cache-storage, minimum token, expiration, and provider-specific cache rules are not added to this simple scenario unless you model them separately from the official provider documentation.
Caching can help when the same system instructions, tool definitions, or document context recur often enough and the provider's cached-input rate is lower than its standard input rate. Validate cache eligibility and hit rate with your own traffic.