What affects cost
- Selected model endpoint
- Input and output token volume
- Text, speech, transcription, or other service unit
- Retries and application-level agent loops
- Production rate limits and enterprise-only model requirements
GroqCloud API pricing lists per-model input and output token rates for hosted open models, with spend controls and separate rates for audio services.
Usage-based model inference priced per million tokens, plus separate per-character or per-audio-hour units for non-text services.
Example cost uses 1,000,000 input tokens and 300,000 output tokens at each model's default tracked rate, with no cache discount.
| Model | Input price | Cached input | Output price | Context | Example cost |
|---|---|---|---|---|---|
| GPT-OSS 120B on Groq Stable | $0.15 / 1M | $0.075 / 1M | $0.60 / 1M | 131.1K tokens | $0.33 |
| GPT-OSS 20B on Groq Stable | $0.075 / 1M | $0.037 / 1M | $0.30 / 1M | 131.1K tokens | $0.165 |
| Llama 4 Scout on Groq Preview | $0.11 / 1M | Not tracked | $0.34 / 1M | 131.1K tokens | $0.212 |
| Qwen3-32B on Groq Preview | $0.29 / 1M | Not tracked | $0.59 / 1M | 131.1K tokens | $0.467 |
Cost boundary: this example compares token charges only. It does not assume equal output quality, latency, retry rates, or output length across models.
No. Each listed model has its own input and output rate, while speech and transcription services use different billing units.
No. Token cost should be evaluated with task quality, output length, retries, context limits, endpoint availability, and rate limits.