G
What is Groq?
Groq runs open-weight models on its LPU (and, in 2026 messaging, LPX alongside Nvidia) for very high tokens/second. The API is OpenAI-compatible.
Important September 2026 catalog fact: llama-3.1-8b-instant and llama-3.3-70b-versatile shut down for free and Developer tiers on 16 August 2026; they remain Enterprise/contact-sales only. The public production prices on console.groq.com/docs/models are openai/gpt-oss-20b at $0.075 input / $0.30 output per 1M tokens (~1000 t/s) and openai/gpt-oss-120b at $0.15 / $0.60 (~500 t/s).
Whisper large-v3 is $0.111 per audio hour; turbo is $0.04. A free tier still exists (no card); limits are per organization.
Developer (card on file) raises RPM/TPM and unlocks Batch and Flex. Cached input is 50% off; Batch is 50% off and does not stack with the cache discount.
There is no ChatGPT-style $20 seat.
A concrete walkthrough: replace a dead Llama 8B Instant route without blowing the bill
If production still calls llama-3.1-8b-instant on a Free or Developer key, that route has been shut since 16 August 2026. Point it at openai/gpt-oss-20b, rerun your eval set, and only then consider gpt-oss-120b.
At 10M input + 2M output tokens/month, 20B is about $0.75 + $0.60 = $1.35; 120B is $1.50 + $1.20 = $2.70. Latency is why you are here — measure p95 before you 'optimize' to a slower host.
If the same system prompt is huge, enable caching (50% input) rather than jumping models. Batch the overnight backfill; it is another 50%, and it does not stack with cache.
Key features
- OpenAI-compatible API; swap the base URL if you already speak that SDK
- Production: GPT-OSS 20B $0.075/$0.30, GPT-OSS 120B $0.15/$0.60 per 1M, plus Whisper
- Speed on the order of 500–1000 tokens/sec on those two production text models
- Free org-level rate limits; Developer plan for higher caps, Batch, Flex
- Prompt cache at 50% of input; Batch at 50% of on-demand (no double dip)
- Llama 3.1 8B / 3.3 70B are Enterprise-only after the August 2026 self-serve shutdown
How to get started
- Create a key at console.groq.com — Free is enough to see whether latency is the win
- Call openai/gpt-oss-20b first; move to 120B only if quality fails on your eval set
- Put a card on the account (Developer) before you need Batch or headroom past Free RPM
- If your code still hard-codes llama-3.1-8b-instant, it will 404 for self-serve — migrate
- Do not pick Groq for GPT-4/Claude/Gemini; the catalog is open weights
Groq pricing
No seat fee. Official production list, September 2026: GPT-OSS 20B $0.075 / $0.30 per 1M; GPT-OSS 120B $0.15 / $0.60; Whisper v3 $0.111/hour, turbo $0.04/hour.
Free tier has org rate limits. Developer unlocks Batch/Flex and higher RPM.
Cache or Batch each cut 50%, not both. Llama 8B/70B: sales only.
| Plan | Price | Best for |
|---|---|---|
| Free | $0 + org rate limits | Latency tests and light prototypes |
| GPT-OSS 20B | $0.075 / $0.30 per 1M | Fast default text (~1000 t/s) |
| GPT-OSS 120B | $0.15 / $0.60 per 1M | Heavier open-weight quality (~500 t/s) |
| Whisper turbo | $0.04 / audio hour | Cheap transcription |
| Developer / Enterprise | Card or commit | Batch, Flex, leftover Llama SKUs |
Our take: Groq is a speed and price play on open models, not a frontier-lab catalog. Start on GPT-OSS 20B. If your repo still says Llama 8B Instant, fix the model id before you debug 'the API is down'.
Prices verified 2026-09. Plans change often — confirm on the official site above before you buy.
Use cases
- Voice or chat UIs where 200 ms of model time is the budget
- High-QPS classification on GPT-OSS 20B at $0.075/$0.30
- Whisper transcription when you want Groq's audio SKUs instead of OpenAI's
- Batch overnight jobs at half the token rate on Developer+
Strengths and tradeoffs
Where it stands out
- Real published token prices and 500–1000 t/s on the two production GPT-OSS models
- Free key with no card to prove the latency claim on your workload
- OpenAI-compatible enough that a base-URL swap is a day's work
Tradeoffs
- Self-serve Llama 8B/70B is gone — tutorials from 2025 will break
- No Claude/GPT-4o; if you need those names you want OpenAI/Anthropic or OpenRouter
- Org-level rate limits mean a second API key does not add quota
Groq alternatives
- Together AI: Wider open-model menu and fine-tunes; usually slower than Groq's LPU path.
- OpenRouter: One key that can route to Groq plus closed labs if you do not want to lock in.
- OpenAI API: GPT-family quality and tools; not 1000 tokens/sec on a $0.075 model.
- Replicate: Per-second GPUs for image/video models Groq does not host.
Pick Groq when tokens/sec and cheap open weights are the requirement. Pick Together for a broader catalog or fine-tunes.
Pick OpenAI/Anthropic for closed frontier models. Pick OpenRouter to keep a spare on-ramp.
FAQ
Is Groq free to use?
Groq's plans: Free tier then pay-per-token; GPT-OSS is the public catalog now. Pricing and free-tier limits change over time, so check the official site above for the latest details.
What is Groq used for?
LPU inference for open models — Llama 8B/70B left the self-serve catalog; GPT-OSS 20B/120B are the listed production prices.
Who is Groq best for?
Groq is a good fit for voice or chat UIs where 200 ms of model time is the budget, or high-QPS classification on GPT-OSS 20B at $0.075/$0.30.
What are the alternatives to Groq?
Commonly compared alternatives include Together AI, OpenRouter, and OpenAI API, see the comparison above for how they differ.
How much does Groq cost?
There is no monthly Groq subscription. Self-serve production text is GPT-OSS 20B at $0.075/$0.30 per million tokens or 120B at $0.15/$0.60. A free key exists with rate limits. Llama 3.1 8B and 3.3 70B are no longer on Free/Developer as of 16 August 2026.
What's the best Groq alternative?
Together AI is the most commonly recommended swap: wider open-model menu and fine-tunes; usually slower than Groq's LPU path. Which one is actually best depends on which of those tradeoffs matters more for what you're doing.
Can I still call Llama 3.1 8B Instant on Groq?
Not on Free or Developer. Groq shut llama-3.1-8b-instant and llama-3.3-70b-versatile for those tiers on 16 August 2026. The docs point you at openai/gpt-oss-20b (and 120B or Qwen 3.x for the larger slot). Enterprise committed-spend contracts can still get the Llama SKUs via sales.
What does Groq actually cost per million tokens now?
The official production table lists GPT-OSS 20B at $0.075 input and $0.30 output, and GPT-OSS 120B at $0.15 and $0.60. Whisper turbo is $0.04 per audio hour. Cached input is 50% off; Batch is 50% off and does not stack with cache.
Related tools
Last updated: 2026-09-21 · Reviewed by AIKetra editors · Domain registered 2007 (19 years old) · How we evaluate tools · Embed a Featured badge
T
O
O
R