R
What is Replicate?
Replicate hosts other people's (and your) models behind a single HTTP API. Most public models bill only the seconds they spend processing your request.
Official hardware list as of September 2026: CPU about $0.000025–$0.000100/sec; Nvidia T4 $0.000225/sec ($0.81/hour); L40S $0.000975/sec ($3.51/hour); A100 80GB $0.001400/sec ($5.04/hour); H100 and H200 $0.001525/sec ($5.49/hour). Multi-GPU rows scale linearly (8× H100 $0.012200/sec).
Private models and Deployments charge for setup and idle time while instances are online — a forgotten warm H100 is about $5.49 × 24 × 30 ≈ $4,000/month. Some language models are token-priced instead; the model's page lists the meter.
Cloudflare acquired Replicate; published per-second rates had not been the story changing as of early-2026 write-ups — still verify replicate.com/pricing.
A concrete walkthrough: 10,000 image gens without renting an H100
Use a public image model on a T4 or L40S. If each image takes 8 seconds on a T4, 10,000 images ≈ 22.2 GPU-hours × $0.81 ≈ $18.
On an H100 at 3 seconds you pay 8.3 hours × $5.49 ≈ $46 for speed you may not need. Do not create a Deployment with min replicas = 1 'for the launch week' unless you have measured cold-start pain — that T4 left on is $0.81 × 24 × 7 ≈ $136/week on top of the $18.
Fine-tunes that cannot use the fast-boot shared pool follow private-model billing: you pay while the box is up.
Key features
- Catalog of image, video, audio, and language models with a playground
- Public-model billing: active processing only; setup and idle are free
- Private models and Deployments: you pay while instances are up, including idle
- Hardware picker from T4 ($0.81/hr) to 8× H100 ($43.92/hr) on the public table
- Fine-tunes and custom Cog images if the catalog model is wrong
- Scale-to-zero on traffic-shaped deployments if you configure it; misconfig = rent
How to get started
- Create an account, add a card, and run a public model in the playground first
- Read the model's 'Run time and cost' line — that hardware is your meter
- Call the API; keep using public models until you need isolation or a pinned version
- Create a Deployment only after you know QPS; set min instances to 0 unless latency requires warm GPUs
- Turn idle private instances off — Replicate will not do that as a courtesy
Replicate pricing
No monthly platform fee. September 2026 official hardware: T4 $0.81/hr, L40S $3.51/hr, A100 80GB $5.04/hr, H100/H200 $5.49/hr, billed per second.
Public models: processing time only. Private models and Deployments: setup + idle + processing.
Some LLMs bill per token. Check each model page.
| Plan | Price | Best for |
|---|---|---|
| Public model | Hardware × seconds busy | Prototypes and spiky traffic |
| T4 | $0.000225/sec ($0.81/hr) | Light vision / small models |
| A100 80GB | $0.001400/sec ($5.04/hr) | Heavier image or mid LLMs |
| H100 / H200 | $0.001525/sec ($5.49/hr) | Fast large models; deadly if left warm |
| Private / Deployment | Same rates while online | Isolation — budget idle hours |
Our take: Stay on public models until you have a latency SLA. The first production mistake is a min-1 H100 Deployment 'just to be safe'. That is a rent, not an API.
Prices verified 2026-09. Plans change often — confirm on the official site above before you buy.
Use cases
- Calling Flux, Whisper, or an SD checkpoint without buying a GPU
- A/B testing three community models in a weekend
- Fine-tuning a public image model on a small brand set
- A production image endpoint that scales from zero
Strengths and tradeoffs
Where it stands out
- Fastest path from 'I saw this model on Twitter' to an API call
- Public-model idle is $0 — the opposite of a forgotten cloud VM
- Hardware prices are on one page, not hidden behind a sales PDF
Tradeoffs
- Private/Deployment idle is easy to ignore until the invoice
- Catalog quality varies; a trending model is not a supported product
- Always-on LLM traffic is often cheaper on Together or a reserved GPU
Replicate vs. alternatives
| Tool | How it compares |
|---|---|
| Replicate | One API for open models — you pay GPU-seconds, and a warm private deploy is a 24/7 rent. |
| Hugging Face | Hub + Inference Endpoints if you already live in model repos and want org billing. |
| Together AI | Better default for always-on open LLMs with token pricing. |
| OpenAI API | Closed GPT models and a token invoice, not community checkpoints. |
Use Replicate to try and to serve spiky media models. Use Hugging Face when the artifact is a repo + community.
Use Together or Groq for token-priced text. Use OpenAI when you want a closed model and SLAs.
FAQ
Is Replicate free to use?
Replicate's plans: Pay per second of hardware; public models skip idle/setup. Pricing and free-tier limits change over time, so check the official site above for the latest details.
What is Replicate used for?
One API for open models — you pay GPU-seconds, and a warm private deploy is a 24/7 rent.
Who is Replicate best for?
Replicate is a good fit for calling Flux, Whisper, or an SD checkpoint without buying a GPU, or a/B testing three community models in a weekend.
What are the alternatives to Replicate?
Commonly compared alternatives include Hugging Face, Together AI, and OpenAI API, see the comparison above for how they differ.
How much does Replicate cost?
You pay per second of the hardware the model uses. A T4 is $0.81/hour, an H100 is $5.49/hour. Public catalog models only bill while they run your request. A private or deployed instance also bills idle time.
Do I pay when nobody is calling my public model?
No. Replicate's billing docs say public models bill only active processing. Private models and Deployments bill setup and idle while instances are online.
How much is an H100 on Replicate?
The official table lists $0.001525 per second, $5.49 per hour, for a single H100 or H200. Eight H100s are $0.012200/sec ($43.92/hour).
Related tools
Last updated: 2026-09-21 · Reviewed by AIKetra editors · Domain registered 1998 (28 years old) · How we evaluate tools · Embed a Featured badge
H
T
O