GPU cloud · 4 GPU models tracked
Replicate
$0.81–$5.49per GPU-hour, the cheapest row to the dearest
Models run on hardware billed by the second - public models for the seconds a request runs, private ones for all the time their instances are up - shown here as an hour; each price is an instance with its CPU and memory. H200 and 4- or 8-GPU machines are sold by contract and are not listed.
As of , Replicate lists 4 GPU models we track, from $0.81 per GPU-hour (NVIDIA T4, on-demand) to $5.49 (NVIDIA H100 80GB (form factor not stated), on-demand).
| Website | replicate.com/ |
|---|---|
| Pricing page | replicate.com/pricing |
| Regions and tiers | No region or tier named in the rates below: one global rate card |
| Models tracked | 4 |
Price by GPU
As of , per GPU-hour, cheapest tier first named:
- Replicate lists the NVIDIA H100 80GB (form factor not stated) from $5.49 per GPU-hour on-demand. The on-demand rate is unchanged since first read, Sep 23, 2026. Against the market: 97% above the median rate card ($2.79 across 19 providers); 16 of the 19 providers we track list a lower on-demand rate, the lowest Gcore at $1.74 (Gcore’s rates, Replicate vs Gcore).
- Replicate lists the NVIDIA A100 80GB from $5.04 per GPU-hour on-demand. The on-demand rate is unchanged since first read, Sep 23, 2026. Against the market: 179% above the median rate card ($1.81 across 40 providers); 39 of the 40 providers we track list a lower on-demand rate, the lowest Deep Infra at $0.89 (Deep Infra’s rates, Replicate vs Deep Infra).
- Replicate lists the NVIDIA L40S from $3.51 per GPU-hour on-demand. The on-demand rate is unchanged since first read, Sep 23, 2026. Against the market: 152% above the median rate card ($1.40 across 28 providers); 27 of the 28 providers we track list a lower on-demand rate, the lowest Beam at $0.76 (Beam’s rates, Replicate vs Beam).
- Replicate lists the NVIDIA T4 from $0.81 per GPU-hour on-demand. The on-demand rate is unchanged since first read, Sep 23, 2026. Against the market: 54% above the median rate card ($0.53 across 9 providers); 8 of the 9 providers we track list a lower on-demand rate, the lowest immers.cloud at $0.23 (immers.cloud’s rates, Replicate vs immers.cloud).
All listed prices
One table per price type, never mixed. Every rate normalized to one GPU for one hour.
On-demand
| GPU | USD / GPU-hour | 7-day change | Type | Region | Min. commitment | Source, captured | Deploy |
|---|---|---|---|---|---|---|---|
| NVIDIA T4 | $0.81 | — | on-demand | — | — | Price page unchanged since first read, Sep 23, 2026 | Deploy ↗ |
| NVIDIA L40S | $3.51 | — | on-demand | — | — | Price page unchanged since first read, Sep 23, 2026 | Deploy ↗ |
| NVIDIA A100 80GB | $5.04 | — | on-demand | — | — | Price page unchanged since first read, Sep 23, 2026 | Deploy ↗ |
| NVIDIA H100 80GB (form factor not stated) | $5.49 | — | on-demand | — | — | Price page unchanged since first read, Sep 23, 2026 | Deploy ↗ |
Every rate on this page as data: /providers/replicate/prices.json — the same fields as prices.json, narrowed to this provider. CC BY 4.0, no key, no rate limit. On your own page: embed this table (/embed/providers/replicate, an iframe, no script).
Compare Replicate with
Every provider that lists at least two of the same GPUs on-demand, most shared GPUs first; each opens the side-by-side page.
- Cerebrium 4 shared GPUs
- Hugging Face Inference Endpoints 4 shared GPUs
- Amazon Web Services 3 shared GPUs
- Baseten 3 shared GPUs
- E2E Networks 3 shared GPUs
- EmpirioLabs 3 shared GPUs
- Gcore 3 shared GPUs
- immers.cloud 3 shared GPUs
- Koyeb 3 shared GPUs
- Lyceum 3 shared GPUs
- Massed Compute 3 shared GPUs
- Modal 3 shared GPUs
- Beam 2 shared GPUs
- Civo 2 shared GPUs
- Cloud.ru Evolution 2 shared GPUs
- CoreWeave 2 shared GPUs
- Crusoe 2 shared GPUs
- Cudo Compute 2 shared GPUs
- Deep Infra 2 shared GPUs
- Google Cloud 2 shared GPUs
- HexGrid 2 shared GPUs
- Leafcloud 2 shared GPUs
- Microsoft Azure 2 shared GPUs
- Northflank 2 shared GPUs
- Oracle Cloud (OCI) 2 shared GPUs
- OVHcloud 2 shared GPUs
- packet.ai 2 shared GPUs
- RunPod 2 shared GPUs
- Seeweb 2 shared GPUs
- Sesterce 2 shared GPUs
- Vast.ai 2 shared GPUs
- Verda 2 shared GPUs
Data updated: