GPU cloud · 8 GPU models tracked
Hugging Face Inference Endpoints
$0.50–$9.25per GPU-hour, the cheapest row to the dearest
A dedicated inference endpoint on the cloud named in the region column, billed by the minute at the hourly rate while it runs, CPU and memory included. It serves a model; it is not a machine you log into.
As of , Hugging Face Inference Endpoints lists 8 GPU models we track, from $0.50 per GPU-hour (NVIDIA T4, on-demand) to $9.25 (NVIDIA B200, on-demand).
| Website | huggingface.co/ |
|---|---|
| Pricing page | huggingface.co/pricing |
| Regions and tiers | aws — as labelled in the rates below |
| Models tracked | 8 |
Price by GPU
As of , per GPU-hour, cheapest tier first named:
- Hugging Face Inference Endpoints lists the NVIDIA B200 from $9.25 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: 36% above the median rate card ($6.79 across 27 providers); 21 of the 27 providers we track list a lower on-demand rate, the lowest Deep Infra at $3.69 (Deep Infra’s rates, Hugging Face Inference Endpoints vs Deep Infra).
- Hugging Face Inference Endpoints lists the NVIDIA H200 from $5.00 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: 14% above the median rate card ($4.38 across 38 providers); 28 of the 38 providers we track list a lower on-demand rate, the lowest Beam at $2.09 (Beam’s rates, Hugging Face Inference Endpoints vs Beam).
- Hugging Face Inference Endpoints lists the NVIDIA H100 80GB (form factor not stated) from $4.50 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: 61% above the median rate card ($2.79 across 19 providers); 14 of the 19 providers we track list a lower on-demand rate, the lowest Gcore at $1.74 (Gcore’s rates, Hugging Face Inference Endpoints vs Gcore).
- Hugging Face Inference Endpoints lists the NVIDIA RTX PRO 6000 Blackwell from $2.75 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: 25% above the median rate card ($2.20 across 27 providers); 17 of the 27 providers we track list a lower on-demand rate, the lowest Beam at $1.09 (Beam’s rates, Hugging Face Inference Endpoints vs Beam).
- Hugging Face Inference Endpoints lists the NVIDIA A100 80GB from $2.50 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: 38% above the median rate card ($1.81 across 40 providers); 28 of the 40 providers we track list a lower on-demand rate, the lowest Deep Infra at $0.89 (Deep Infra’s rates, Hugging Face Inference Endpoints vs Deep Infra).
- Hugging Face Inference Endpoints lists the NVIDIA L40S from $1.80 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: 29% above the median rate card ($1.40 across 28 providers); 19 of the 28 providers we track list a lower on-demand rate, the lowest Beam at $0.76 (Beam’s rates, Hugging Face Inference Endpoints vs Beam).
- Hugging Face Inference Endpoints lists the NVIDIA L4 from $0.80 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: at the median rate card ($0.80 across 17 providers); 9 of the 17 providers we track list a lower on-demand rate, the lowest Seeweb at $0.43 (Seeweb’s rates, Hugging Face Inference Endpoints vs Seeweb).
- Hugging Face Inference Endpoints lists the NVIDIA T4 from $0.50 per GPU-hour on-demand (aws). The on-demand rate is unchanged since first read, Sep 25, 2026. Against the market: 5% below the median rate card ($0.53 across 9 providers); 2 of the 9 providers we track list a lower on-demand rate, the lowest immers.cloud at $0.23 (immers.cloud’s rates, Hugging Face Inference Endpoints vs immers.cloud).
All listed prices
One table per price type, never mixed. Every rate normalized to one GPU for one hour.
On-demand
| GPU | USD / GPU-hour | 7-day change | Type | Region | Min. commitment | Source, captured | Deploy |
|---|---|---|---|---|---|---|---|
| NVIDIA T4 | $0.50 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
| NVIDIA L4 | $0.80 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
| NVIDIA L40S | $1.80 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
| NVIDIA A100 80GB | $2.50 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
| NVIDIA RTX PRO 6000 Blackwell | $2.75 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
| NVIDIA H100 80GB (form factor not stated) | $4.50 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
| NVIDIA H200 | $5.00 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
| NVIDIA B200 | $9.25 | — | on-demand | aws | — | Price page unchanged since first read, Sep 25, 2026 | Deploy ↗ |
Every rate on this page as data: /providers/huggingface/prices.json — the same fields as prices.json, narrowed to this provider. CC BY 4.0, no key, no rate limit. On your own page: embed this table (/embed/providers/huggingface, an iframe, no script).
Compare Hugging Face Inference Endpoints with
Every provider that lists at least two of the same GPUs on-demand, most shared GPUs first; each opens the side-by-side page.
- Cerebrium 8 shared GPUs
- Amazon Web Services 7 shared GPUs
- E2E Networks 7 shared GPUs
- EmpirioLabs 7 shared GPUs
- Koyeb 7 shared GPUs
- Modal 7 shared GPUs
- RunPod 6 shared GPUs
- Vast.ai 6 shared GPUs
- Baseten 5 shared GPUs
- Beam 5 shared GPUs
- CoreWeave 5 shared GPUs
- Google Cloud 5 shared GPUs
- HexGrid 5 shared GPUs
- Lyceum 5 shared GPUs
- Oracle Cloud (OCI) 5 shared GPUs
- Seeweb 5 shared GPUs
- Verda 5 shared GPUs
- Daytona 4 shared GPUs
- Deep Infra 4 shared GPUs
- fal 4 shared GPUs
- Gcore 4 shared GPUs
- Hyperstack 4 shared GPUs
- immers.cloud 4 shared GPUs
- Jarvislabs 4 shared GPUs
- Massed Compute 4 shared GPUs
- Microsoft Azure 4 shared GPUs
- Nebius 4 shared GPUs
- Northflank 4 shared GPUs
- OVHcloud 4 shared GPUs
- packet.ai 4 shared GPUs
- Replicate 4 shared GPUs
- Sesterce 4 shared GPUs
- Civo 3 shared GPUs
- Crusoe 3 shared GPUs
- GMI Cloud 3 shared GPUs
- Leafcloud 3 shared GPUs
- Neysa 3 shared GPUs
- Selectel 3 shared GPUs
- Cloud.ru Evolution 2 shared GPUs
- Cudo Compute 2 shared GPUs
- Hyperbolic 2 shared GPUs
- Lambda 2 shared GPUs
- Paperspace 2 shared GPUs
- Scaleway 2 shared GPUs
- Together AI 2 shared GPUs
Data updated: