Google Cloud vs Hugging Face Inference Endpoints
As of , Hugging Face Inference Endpoints lists the lower on-demand rate on 3 of the 5 GPUs both offer on-demand; Google Cloud wins on 2. On the priciest shared card both list on-demand, NVIDIA B200: Google Cloud $16.11 (us-west1) vs Hugging Face Inference Endpoints $9.25 (aws) per GPU-hour on-demand — Hugging Face Inference Endpoints is $6.86 (43%) cheaper.
Cloud GPU rental prices of Google Cloud and Hugging Face Inference Endpoints, compared per GPU-hour on the GPUs both list. Where a provider has several tiers or regions, its cheapest on-demand rate is used and labelled. Community capacity (RunPod Community Cloud) and marketplace floors are shown only where a provider has nothing else, and never decide which is cheaper.
On-demand, per GPU-hour
| GPU | Google Cloud | Hugging Face Inference Endpoints | Difference | Cheaper |
|---|---|---|---|---|
| NVIDIA B200 | $16.11 us-west1 · GPU only captured Deploy ↗ | $9.25 aws captured Deploy ↗ | $6.86 (43%) | Hugging Face Inference Endpoints |
| NVIDIA H200 | $9.31 us-west1 · GPU only captured Deploy ↗ | $5.00 aws captured Deploy ↗ | $4.31 (46%) | Hugging Face Inference Endpoints |
| NVIDIA A100 80GB | $3.93 us-central1 · GPU only captured Deploy ↗ | $2.50 aws captured Deploy ↗ | $1.43 (36%) | Hugging Face Inference Endpoints |
| NVIDIA L4 | $0.56 us-east4 · GPU only captured Deploy ↗ | $0.80 aws captured Deploy ↗ | $0.24 (30%) | Google Cloud |
| NVIDIA T4 | $0.35 us-east5 · GPU only captured Deploy ↗ | $0.50 aws captured Deploy ↗ | $0.15 (30%) | Google Cloud |
Spot / interruptible, per GPU-hour
| GPU | Google Cloud | Hugging Face Inference Endpoints |
|---|---|---|
| NVIDIA B200 | $1.45 captured | — |
| NVIDIA H200 | $5.57 captured | — |
| NVIDIA A100 80GB | $0.40 captured | — |
GPU by GPU
- NVIDIA B200: Google Cloud $16.11 (us-west1) vs Hugging Face Inference Endpoints $9.25 (aws) per GPU-hour on-demand — Hugging Face Inference Endpoints is $6.86 (43%) cheaper.
- NVIDIA H200: Google Cloud $9.31 (us-west1) vs Hugging Face Inference Endpoints $5.00 (aws) per GPU-hour on-demand — Hugging Face Inference Endpoints is $4.31 (46%) cheaper.
- NVIDIA A100 80GB: Google Cloud $3.93 (us-central1) vs Hugging Face Inference Endpoints $2.50 (aws) per GPU-hour on-demand — Hugging Face Inference Endpoints is $1.43 (36%) cheaper.
- NVIDIA L4: Google Cloud $0.56 (us-east4) vs Hugging Face Inference Endpoints $0.80 (aws) per GPU-hour on-demand — Google Cloud is $0.24 (30%) cheaper.
- NVIDIA T4: Google Cloud $0.35 (us-east5) vs Hugging Face Inference Endpoints $0.50 (aws) per GPU-hour on-demand — Google Cloud is $0.15 (30%) cheaper.
Not compared: GPUs only one of the two lists, reserved/committed tiers, and anything either provider prices only on request. See each provider's page for its full list.
Data updated: