GPUs for LLM inference
As of — An 8B model at 16-bit: from $0.16 per GPU-hour on the NVIDIA V100 32GB (Geodd), 40 cards qualify; A 70B model at 8-bit: from $0.65 per GPU-hour on the NVIDIA RTX PRO 6000 Blackwell (Hostinger), 18 cards qualify; A 70B model at 16-bit on one GPU: from $1.71 per GPU-hour on the AMD Instinct MI300X (TensorWave), 10 cards qualify. Rates are each card’s cheapest live on-demand rate card across the providers we track, read hourly, every one linked to the page it was read from.
Serving a model keeps its weights in GPU memory: parameters times bytes per parameter - 2 bytes at 16-bit (FP16, BF16), 1 byte at 8-bit (FP8, INT8). The KV cache of every request being served comes on top, and grows with context length and concurrency, so the figures below are floors, not sizes.
An 8B model at 16-bit
8 billion parameters × 2 bytes = 16 GB of weights: a card with 24 GB or more.
| Card, memory | Cheapest on-demand | Provider, read | Median rate card | $ per GB-hour | Deploy |
|---|---|---|---|---|---|
| NVIDIA V100 32GB · 32 GB | $0.16 | captured | $0.76 8 providers | $0.005 | Deploy ↗ |
| NVIDIA GeForce RTX 3090 · 24 GB | $0.17 | captured | $0.50 8 providers | $0.007 | Deploy ↗ |
| NVIDIA TITAN RTX · 24 GB | $0.24 | captured | — 1 provider | $0.010 | Deploy ↗ |
| NVIDIA Quadro RTX 6000 · 24 GB | $0.27 GPU only | captured | $0.51 4 providers | $0.011 | Deploy ↗ |
| NVIDIA RTX A5000 · 24 GB | $0.27 from | captured | $0.57 8 providers | $0.011 | Deploy ↗ |
| NVIDIA L4 · 24 GB | $0.27 GPU only | captured | $0.75 22 providers | $0.011 | Deploy ↗ |
| NVIDIA A30 · 24 GB | $0.32 | captured | $0.41 7 providers | $0.013 | Deploy ↗ |
| NVIDIA A10 · 24 GB | $0.32 GPU only | captured | $1.22 9 providers | $0.013 | Deploy ↗ |
| NVIDIA RTX A6000 · 48 GB | $0.35 | captured | $0.55 17 providers | $0.007 | Deploy ↗ |
| NVIDIA RTX 4090 · 24 GB | $0.39 | captured | $0.48 15 providers | $0.016 | Deploy ↗ |
30 more cards qualify, dearer: NVIDIA RTX PRO 4500 Blackwell $0.39, NVIDIA RTX 5000 Ada Generation $0.45, NVIDIA GeForce RTX 3090 Ti $0.46, NVIDIA L40 $0.53, NVIDIA RTX 5090 $0.54, NVIDIA RTX 4500 Ada Generation $0.54, NVIDIA A40 $0.59, NVIDIA RTX 6000 Ada $0.65, NVIDIA RTX PRO 6000 Blackwell $0.65, NVIDIA RTX PRO 5000 Blackwell $0.71, NVIDIA L40S $0.72, NVIDIA A100 40GB $0.76, NVIDIA A100 80GB $0.89, NVIDIA A10G $1.00, AMD Instinct MI300X $1.71, NVIDIA H100 80GB (form factor not stated) $1.71, NVIDIA H100 SXM $1.79, NVIDIA RTX PRO 6000 Blackwell Max-Q $1.80, NVIDIA H100 PCIe $1.83, NVIDIA H200 $2.09, AMD Instinct MI325X $2.25, NVIDIA GH200 $2.29, NVIDIA H100 NVL $2.92, AMD Instinct MI355X $2.95, NVIDIA H200 NVL $3.29, NVIDIA B200 $3.69, AMD Instinct MI350X $4.00, NVIDIA B300 $4.89, NVIDIA GB200 NVL72 $6.94, NVIDIA GB300 NVL72 $10.74.
A 70B model at 8-bit
70 billion parameters × 1 byte = 70 GB of weights: a card with 80 GB or more.
| Card, memory | Cheapest on-demand | Provider, read | Median rate card | $ per GB-hour | Deploy |
|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell · 96 GB | $0.65 from | captured | $2.20 35 providers | $0.007 | Deploy ↗ |
| NVIDIA A100 80GB · 80 GB | $0.89 | captured | $1.89 53 providers | $0.011 | Deploy ↗ |
| AMD Instinct MI300X · 192 GB | $1.71 from | captured | $2.81 8 providers | $0.009 | Deploy ↗ |
| NVIDIA H100 80GB (form factor not stated) · 80 GB | $1.71 | captured | $3.44 22 providers | $0.021 | Deploy ↗ |
| NVIDIA H100 SXM · 80 GB | $1.79 GPU only | captured | $3.55 33 providers | $0.022 | Deploy ↗ |
| NVIDIA RTX PRO 6000 Blackwell Max-Q · 96 GB | $1.80 | captured | $1.86 2 providers | $0.019 | Deploy ↗ |
| NVIDIA H100 PCIe · 80 GB | $1.83 from | captured | $2.82 10 providers | $0.023 | Deploy ↗ |
| NVIDIA H200 · 141 GB | $2.09 from | captured | $4.54 45 providers | $0.015 | Deploy ↗ |
| AMD Instinct MI325X · 256 GB | $2.25 | captured | $2.25 3 providers | $0.009 | Deploy ↗ |
| NVIDIA GH200 · 96 GB | $2.29 | captured | $3.87 3 providers | $0.024 | Deploy ↗ |
8 more cards qualify, dearer: NVIDIA H100 NVL $2.92, AMD Instinct MI355X $2.95, NVIDIA H200 NVL $3.29, NVIDIA B200 $3.69, AMD Instinct MI350X $4.00, NVIDIA B300 $4.89, NVIDIA GB200 NVL72 $6.94, NVIDIA GB300 NVL72 $10.74.
A 70B model at 16-bit on one GPU
70 billion parameters × 2 bytes = 140 GB of weights: a card with 141 GB or more (or the model split over two 80 GB cards).
| Card, memory | Cheapest on-demand | Provider, read | Median rate card | $ per GB-hour | Deploy |
|---|---|---|---|---|---|
| AMD Instinct MI300X · 192 GB | $1.71 from | captured | $2.81 8 providers | $0.009 | Deploy ↗ |
| NVIDIA H200 · 141 GB | $2.09 from | captured | $4.54 45 providers | $0.015 | Deploy ↗ |
| AMD Instinct MI325X · 256 GB | $2.25 | captured | $2.25 3 providers | $0.009 | Deploy ↗ |
| AMD Instinct MI355X · 288 GB | $2.95 from | captured | $5.99 3 providers | $0.010 | Deploy ↗ |
| NVIDIA H200 NVL · 141 GB | $3.29 from | captured | $3.50 6 providers | $0.023 | Deploy ↗ |
| NVIDIA B200 · 192 GB | $3.69 | captured | $7.20 32 providers | $0.019 | Deploy ↗ |
| AMD Instinct MI350X · 288 GB | $4.00 | captured | $5.08 2 providers | $0.014 | Deploy ↗ |
| NVIDIA B300 · 288 GB | $4.89 | captured | $8.99 20 providers | $0.017 | Deploy ↗ |
| NVIDIA GB200 NVL72 · 186 GB | $6.94 from | captured | $13.25 6 providers | $0.037 | Deploy ↗ |
| NVIDIA GB300 NVL72 · 288 GB | $10.74 | captured | $18.00 3 providers | $0.037 | Deploy ↗ |
How these figures are read
A card qualifies by its memory alone; speed, interconnect and software support are not ranked here. The rate is the card’s cheapest live on-demand rate card - each provider once, at its cheapest on-demand rate, community capacity and marketplace floors out, read within 48 hours - and the median is taken over the same rate cards, as on every card page. Memory figures are floors from the arithmetic above, not measurements: context length, batch size and the serving stack change them. Methodology.
Other workloads: GPUs for fine-tuning · GPUs for training · GPUs for image generation · Cheapest GPU memory per hour.
Data updated: