GPU cloud · 3 GPU models tracked
Inferless
$0.66–$5.36per GPU-hour, the cheapest row to the dearest
Serverless inference endpoints: billed by the second, scaled to zero when idle. Each price is one whole card with the vCPUs and memory listed beside it. The same page also prices a shared half of each card, which this site does not publish - half a card is not a card.
As of , Inferless lists 3 GPU models we track, from $0.66 per GPU-hour (NVIDIA T4, serverless) to $5.36 (NVIDIA A100 80GB, serverless).
- Cheapest row$0.66NVIDIA T4 · serverless
- A100 80GB$5.36on-demand · 184% above the median $1.89 (53)
| Website | www.inferless.com/ |
|---|---|
| Pricing page | www.inferless.com/pricing |
| Regions and tiers | No region or tier named in the rates below: one global rate card |
| Models tracked | 3 |
Price by GPU
As of , per GPU-hour: Inferless’s on-demand rate for each card it lists, a cheaper market it also sells, and where the on-demand rate stands among the providers we track.
| GPU | Inferless on-demand, per GPU-hour | Against the market (median rate card) |
|---|---|---|
| NVIDIA A100 80GB | $5.36 captured | 184% above the median $1.89 · 53 providers 51 of 53 list lower · lowest Deep Infra $0.89 |
| NVIDIA A10 | $1.22 captured | at the median $1.22 · 9 providers 4 of 9 list lower · lowest Intelion Cloud $0.32 GPU only · compare |
| NVIDIA T4 | $0.66 captured | 25% above the median $0.53 · 13 providers 10 of 13 list lower · lowest Edgevana $0.15 |
The same, in words
- Inferless lists the NVIDIA A100 80GB from $5.36 per GPU-hour on-demand. The on-demand rate is unchanged since first read, Sep 30, 2026 (7-day change: none). Against the market: 184% above the median rate card ($1.89 across 53 providers); 51 of the 53 providers we track list a lower on-demand rate, the lowest Deep Infra at $0.89 (Deep Infra’s rates).
- Inferless lists the NVIDIA A10 from $1.22 per GPU-hour on-demand. The on-demand rate is unchanged since first read, Sep 30, 2026 (7-day change: none). Against the market: at the median rate card ($1.22 across 9 providers); 4 of the 9 providers we track list a lower on-demand rate, the lowest Intelion Cloud at $0.32 (Intelion Cloud’s rates, Inferless vs Intelion Cloud).
- Inferless lists the NVIDIA T4 from $0.66 per GPU-hour on-demand. The on-demand rate is unchanged since first read, Sep 30, 2026 (7-day change: none). Against the market: 25% above the median rate card ($0.53 across 13 providers); 10 of the 13 providers we track list a lower on-demand rate, the lowest Edgevana at $0.15 (Edgevana’s rates).
All listed prices
One table per price type, never mixed. Every rate normalized to one GPU for one hour; a price opens the page it was read from, and 7d marks a rate that moved in the past seven days. The day each figure was first listed is in this page’s data, below.
On-demand
| GPU · what is sold | $ / GPU-hour | Read | Source |
|---|---|---|---|
| NVIDIA T4serverless | $0.66 | captured | Deploy ↗ |
| NVIDIA A10serverless | $1.22 | captured | Deploy ↗ |
| NVIDIA A100 80GBserverless | $5.36 | captured | Deploy ↗ |
Every rate on this page as data: /providers/inferless/prices.json — the same fields as prices.json, narrowed to this provider. CC BY 4.0, no key, no rate limit. On your own page: embed this table (/embed/providers/inferless, an iframe, no script).
Compare Inferless with
Every provider that lists at least two of the same GPUs on-demand, most shared GPUs first; each opens the side-by-side page.
- Cerebrium 3 shared GPUs
- immers.cloud 3 shared GPUs
- Microsoft Azure 3 shared GPUs
- Modal 3 shared GPUs
- Amazon Web Services 2 shared GPUs
- Baseten 2 shared GPUs
- Google Cloud 2 shared GPUs
- Hugging Face 2 shared GPUs
- Intelion Cloud 2 shared GPUs
- Lambda 2 shared GPUs
- Oracle Cloud (OCI) 2 shared GPUs
- Replicate 2 shared GPUs
- Sesterce 2 shared GPUs
- Yandex Cloud 2 shared GPUs
Data updated: