Find the cheapest GPU for the job.

Pick a model. We size it, match it to hardware, and compare the published rates we have archived evidence for.

Tracking
VerdaMassed ComputeAceCloudSaladCloudVast.aiAmazon Web ServicesMicrosoft AzureCoreWeaveCyfuture AIGPU.aiLambdaGPUhubExoscaleGeoddRunPod Community CloudRunPod Secure CloudScalewayHyperstackCloudzyCrusoe CloudNebius AI CloudTogether AIImpossible CloudCivoRuncrateDigitalOcean GPU DropletsKoyebSpheronThunder ComputeLiumLeafcloudBeamModalCanopy WaveLinode / AkamaiReplicateGcoreOblivusCerebriumOVHcloudSwissGPUpacket.aiJarvisLabsNeysaEnvergeElastxFalNorthflankHot AisleHetznerSesterceLatitude.sh
+4 more
Popular this week on Hugging Face
Llama 3.3 70B
meta-llama/Llama-3.3-70B-Instruct · 70B dense · 128k ctx · text generation
Precision
Context
Task
Memory required
169GB
Weights 140 GBKV cache 9.9 GBOverhead 18.8 GB
Weights plus KV cache at 32k context, batch size 1. Larger batches raise the KV cache proportionally.
estimatedSized from a transcribed parameter count and a KV coefficient. Paste this model's Hugging Face URL to size it from its own config instead.
Where it fits, cheapest first

How this ranking works

The order

Cheapest total per hour. The price is for the whole configuration — four cards at $0.40 each shows as $1.60.

Only what you can rent

Node sizes come from the offers themselves. If a provider rents a card ×1 and ×8, we won’t price ×3 — separate rentals don’t pool their memory.

The pills

  • No FP8 in hardware runs in FP16 instead — about twice the memory shown, so the fit may not hold.
  • No BF16 FP16 only. BF16 weights usually convert, but can overflow.
  • No FlashAttention-2 falls back to attention that uses more memory than estimated.
  • Pools over PCIe no NVLink, so cross-GPU traffic goes through host memory.
  • Interconnect not recorded this card ships with and without NVLink and the provider doesn’t say which. Ask before a multi-card run.
  • ROCm / Habana, not CUDA different runtime. CUDA kernels need a port.

What reorders the list

Only one thing: a card that can’t do your selected precision is ranked on memory it wouldn’t use, so it sorts below cards that can.

Nothing is ordered by speed. We publish no throughput figures because we measure none — so a cheap old card can still cost more for the same job.

Sources

Prices are scraped from each provider and trace to an archived copy of the page. Architecture and format support come from vendor documentation, maintained by hand — capability only, never performance.

44 other GPUs we track either cost more, or are not rented in a node big enough to hold this model — pooling VRAM needs the cards in one machine, so the sizes here are ones a provider actually sells.

See every GPU we track →Batch size 1. Compares published on-demand, per-GPU rates only — spot and per-node quotes are not ranked against them.