This blog was dark for a while. Not dead - parked. We are back. First post after the pause is not a launch. It is a spreadsheet with teeth.
Cloud GPU listings mix dense and sparse BF16 numbers like they are the same unit. They are not. NVIDIA datasheets often lead with 2:4 sparsity (roughly 2x dense). Sort by that mix and you pick the wrong card for a long run ðŸ˜. We re-checked VRAM, memory bandwidth, architecture, and BF16 dense TFLOPS against vendor datasheets. Hourly prices were already trusted (in Runpod). Everything else got rebuilt.
The dense / sparse trap
Training almost never uses structured sparsity the way the marketing peak assumes. If you pay for FLOPs, you should compare BF16 dense (FP32 accumulate). Sparse peaks belong in a footnote, not in the ranking column.
Biggest corrections from the first pass:
- RTX PRO 4000 is ~161 dense TFLOPS, not ~358. Bandwidth is 672 GB/s.
- RTX 5090 419 is sparse; dense is 209.5.
- H100 NVL is 835.5 dense, not the SXM 989 number copied across SKUs.
- L40 dense is 181, not the L40S 362 figure.
- H200 is 141 GB, not 143. SXM and NVL share the same die and HBM3e stack.
- B200 bandwidth is 7.7 TB/s (often rounded to 8). B300 keeps ~2250 BF16 dense; the Ultra bump is mostly FP4 and 288 GB.
- MI300X 1307 BF16 dense was already right - and that moves it to the top of $/FLOP once everyone else is densified.
Ada workstation cards were the worst offenders: Tensor numbers in the PDF are often FP8-with-sparsity. Divide by 8 and you get a usable BF16 dense estimate (RTX 4000 Ada ~41, RTX 2000 Ada ~24).
How to pick a GPU for a long run
Fixed work W (tokens, epochs, whatever). Price p in $/h. Peak dense BF16 F. Bandwidth B. Arithmetic intensity I (FLOPs per byte). Utilization η (MFU, often 0.3-0.5 in training).
F_eff = η · min(F, I · B)
pick argmin p / F_eff subject to VRAM ≥ model + opt + acts
That collapses to two sorts:
- Compute-bound (big-batch training): minimize $ per BF16 TFLOP.
- Memory-bound (decode, tiny batches): minimize $ per TB/s.
If the model does not fit, multiply by a parallel tax. PCIe without NVLink is ugly: TP=2 on a 70B BF16 decode can lose 35-55% vs a single fat card. Consumer GDDR also has no ECC - a bit flip in a 48-hour run is a real failure mode. Checkpoint.
Peak is not MFU. L40S looks cheap on 362 TFLOPS sitting on 0.86 TB/s until the kernel is bandwidth-starved. MI300X wins on paper if ROCm actually delivers; measured FP16/BF16 is often 45-85% of peak depending on the stack.
The table
Hourly $ kept as-is. Specs renormalized to BF16 dense. Sorted by $ per BF16 TFLOP. Scroll sideways if your viewport is not a cinema screen ... lol.
| GPU | Architecture | VRAM | Data throughput in TB/s | BF16 TFLOPS | Price in $ per hour | Price in $ per TB data throughput | Price in $ per BF16 TFLOP |
|---|---|---|---|---|---|---|---|
| MI300X | CDNA 3 | 192 GB | 5.3 | 1307 | 2.39 | 0.000125 | 0.00000051 |
| RTX A5000 | Ampere | 24 GB | 1 | 111 | 0.27 | 0.000098 | 0.00000068 |
| RTX A4500 | Ampere | 20 GB | 0.64 | ~95 | 0.25 | 0.000109 | 0.00000073 |
| L40S | Ada Lovelace | 48 GB | 1 | 362 | 0.99 | 0.000318 | 0.00000076 |
| A40 | Ampere | 48 GB | 1 | 150 | 0.44 | 0.000176 | 0.00000082 |
| B200 | Blackwell | 180 GB | 7.7 | 2250 | 6.79 | 0.000245 | 0.00000084 |
| RTX A4000 | Ampere | 16 GB | 0 | ~77 | 0.25 | 0.000155 | 0.00000091 |
| H100 SXM | Hopper | 80 GB | 3.35 | 989 | 3.29 | 0.000273 | 0.00000092 |
| RTX PRO 4500 (+SE) | Blackwell | 32 GB | 1 | ~215 | 0.72 | 0.000223 | 0.00000093 |
| RTX A6000 | Ampere | 48 GB | 1 | 155 | 0.53 | 0.000192 | 0.00000095 |
| B300 | Blackwell Ultra | 288 GB | 8 | 2250 | 7.89 | 0.000274 | 0.00000097 |
| RTX PRO 4000 | Blackwell | 24 GB | 1 | ~161 | 0.57 | 0.000236 | 0.00000098 |
| RTX PRO 6000 WK | Blackwell | 96 GB | 1.79 | ~500 | 1.89 | 0.000293 | 0.00000105 |
| H100 NVL | Hopper | 94 GB | 3.9 | 836 | 3.19 | 0.000227 | 0.00000106 |
| H100 PCIe | Hopper | 80 GB | 2 | 756 | 2.89 | 0.000401 | 0.00000106 |
| H200 NVL | Hopper | 141 GB | 4.8 | 989 | 3.79 | 0.000219 | 0.00000106 |
| L4 | Ada Lovelace | 24 GB | 0.3 | 121 | 0.49 | 0.000454 | 0.00000112 |
| RTX PRO 6000 SE | Blackwell | 96 GB | 1.79 | ~500 | 2.09 | 0.000324 | 0.00000116 |
| PRO 6000 MIG 48GB | Blackwell (MIG) | 48 GB | ~0.90 | ~250 | 1.09 | 0.000338 | 0.00000121 |
| A100 PCIe | Ampere | 80 GB | 1.94 | 312 | 1.39 | 0.000200 | 0.00000124 |
| RTX 4090 | Ada Lovelace | 24 GB | 1.01 | 165 | 0.74 | 0.000204 | 0.00000124 |
| L40 | Ada Lovelace | 48 GB | 1 | 181 | 0.82 | 0.000264 | 0.00000126 |
| RTX 6000 Ada | Ada Lovelace | 48 GB | 0.96 | 182.5 | 0.84 | 0.000243 | 0.00000128 |
| H200 SXM | Hopper | 141 GB | 4.8 | 989 | 4.59 | 0.000266 | 0.00000129 |
| PRO 6000 MIG 24GB | Blackwell (MIG) | 24 GB | ~0.45 | ~125 | 0.59 | 0.000366 | 0.00000131 |
| RTX 5090 | Blackwell | 32 GB | 1.79 | 209.5 | 0.99 | 0.000153 | 0.00000131 |
| A100 SXM | Ampere | 80 GB | 2.04 | 312 | 1.59 | 0.000217 | 0.00000142 |
| RTX 4000 Ada | Ada Lovelace | 20 GB | 0.36 | ~41 | 0.28 | 0.000216 | 0.00000190 |
| RTX 3090 | Ampere | 24 GB | 1 | 71 | 0.5 | 0.000148 | 0.00000196 |
| RTX 2000 Ada | Ada Lovelace | 16 GB | 0 | ~24 | 0.24 | 0.000298 | 0.00000278 |
What we would actually rent
Jobs that fit in 20-24 GB: A4500 / A5000 still crush $/FLOP. Real training: MI300X if the software stack is ROCm-shaped, otherwise B200 or H100 SXM among NVIDIA. H200 NVL is the sleeper - H100-SXM compute, 141 GB, 4.8 TB/s, almost the same $ per dense TFLOP as H100 PCIe/NVL.
Do not sort sparse marketing peaks and call it research. We did that once. Then we fixed it. Life is crazy ðŸ˜ðŸ˜‚
FLOPS --> BF16 dense, vendor datasheets
BW / VRAM --> NVIDIA / AMD product pages
Formula --> roofline, not vibes
SupraLabs_