This blog was dark for a while. Not dead - parked. We are back. First post after the pause is not a launch. It is a spreadsheet with teeth.

Cloud GPU listings mix dense and sparse BF16 numbers like they are the same unit. They are not. NVIDIA datasheets often lead with 2:4 sparsity (roughly 2x dense). Sort by that mix and you pick the wrong card for a long run 😭. We re-checked VRAM, memory bandwidth, architecture, and BF16 dense TFLOPS against vendor datasheets. Hourly prices were already trusted (in Runpod). Everything else got rebuilt.

The dense / sparse trap

Training almost never uses structured sparsity the way the marketing peak assumes. If you pay for FLOPs, you should compare BF16 dense (FP32 accumulate). Sparse peaks belong in a footnote, not in the ranking column.

Biggest corrections from the first pass:

  • RTX PRO 4000 is ~161 dense TFLOPS, not ~358. Bandwidth is 672 GB/s.
  • RTX 5090 419 is sparse; dense is 209.5.
  • H100 NVL is 835.5 dense, not the SXM 989 number copied across SKUs.
  • L40 dense is 181, not the L40S 362 figure.
  • H200 is 141 GB, not 143. SXM and NVL share the same die and HBM3e stack.
  • B200 bandwidth is 7.7 TB/s (often rounded to 8). B300 keeps ~2250 BF16 dense; the Ultra bump is mostly FP4 and 288 GB.
  • MI300X 1307 BF16 dense was already right - and that moves it to the top of $/FLOP once everyone else is densified.

Ada workstation cards were the worst offenders: Tensor numbers in the PDF are often FP8-with-sparsity. Divide by 8 and you get a usable BF16 dense estimate (RTX 4000 Ada ~41, RTX 2000 Ada ~24).

How to pick a GPU for a long run

Fixed work W (tokens, epochs, whatever). Price p in $/h. Peak dense BF16 F. Bandwidth B. Arithmetic intensity I (FLOPs per byte). Utilization η (MFU, often 0.3-0.5 in training).

// cost cost = p · W / F_eff
F_eff = η · min(F, I · B)
pick argmin p / F_eff subject to VRAM ≥ model + opt + acts

That collapses to two sorts:

  • Compute-bound (big-batch training): minimize $ per BF16 TFLOP.
  • Memory-bound (decode, tiny batches): minimize $ per TB/s.

If the model does not fit, multiply by a parallel tax. PCIe without NVLink is ugly: TP=2 on a 70B BF16 decode can lose 35-55% vs a single fat card. Consumer GDDR also has no ECC - a bit flip in a 48-hour run is a real failure mode. Checkpoint.

Peak is not MFU. L40S looks cheap on 362 TFLOPS sitting on 0.86 TB/s until the kernel is bandwidth-starved. MI300X wins on paper if ROCm actually delivers; measured FP16/BF16 is often 45-85% of peak depending on the stack.

The table

Hourly $ kept as-is. Specs renormalized to BF16 dense. Sorted by $ per BF16 TFLOP. Scroll sideways if your viewport is not a cinema screen ... lol.

GPUArchitectureVRAMData throughput in TB/sBF16 TFLOPSPrice in $ per hourPrice in $ per TB data throughputPrice in $ per BF16 TFLOP
MI300XCDNA 3192 GB5.313072.390.0001250.00000051
RTX A5000Ampere24 GB11110.270.0000980.00000068
RTX A4500Ampere20 GB0.64~950.250.0001090.00000073
L40SAda Lovelace48 GB13620.990.0003180.00000076
A40Ampere48 GB11500.440.0001760.00000082
B200Blackwell180 GB7.722506.790.0002450.00000084
RTX A4000Ampere16 GB0~770.250.0001550.00000091
H100 SXMHopper80 GB3.359893.290.0002730.00000092
RTX PRO 4500 (+SE)Blackwell32 GB1~2150.720.0002230.00000093
RTX A6000Ampere48 GB11550.530.0001920.00000095
B300Blackwell Ultra288 GB822507.890.0002740.00000097
RTX PRO 4000Blackwell24 GB1~1610.570.0002360.00000098
RTX PRO 6000 WKBlackwell96 GB1.79~5001.890.0002930.00000105
H100 NVLHopper94 GB3.98363.190.0002270.00000106
H100 PCIeHopper80 GB27562.890.0004010.00000106
H200 NVLHopper141 GB4.89893.790.0002190.00000106
L4Ada Lovelace24 GB0.31210.490.0004540.00000112
RTX PRO 6000 SEBlackwell96 GB1.79~5002.090.0003240.00000116
PRO 6000 MIG 48GBBlackwell (MIG)48 GB~0.90~2501.090.0003380.00000121
A100 PCIeAmpere80 GB1.943121.390.0002000.00000124
RTX 4090Ada Lovelace24 GB1.011650.740.0002040.00000124
L40Ada Lovelace48 GB11810.820.0002640.00000126
RTX 6000 AdaAda Lovelace48 GB0.96182.50.840.0002430.00000128
H200 SXMHopper141 GB4.89894.590.0002660.00000129
PRO 6000 MIG 24GBBlackwell (MIG)24 GB~0.45~1250.590.0003660.00000131
RTX 5090Blackwell32 GB1.79209.50.990.0001530.00000131
A100 SXMAmpere80 GB2.043121.590.0002170.00000142
RTX 4000 AdaAda Lovelace20 GB0.36~410.280.0002160.00000190
RTX 3090Ampere24 GB1710.50.0001480.00000196
RTX 2000 AdaAda Lovelace16 GB0~240.240.0002980.00000278

What we would actually rent

Jobs that fit in 20-24 GB: A4500 / A5000 still crush $/FLOP. Real training: MI300X if the software stack is ROCm-shaped, otherwise B200 or H100 SXM among NVIDIA. H200 NVL is the sleeper - H100-SXM compute, 141 GB, 4.8 TB/s, almost the same $ per dense TFLOP as H100 PCIe/NVL.

Do not sort sparse marketing peaks and call it research. We did that once. Then we fixed it. Life is crazy 😭😂

// notes Prices --> as provided ($/h)
FLOPS --> BF16 dense, vendor datasheets
BW / VRAM --> NVIDIA / AMD product pages
Formula --> roofline, not vibes
#we're-back #gpu #bf16 #roofline #cloud #mi300x #blackwell #research