Diffusion Single File
comfyui

Measured cost per clip on six rented GPUs (4090, L40S, H100 PCIe, RTX 5090): same weights, graph and seeds

#59
by susorokin - opened

Disclosure: I'm building a service around GPU cost measurement, so read me as an interested party. Every number below is from runs I paid for myself ($2.06 total for the six rows).

Job: text-to-video, one 5-second clip at 864x480, 20 steps, the stock ComfyUI T2V graph from v0.35.0 with the int8_convrot weights from this repo (DiT + Qwen3-VL encoder + VAEs, 67 GB total), no LoRA, same prompt and seeds on every card, torch cu130 everywhere. Three clips per card. "Steady" is clips 2 and 3 at the rate paid. "Session" is everything the provider billed, including the 67 GB download and the first cold clip, divided by three clips.

provider card host vCPU / RAM s per clip $ per clip, steady $ per clip, session
RunPod Community, $0.69/h RTX 5090 32 GB 16 / 94 GB 69 $0.0135 $0.049
RunPod Secure, $0.99/h RTX 5090 32 GB 15 / 117 GB 100 $0.028 $0.083
RunPod Secure, $0.74/h RTX 4090 24 GB 15 / 86 GB 92 $0.019 $0.07
Vast.ai spot, bid $0.40/h RTX 4090 24 GB 32 / 108 GB 93 $0.013 $0.17
Nebius preemptible, $0.92/h L40S 48 GB 24 / 94 GB 89 $0.024 $0.19
Hyperstack spot, $2.00/h H100 PCIe 80 GB 28 / 177 GB 67 $0.038 $0.13

Per step: 5090 2.63–2.88 s, H100 PCIe 2.99 s, L40S and both 4090s 3.72–3.97 s.

Notes:

  • The 5090 does not hold the whole DiT: VRAM sat at 31.8 GB through the sampler and 1–3 GB still crossed PCIe every step, against 9–10 GB on a 4090. The Secure 5090 lost 30 s per clip to the mp4 encode on that pod's CPU (38 s against 2 s on the Community pod).
  • Host RAM decides whether the job runs at all: the weights stage in RAM, peak 68 GB on the 4090 hosts. Most Community 5090 pods on offer had 46–54 GB.
  • Download speed dominates short sessions: 450 MB/s on the Vast host, 370 on Hyperstack, 210 on RunPod, 104 on Nebius; on Nebius the first clip took 11.5 minutes because the weights come back from a network disk.
  • Steady-state numbers exclude the text encoder (same prompt, cached): a new prompt adds 15–23 s per clip on the 24 GB cards and 6–9 s on the 5090.

Per-run table, method, caveats and raw CSVs: https://qrun.cloud/measurements

Update, 13 September: a seventh row. RunPod Secure, RTX PRO 6000 Blackwell Server Edition 96 GB, $2.09/h: 47 s per clip, 1.98 s per step, $0.0275 per clip steady and $0.164 over a three-clip session. Peak VRAM 63.7 GB: the DiT and the text encoder sit in the card together, nothing streams; text encode on the first clip drops to 3.5 s. Cheaper per clip than the H100 PCIe ($0.0375) despite the higher hourly rate. Three clips, one pod, one day.

Sign up or log in to comment