SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs
Abstract
Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64times over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.
Community
SlideDP is the first system to make host‑resident full‑parameter fine‑tuning scale efficiently under synchronous DP, by: 1. coordinating multiple GPUs over a single shared host state 2. overlapping CPU work with GPU computation at fine granularity 3. adapting communication routes to topology and workload 4. using GPU memory strategically to reduce exposed pipeline work. This transforms host‑memory‑centric fine‑tuning from a single‑GPU workaround into a high‑performance multi‑GPU training paradigm.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training (2026)
- TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training (2026)
- Towards Training Private LLMs: Exploring Fine-Tuning Language Models on Apple Silicon with RDMA over Thunderbolt (2026)
- Analytical Resource Management for Fine-grained MoE Computation-Communication Overlap (2026)
- HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training (2026)
- HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management (2026)
- Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper