GRACE: Generation-aware latent compression for efficient video generation
Abstract
Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.
Community
This work was done while the KAIST AI authors were interns at Kakao Corp..๐ฆ
๐ฌ Project page (with extensive qualitative results) : https://cvlab-kaist.github.io/GRACE/
๐ป Code: https://github.com/cvlab-kaist/GRACE
TL;DR: GRACE compresses the VAE of Wan2.1-I2V-14B so the DiT runs on 8ร fewer tokens, giving 11.1ร lower latency at 480ร832ร81 while matching the original pipeline's quality on VBench.
How: We keep the frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression. The compressed latent is aligned with the pretrained one in the frozen DiT's feature space, so the autoencoder is optimized for generation rather than reconstruction alone. After that, the DiT only needs lightweight fine-tuning with asymmetric denoising, where the base is denoised ahead of the residual.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence (2026)
- V-RAE: Rethinking Video Latent Spaces for Generation (2026)
- LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation (2026)
- Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation (2026)
- Enhancing Autoregressive Video Generation via Representation Adversarial Distillation (2026)
- FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders (2026)
- When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.10524 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper