Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 1 day ago • 12
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 6 days ago • 108
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 6 days ago • 171
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 6 days ago • 69
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 9 days ago • 50
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 7 days ago • 61
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 7 days ago • 76
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 12 days ago • 65
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 15 days ago • 79
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 9 days ago • 243
HyQuant: Hybrid-Precision Quantization for LLM Attention Paper • 2608.27875 • Published 26 days ago • 30
Erase-then-Delta Attention: Decoupling Erase and Write Addresses in Delta-Rule Linear Attention Paper • 2606.26560 • Published Jun 25 • 3
Kalman Delta Networks: Uncertainty-aware Associative Memory Paper • 2609.07816 • Published 16 days ago • 29
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation Paper • 2609.02998 • Published 21 days ago • 17
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Paper • 2609.05295 • Published 19 days ago • 17