Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 2 days ago • 14
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization Paper • 2609.15991 • Published 6 days ago • 11
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 5 days ago • 8
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 2 days ago • 115
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention Paper • 2609.24797 • Published 3 days ago • 7
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 3 days ago • 42
LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay Paper • 2609.25053 • Published 17 days ago • 4
PACT: From Credit Assignment to Critic Alignment Paper • 2609.26355 • Published 2 days ago • 14
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback Paper • 2606.14368 • Published Jun 12 • 3
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 3 days ago • 12
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 20 days ago • 113
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 6 days ago • 130
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup Paper • 2609.15126 • Published 10 days ago • 12
Calibrating Teacher--Student Discrepancy for On-Policy Distillation Paper • 2609.21619 • Published 6 days ago • 13
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling Paper • 2609.19499 • Published 8 days ago • 37
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 7 days ago • 42
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 7 days ago • 129
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 7 days ago • 108