Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 6 days ago • 24
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure Paper • 2609.30074 • Published 8 days ago • 2
Self-Play Search Distillation for Large Language Model Reasoning Paper • 2609.30936 • Published 7 days ago • 4
WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing Paper • 2608.18486 • Published 5 days ago • 4
Allspark: Weak to Strong Transfer via Alternating Chain of Thought Paper • 2609.32913 • Published 6 days ago • 4
Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization Paper • 2606.29521 • Published 11 days ago • 5
G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation Paper • 2609.31009 • Published 7 days ago • 4
Reinforcing Agentic Creativity in Scientific Ideation with Night Science Paper • 2609.35706 • Published 4 days ago • 4
Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models Paper • 2609.34065 • Published 4 days ago • 3
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 5 days ago • 4
KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation Paper • 2609.33074 • Published 5 days ago • 5
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models Paper • 2609.35578 • Published 4 days ago • 4
WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models Paper • 2609.23033 • Published 6 days ago • 7
Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation Paper • 2609.34848 • Published 4 days ago • 6
Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction Paper • 2609.32353 • Published 6 days ago • 8
ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis Paper • 2609.32630 • Published 6 days ago • 14
Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers Paper • 2609.32814 • Published 6 days ago • 21