Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published 12 days ago • 69
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 12 days ago • 95
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 12 days ago • 83
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 12 days ago • 180
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 14 days ago • 121
Running on CPU Upgrade Featured 107 H3 Acceleration Arena 🥇 107 Blind A/B ranking of MiniMax-H3 acceleration variants
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published 14 days ago • 51
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 14 days ago • 104
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 15 days ago • 59
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 15 days ago • 154
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling Paper • 2608.29335 • Published 17 days ago • 72
FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree Text-to-Video • 35B • Updated 10 days ago • 272k • 304
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation Paper • 2608.19098 • Published 27 days ago • 23