view article Article LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation LiquidAI • about 19 hours ago • 25
view reply Have you experimented with or tested different vocabulary sizes to see if the vocab size directly impacts the overtraining threshold and benchmark peak?
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 147
view article Article Distillation in 2026 (so far): which frontier models use it and how sergiopaniego • Jul 8 • 21
view article Article Making Knowledge Distillation Cheap Enough to Run at Scale MultiverseComputingCAI • 10 days ago • 32
view post Post 6214 Gemma 4 is now faster and much more accurate! 🚀Google made huge improvements to tool-calling and chat accuracy, reliability + speed.To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4 See translation 8 replies · 🚀 30 30 👍 15 15 😎 6 6 🤗 5 5 🤝 1 1 + Reply
view article Article VLX-Seek: Improving VLM Fine-Grained Perception via Region Reference Instead of Coordinate Generation omlab • Jun 27 • 14