facebook/dinov3-vitl16-pretrain-lvd1689m Image Feature Extraction • 0.3B • Updated Aug 19, 2025 • 748k • 477
view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 18 days ago • 192
view article Article Native-speed vLLM transformers modeling backend hmellor, lysandre • 28 days ago • 66
Running on CPU Upgrade Featured 3.25k The Smol Training Playbook 📚 3.25k The secrets to building world-class LLMs
DFlash: Block Diffusion for Flash Speculative Decoding Paper • 2602.06036 • Published Feb 5 • 90
CohereLabs/command-a-plus-05-2026-bf16 Image-Text-to-Text • 219B • Updated Jun 15 • 90.6k • • 142
CohereLabs/command-a-plus-05-2026-w4a4 Image-Text-to-Text • 126B • Updated Jun 16 • 4.73k • • 237