Running 2 RIS-Kernel Long Context Inference Demo 🔬 2 Run long-context LLM inference with O(N log N) sparse attention
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF Text Generation • 33B • Updated Aug 13 • 82.5k • 98
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 Any-to-Any • 33B • Updated 22 days ago • 286k • 424
Running Featured 244 Gemma 4 WebGPU 🚀 244 Run Gemma 4 locally in-browser on WebGPU w/ Transformers.js
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 Text Generation • 124B • Updated 22 days ago • 84.9k • 279