nota-ai/GLM-5.3-Nota-NVFP4-Global-Pruned-17.75 Text Generation • 625B • Updated 17 days ago • 751 • 18
nota-ai/Qwen3.8-2.4T-A95B-Nota-NVFP4-Global-Pruned-40 Text Generation • 1.5T • Updated about 1 month ago • 544 • 23
nota-ai/Nemotron-3.5-Lightning-30B-A3B-NVFP4-Global-Pruned-15 Text Generation • 16B • Updated Aug 17 • 234 • 20
Efficient NVIDIA Model Family Collection Compressed NVIDIA Models (Nemotron and Cosmos) • 1 item • Updated 19 days ago • 1
Efficient Solar Open Family Collection Solar Open models officially optimized by Nota AI for Korea’s government-led Sovereign AI initiative as part of the Upstage consortium. • 7 items • Updated 19 days ago • 26
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B Paper • 2607.04244 • Published Jul 5 • 1
Efficient MoE-based LLM Collection Mixture-of-Experts Large Language Models with Advanced Quantization • 5 items • Updated Mar 11 • 25
Efficient Large Vision-Language Model Collection ERGO: LVLM trained with RL on efficiency objectives; https://github.com/nota-github/ERGO • 3 items • Updated Feb 22 • 27