AlexCaswen/Nemotron-3-Super-120B-A12B-MTP-GGUF Text Generation โข 4B โข Updated 11 days ago โข 198 โข 1
deepseek-ai/DeepSeek-V4-Flash-0731 Text Generation โข 304B โข Updated 20 days ago โข 2.83M โข โข 3.59k
Running Agents 89 Accurate GGUF Memory Calculator ๐ 89 Calculate memory for GGUF models using GPU layers + context
view post Post 2989 Good news, llama.cpp seems to be close to supporting MTP on qwen models. Bad news, every single gguf will have to be redone when it is. See translation 1 reply ยท ๐ 15 15 + Reply