patrickbdevaney commited on
Commit
33cb67e
·
verified ·
1 Parent(s): 871d152

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +6 -0
README.md CHANGED
@@ -19,6 +19,12 @@ NVIDIA Jetson AGX Thor (117 GiB unified memory). 86.1 GiB across 65 shards.
19
 
20
  Vision, audio and video input are preserved; `audio_tokenizer/` ships with the checkpoint.
21
 
 
 
 
 
 
 
22
  ## How the experts were chosen
23
 
24
  Not by activation frequency. Expert saliency was accumulated over a calibration corpus and the
 
19
 
20
  Vision, audio and video input are preserved; `audio_tokenizer/` ships with the checkpoint.
21
 
22
+ ## GGUF Quantizations (llama.cpp)
23
+
24
+ Official llama.cpp GGUF quantizations (including native **MXFP4_MOE**, optimal hybrid **Q2_K**, multimodal **mmproj**, and speculative **mtp** draft towers) are available at:
25
+ 👉 **[patrickbdevaney/MiMo-V2.6-Flash-REAP50-GGUF](https://huggingface.co/patrickbdevaney/MiMo-V2.6-Flash-REAP50-GGUF)**
26
+
27
+
28
  ## How the experts were chosen
29
 
30
  Not by activation frequency. Expert saliency was accumulated over a calibration corpus and the