Third-party fidelity measurement: KL(reference ‖ NVFP4) = 0.0548 nats
#14 opened 26 days ago
by
malaiwah
nvidia daddy , please update GLM5.3-flash-nvfp4
👍 1
1
#12 opened about 1 month ago
by
mickeypro
Waiting for nvidia/GLM-5.3-NVFP4
1
#10 opened about 1 month ago
by
ghostplant
Even though the accuracy seems comparable, there is significant token inflation on a per task basis
#8 opened 2 months ago
by
jakubjaniak
tool_choice: "required" causes xgrammar FSM crash / infinite hang with GLM-5.2 on vLLM 0.24.0
1
#7 opened 2 months ago
by
paolovic
How is the actual programming performance?
#6 opened 3 months ago
by
Artom
Some accuracy benchmark results are not as good as GLM-5.2-FP8
#5 opened 3 months ago
by
Tianjiu
Does it support MTP?
2
#4 opened 3 months ago
by
yz342
Which version of VLLM should I use to run this checkpoint?
3
#2 opened 3 months ago
by
gameofdimension
Can we use this model with nvfp4 kv cache?
6
#1 opened 3 months ago
by
positiveone