Jan Reges
janreges3
ยท
AI & ML interests
None yet
Organizations
None yet
NVFP4 vs FP8 throughput on an RTX 6000 Pro 96 GB (Blackwell) - real vLLM numbers
๐๐ 3
2
#9 opened 2 months ago
by
janreges3
Chat-template fix: occasional </think> leak + body duplication on long-context tasks (patched template attached)
1
#3 opened 5 months ago
by
janreges3
Thanks and request for FP8 version
3
#2 opened 5 months ago
by
janreges3
vLLM - Looping prevention
๐ 1
1
#39 opened 7 months ago
by
janreges3
How to set MRL (variable dimensions) in vLLM
โค๏ธ 1
5
#21 opened 10 months ago
by
janreges3
Request for NVFP4A16 model Qwen3-30B-A3B 2507 (Thinking & Instruct)
#1 opened about 1 year ago
by
janreges3