--- license: other license_name: bcz-proprietary license_link: https://huggingface.co/sprappcom/kat-coder-v25-dev-pqm/blob/main/NOTICE language: - en tags: - moe - code - tool-calling - gated-deltanet - pqm - prism-engine - q4_k base_model: Kwaipilot/KAT-Coder-V2.5-Dev --- # kat-coder-v25-dev-pqm Standalone **[Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev)** (clean upstream code / tool-calling model, Qwen3.6-35B-A3B MoE) exported to `.pqm` for the `prism-engine` inference server. **GGUF-free standalone boot** — this single `.pqm` container is self-contained and boots directly on `prism-engine`, with no GGUF, safetensors, or transformers checkpoint alongside it. ## License & attribution This repository distributes a **composite work** under a dual arrangement: - **Proprietary layer — © 2026 BCZ Singapore Pte Ltd. All rights reserved.** The `.pqm` container format and packaging, and the `prism-engine` inference server (Rust/CUDA) are proprietary and are **not** licensed under Apache-2.0. - **Base model weights — Apache-2.0.** The underlying weights are an **unmodified** (repacked-only) derivative of a single Apache-2.0 upstream: | Component | Source | Copyright | License | Modification | |---|---|---|---|---| | Full model (backbone + experts) | [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) | © 2026 Kwaipilot / Kuaishou | Apache-2.0 | weights repacked to `.pqm` (mixed Q4_K/Q6_K/F32); no fine-tune, no merge | `KAT-Coder-V2.5-Dev` (the open Apache-2.0 weights, itself Qwen3.6-35B-A3B-based) — **not** the separate closed API-only "KAT-Coder" product. These base weights remain licensed under **Apache-2.0**. The full Apache-2.0 license text and per-component attribution are in the [`NOTICE`](https://huggingface.co/sprappcom/kat-coder-v25-dev-pqm/blob/main/NOTICE) file in this repo. This is **not** an original, trained-from-scratch foundation model — it is a quantized repack of the Apache-2.0 upstream named above. Use of the base weights is governed by Apache-2.0; use of the proprietary `.pqm` packaging and `prism-engine` requires a separate license from BCZ Singapore Pte Ltd. ## Requirements — read before downloading - **Inference engine:** `.pqm` is a proprietary container, **not** readable by vLLM, llama.cpp, Ollama, or transformers. Runs only on `prism-engine` (proprietary Rust/CUDA server, **not included in this repo**). Contact sprappcom for engine access. - **Tokenizer sidecar:** `.pqm` does **not** embed a tokenizer. Use the `kat_v25.tok` shipped in this repo via `PRISMX_TOKENIZER`. - **Hardware:** A100 (sm_80) / Ada-Ampere (sm_86/sm_89) validated. Full residency ~20 GB VRAM, or run with CPU-RAM expert offload on smaller cards. ## Architecture - Qwen3.6-35B-A3B MoE, hidden_size=2048, 40 layers (30 Gated DeltaNet + 10 full attention) - 256 routed experts + 1 shared expert, top-8 routing - vocab_size=248320 ## Quantization | Tensor group | Type | |---|---| | token embedding, attn_qkv/attn_output/attn_gate, routed experts (ffn_*_exps) | Q4_K | | shared expert (ffn_*_shexp), output.weight | Q6_K | | norms, SSM gates | F32 | ~4.56 BPW. Container ~39.9 GB (routed experts stored raw for CPU-offload support). ## Usage ```bash export PRISMX_PQM_STANDALONE=1 export PRISMX_PQM=/path/to/kat_v25.pqm export PRISMX_TOKENIZER=/path/to/kat_v25.tok prism_server 0.0.0.0:8080 \ --n-cpu-moe 28 \ --max-prefill-chunk 8 \ --moe-cache-experts 2048 \ --max-batch 1 ``` Standalone mode always binds `0.0.0.0:8080`. OpenAI-compatible `/v1/completions` and `/v1/chat/completions`. Use `temperature >= 0.15` (greedy decoding can collapse). ## Verification Boot-verified GGUF-free on A100 (sm_80) with `prism_server` (`PRISMX_PQM_STANDALONE=1 --n-cpu-moe 28 --max-prefill-chunk 8 --moe-cache-experts 2048 --max-batch 1`, tokenizer `kat_v25.tok` from this repo): coherent English code and tool-calling completions, no NaN/garbage output. ## Known limitations - **Text-only.** This `.pqm` carries no vision weights; image input is available on the separate AWQ variant, not this file. - Intermittent sub-word token drop/duplication on some code/arithmetic prompts is a known open issue; validate structured output downstream. ## Integrity | File | Bytes | SHA-256 | |---|---|---| | `kat_v25.pqm` | 39870623744 | `0050a596155bee84d162148a3e9248099f2dc838c41dca1391eb5a33ef693b15` | | `kat_v25.tok` | 8949400 | md5 `907b88ec3a1f5aed0e234472eb441520` | Check: `sha256sum kat_v25.pqm` must print the hash above.