--- license: apache-2.0 license_link: https://ai.google.dev/gemma/docs/gemma_4_license base_model: unsloth/gemma-4-12b-it base_model_relation: finetune pipeline_tag: text-generation language: - en tags: - gguf - llama.cpp - gemma - python - code - reasoning - conversational --- # Instinct-Python-Coder-Gemma4-12B-KimiK3 Instinct-Python-Coder-Gemma4-12B-KimiK3 is a general-purpose Python coder that thinks briefly, then answers. We took the highly capable Gemma 4 12B model and taught it to perform better in Python coding with _much_ more concise reasoning in its `` channel, leading to faster inference, a shorter context window, and cost savings for the end user. The reasoning style is distilled from Kimi K3: the model works through the approach in a few lines, then hands back the code. ## Evaluation We measured first-attempt accuracy on a held-out set of 228 Python tasks: one greedy completion per task, each graded automatically. | | Gemma 4 12B (base) | Instinct | | --------------------------------------- | ------------------ | --------------- | | Solved | 40 / 228 (17.5%) | 81 / 228 (35.5%) | | Ran out of budget without writing code | 90 | 28 | | Time to run the full set (batched) | 233.6 min | 126.8 min | With a 35.5% first-attempt accuracy, our Python Coder Instinct model preserved 62% of Kimi K3's thinking capability: Kimi K3 itself reaches about 57.3% first-attempt accuracy on a broader Python pool under the same kind of check. For comparison, the base Gemma 4 12B model has only 17.5% accuracy, so we more than doubled it. On 90 of the 228 tasks the base model ran out of its budget without ever writing code; the finetune cut that to 28. These gains trace to the same change: the model reaches the answer instead of thinking until it runs out of room. ## Training Fine-tuned on 1.87M post-training tokens, passed over twice for 3.73M tokens in total, at a sequence length of 8,192. ## Limitations This is one 12B model measured once with greedy decoding, so treat 35.5% as a single reading with no error bar. It was tuned and tested on Python, and nothing else was measured here. It inherits Gemma 4's behavior and limitations. The 17.5% base-model figure above was measured with a plain "explain your approach, then write code" prompt — not the `` instruction this fine-tune was trained on. Scoring the same base weights with a prompt that explicitly asks for a `` block (the format this model actually uses) gives 3.5%, not 17.5%: the untuned model mostly rambles trying to follow an unfamiliar instruction instead of writing code. So part of the 17.5% → 35.5% gap is this model learning to follow that instruction format, not purely a coding-skill improvement. The absolute 35.5% is unaffected either way. ## License and lineage Base model: [unsloth/gemma-4-12b-it](https://huggingface.co/unsloth/gemma-4-12b-it), which ships under the **Apache 2.0** license (see its own model card's frontmatter — `license: apache-2.0`), not the standard Gemma Terms of Use. This fine-tune inherits that Apache 2.0 license. Training data is distilled from Kimi K3 (Moonshot AI); Moonshot gave written permission for this project to train on and redistribute Kimi K3-derived outputs and models (recorded in this repo's `docs/kimi-k3-distillation.md`). ## Usage It ships as a single Q8_0 GGUF, roughly 13 GB on disk, and runs on a GPU or Mac with about 16 GB of memory. Serve it with llama.cpp: ```bash llama-server -m Instinct-Python-Coder-Gemma4-12B-KimiK3-Q8_0.gguf -ngl 99 -c 8192 ``` That exposes an OpenAI-compatible endpoint at `http://localhost:8080/v1`, and the Gemma 4 chat template is baked into the GGUF, so turns and the thinking channel are formatted for you. It runs anywhere GGUF runs, and any tool that speaks the OpenAI chat API can drive it: - **Runtimes and apps**: llama.cpp, Ollama, LM Studio, Jan, KoboldCpp - **Coding agents and harnesses**: opencode, pi, Hermes, Aider, Cline, Continue Describe what you want in plain language and it replies with a short pass of reasoning followed by the code: ```` User: Return the first character in a string that appears only once, or None. Count characters in one pass, then scan again and return the first with count 1. O(n). ```python from collections import Counter def first_unique(s): counts = Counter(s) for ch in s: if counts[ch] == 1: return ch return None ``` ````