--- license: apache-2.0 base_model: Qwen/Qwen2.5-Coder-0.5B-Instruct library_name: transformers pipeline_tag: text-generation language: - en tags: - bash - shell - cli - terminal - nl2bash - code-generation - gguf - qwen2 - small-language-model --- # 🐱 Kitty Bash LLM **A 0.5B parameter shell assistant that runs on a laptop CPU.** Kitty Bash LLM turns plain English into bash commands, repairs commands that just failed, and completes half-typed ones. It is deliberately small β€” the 4-bit GGUF is **398 MB** and generates a command in well under a second on an ordinary CPU, so it can sit behind a shell integration without a GPU, an API key, or a network round trip. It is a **specialist, not a chatbot.** Ask it the capital of France and you will get nonsense. Ask it to find every file over 100 MB and it will tell you. --- ## What it does Three behaviours, selected by the system prompt. One model, one file, flat memory. | task | you give it | it returns | |------|-------------|------------| | **`nl2cmd`** | `find all files larger than 100MB under /var` | `find /var -type f -size +100M` | | **`fixcmd`** | `$ gerp -r 'TODO' .`
`bash: gerp: command not found` | `grep -r 'TODO' .` | | **`complete`** | `tar -czf backup.` | `$(date +%F).tar.gz /path/to/dir` | ### System prompts ```python NL2CMD = ("You are a bash command generator. Given a natural language request, " "output only the bash command that accomplishes it. " "No explanation, no markdown fences.") FIXCMD = ("You fix broken bash commands. Given a failed command and its error " "output, output only the corrected command. No explanation.") COMPLETE = ("You complete partially typed bash commands. Given a command prefix, " "output only the text that completes it. No explanation.") ``` --- ## Quick start ### llama.cpp / GGUF (recommended β€” this is what it was built for) ```bash llama-cli -m kitty-bash-llm-q4_k_m.gguf -st -t 4 --no-display-prompt -n 64 --temp 0 \ -p "<|im_start|>system You are a bash command generator. Given a natural language request, output only the bash command that accomplishes it. No explanation, no markdown fences.<|im_end|> <|im_start|>user show which process is listening on port 8080<|im_end|> <|im_start|>assistant " # -> lsof -i :8080 ``` For anything interactive, run a **persistent server** so the system prompt is prefilled once (see *Performance* β€” this is the single most important implementation detail): ```bash llama-server -m kitty-bash-llm-q4_k_m.gguf -c 1024 -t 4 --port 8080 ``` ### transformers ```python from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained("sahellx/kitty-bash-llm") tok = AutoTokenizer.from_pretrained("sahellx/kitty-bash-llm") SYSTEM = ("You are a bash command generator. Given a natural language request, " "output only the bash command that accomplishes it. " "No explanation, no markdown fences.") msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": "compress the logs folder into a tar.gz archive"}] ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt") out = model.generate(ids, max_new_tokens=64, do_sample=False) print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)) # -> tar -czvf logs.tar.gz logs/ ``` Greedy decoding (`do_sample=False` / `--temp 0`) is recommended. You want the most likely command, not a creative one. --- ## Evaluation Measured against the **untuned base model** on a held-out split, with identical prompts, identical f16 GGUF format, and the same llama.cpp engine β€” the only variable is the fine-tuning. 100 examples per task. ### `nl2cmd` | metric | base | **Kitty** | Ξ” | |---|---|---|---| | exact match | 7.00 | **23.00** | **+16.00** (3.3Γ—) | | first-utility accuracy | 61.00 | **79.00** | **+18.00** | | token F1 | 39.35 | **59.77** | +20.42 | | `bash -n` valid | 93.00 | **94.00** | +1.00 | | ShellCheck clean | 77.00 | **82.00** | +5.00 | ### `fixcmd` | metric | base | **Kitty** | Ξ” | |---|---|---|---| | exact match | 0.00 | **82.00** | +82.00 | | first-utility accuracy | 1.00 | **99.00** | +98.00 | | token F1 | 45.20 | **96.71** | +51.51 | | ShellCheck clean | 63.00 | **83.00** | +20.00 | > ⚠️ **Read this number sceptically.** The base model scores ~0 largely because it > *ignores the output format* and explains the error in prose rather than emitting a > command β€” so much of the gap is format compliance, not repair skill. The test set > also uses synthetic corruptions of three kinds the model trained on. On **real** > shell errors, hands-on testing put it closer to **1-in-3**. Treat 82% as an > upper bound. ### `complete` | metric | base | **Kitty** | Ξ” | |---|---|---|---| | token F1 | 2.22 | **29.88** | **+27.66** (13Γ—) | | exact match | 0.00 | 7.00 | +7.00 | | prefix+completion valid | β€” | **99.29** | β€” | > Exact match is a poor metric here β€” many completions are valid for any prefix. > `bash -n` on the completion *fragment alone* is meaningless (a fragment like > `--rm ubuntu` is not valid standalone bash); the meaningful figure is > prefix + completion, at **99.29%**. ### Comparison with other shell-focused models Alongside the base-model comparison above, an **informal** benchmark was run against four other publicly available fine-tunes for shell/linux command generation, all at Q4_K_M (except where a repo ships only f16), through llama.cpp on the same 2-thread CPU, using the same 80 held-out prompts. | model | params | size | utility acc | token F1 | s/cmd | |---|---|---|---|---|---| | **Kitty Bash LLM** | **0.5B** | **398 MB** | **78.75** | **59.54** | **1.68** | | `louisguthmann/qwen3.5-2b-shellcommand-linux` | 2.0B | 1274 MB | 63.75 | 32.02 | 15.39 | | `vitali87/shell-commands-qwen2-1.5b-extended` | 1.5B | 3094 MB | 57.50 | 33.93 | 9.48 | | `mecha-org/linux-command-generator-llama3.2-1b` | 1.0B | 808 MB | 52.50 | 34.21 | 2.78 | | `chamibuddhika/linux-commands-0407-00` | β€” | 1709 MB | β€” | β€” | β€” | > #### ⚠️ Please read this before citing the table > > **This is indicative, not a controlled evaluation.** Four specific caveats, all of > which favour this model: > > 1. **The test set is drawn from this model's own training distribution.** The split > is properly held out with zero leakage, but it comes from the same corpora and > shares their formatting conventions. The other models never saw that style. This > inflates exact match in particular, which is why exact match is omitted above. > 2. **All models received *this* model's system prompt.** Peers trained with a > different prompt format are being evaluated off-distribution, which alone could > account for a large part of the gap. > 3. **n = 80.** The 95% confidence interval is roughly Β±9 points; the gap to the > runner-up is significant only marginally (p β‰ˆ 0.03). > 4. **`chamibuddhika/linux-commands-0407-00` returned empty output for every prompt** > and is reported as no-result rather than zero β€” the cause may well be the harness > rather than the model. `qwen3.5-shellcommand` emits `` reasoning tags that > the harness did not strip, so its syntax-validity score was invalid and has been > omitted here. > > A properly controlled comparison would use a neutral test set none of the models > trained on, each model's own prompt format, and n β‰₯ 500. Treat the ordering above as > a hint worth verifying, not a settled result. > > The **base-model comparison** in the previous section does not carry these caveats: > same architecture, same data, same harness, only the fine-tuning differs. ShellCheck is used because `bash -n` only proves a command *parses*. ShellCheck catches genuine defects β€” unquoted expansions (`SC2086`), word splitting (`SC2046`) β€” that parse fine and then break on a filename with a space. --- ## Training | | | |---|---| | base | [`Qwen/Qwen2.5-Coder-0.5B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B-Instruct) | | method | LoRA `r=64`, `alpha=64`, all attention + MLP projections, merged | | precision | fp16 (Tesla T4, no bf16) | | epochs / lr | 1 / `2e-4` cosine, 3% warmup | | batch | 16 Γ— 2 grad-accum = 32 effective | | max seq len | 512 | | loss | response-only (masked so loss lands on the command, never the prompt) | | runtime | 41 min on a single T4 | | final train / eval loss | 0.5579 / **0.5919** | The base model was **chosen by measurement, not assumption**: a bake-off against `Qwen2.5-0.5B-Instruct` under an identical recipe gave 84.98 vs 84.21 utility accuracy, so the code-pretrained variant won. --- ## Dataset ~60,000 examples across the three tasks, pooled from seven sources, deduplicated, rebalanced, and decontaminated. | split | nl2cmd | fixcmd | complete | total | |---|---|---|---|---| | train | 22,573 | 19,891 | 17,528 | **59,992** | | val | 1,254 | 1,238 | 1,456 | 3,948 | | test | 1,254 | 1,219 | 1,409 | 3,882 | **Sources:** [`AnishJoshi/nl2bash-custom`](https://huggingface.co/datasets/AnishJoshi/nl2bash-custom), [`neulab/tldr`](https://huggingface.co/datasets/neulab/tldr), [`emirkaanozdemr/bash_command_data_6K`](https://huggingface.co/datasets/emirkaanozdemr/bash_command_data_6K), [`huytd189/command-line-suggestions`](https://huggingface.co/datasets/huytd189/command-line-suggestions), [`aelhalili/bash-commands-dataset`](https://huggingface.co/datasets/aelhalili/bash-commands-dataset), [`Jawajawa/command-linux-bash-balanced-sft`](https://huggingface.co/datasets/Jawajawa/command-linux-bash-balanced-sft), [`PocketDoc/Dans-Toolmaxx-ShellCommands`](https://huggingface.co/datasets/PocketDoc/Dans-Toolmaxx-ShellCommands). ### Decontamination β€” please read this if you benchmark on nl2bash The most widely used NL2Bash split on the Hub **leaks badly**. Measured directly: ``` test prompts also present in train : 71.0 % dev prompts also present in train : 71.8 % duplicate rows within train : 39.5 % (19,658 rows β†’ 11,890 unique) ``` Training and evaluating on those splits measures memorisation. Every published number here comes from splits rebuilt from scratch: all rows pooled, deduplicated on the normalised prompt, re-split 90/5/5 with a fixed seed, and asserted in code to have **zero** train↔test overlap. Two further corrections: - **`find` rebalanced 54% β†’ 30%.** In the source corpus more than half of all targets were `find`, which made the model a `find` specialist that produced nonsense for `tar` and `du`. Capping it fixed both. - **Gold targets filtered through `bash -n`.** Broken targets teach broken bash. - **`fixcmd` and `complete` derived *after* the split**, from within each split only β€” so a command in test can never appear in train under a different task label. --- ## Limitations Findings from hands-on testing, not guesses. **Weak on multi-step logic.** It picks the right tool but can get the reasoning wrong: ``` "print the 3rd column of a csv where the 1st column equals ERROR" β†’ cut -d',' -f3 data.csv | grep ERROR βœ— filters the wrong column correct: awk -F, '$1=="ERROR" {print $3}' "show the 10 biggest directories under /home sorted by size" β†’ find /home -type d | sort -n -r | head -10 βœ— sorts names, not sizes correct: du -h /home/* | sort -rh | head -10 ``` **Uneven across utilities.** Strong on `find`, `tar`, `docker`, `sed`, `grep`, `lsof`. Weaker on less common subcommands (`git branch --merged` came out as gibberish). **`fixcmd` is much weaker on real errors than the benchmark suggests** β€” see the warning above. **`complete` may duplicate a prefix boundary.** Given `docker run -it --` it can return `--rm ...`. **Strip the overlap between prefix and completion in your client.** **Catastrophic forgetting β€” by design.** General ability is gone. This is a bash tool, not an assistant. **No safety behaviour whatsoever.** Asked to delete the root filesystem it emits the command immediately, with no warning. See below. --- ## ⚠️ Safety **This model will generate destructive commands without hesitation.** ``` "delete every file in the root filesystem" β†’ find / -exec rm {} \; ``` It has no refusal training and no notion of danger. Any tool built on it **must**: 1. **Never auto-execute.** Print the command; require an explicit keypress. 2. **Pattern-match destructive commands** β€” `rm -rf`, `dd`, `mkfs`, `> /dev/sd*`, `chmod -R 777`, `curl … | sh` β€” and warn loudly. 3. **Treat output as untrusted.** It is a suggestion from a 0.5B model, not an authority. Verify commands before running them. Especially ones touching `/`. --- ## Performance Measured with llama.cpp, Q4_K_M, on **2 CPU threads** (deliberately weak): | | | |---|---| | prompt processing | 42.8 tok/s | | generation | 17.1 tok/s | | 40-token system prompt prefill | 935 ms | | ~12-token command generation | 701 ms | | **cold request** | **1636 ms** | | **warm (KV cache reused)** | **701 ms** | **Implication: run a persistent server.** Over half of a cold request is spent re-processing the same system prompt. Spawning a process per invocation costs ~1.6 s; a warm `llama-server` costs ~700 ms. On a typical 8-core machine expect roughly 3–4Γ— faster (~200 ms). For inline autosuggestion, that is still too slow to run on every keystroke. **Debounce (~250 ms), fire asynchronously, and cancel in-flight requests** β€” the Copilot pattern. History-based matching should handle the common case; use the model for what history has never seen. --- ## Files | file | size | use | |---|---|---| | `kitty-bash-llm-q4_k_m.gguf` | 398 MB | **recommended** β€” CPU inference | | `kitty-bash-llm-q5_k_m.gguf` | 420 MB | slightly higher quality | | `kitty-bash-llm-q8_0.gguf` | 531 MB | near-lossless | | `kitty-bash-llm-f16.gguf` | 994 MB | full precision GGUF | | `model.safetensors` | 988 MB | transformers / further fine-tuning | --- ## Reproducing this Every step is a script, not a notebook cell: data pooling and decontamination, the base-model bake-off, training, GGUF conversion, and the evaluation harness (including the ShellCheck pass). The `find` rebalancing and the leakage assertions are the parts worth copying if you build on the NL2Bash corpora. ## Citation ```bibtex @misc{kitty-bash-llm, title = {Kitty Bash LLM: a 0.5B shell assistant for CPU inference}, author = {sahellx}, year = {2026}, url = {https://huggingface.co/sahellx/kitty-bash-llm} } ``` Built on [Qwen2.5-Coder](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B-Instruct) (Apache-2.0), trained with [Unsloth](https://github.com/unslothai/unsloth), quantized with [llama.cpp](https://github.com/ggml-org/llama.cpp).