🌶️ Pepper 1 Preview

Lightweight instruction-following code assistant for Python, built on Qwen2.5-Coder-1.5B.

Pepper is a 1.5B parameter code assistant designed for writing functions from natural-language descriptions. It runs on consumer hardware and is one of the few publicly available 1.5B models tuned specifically for instruction-following code generation.

Part of the Minsore family · minsore.com


✨ Highlights

  • 🧠 Instruction-tuned — writes functions from plain English descriptions
  • 🐍 Python-focused — clean, idiomatic code without fluff
  • ⚡ Real-time — designed for interactive use
  • 📦 Compact — 1.5B params, ~1 GB in Q4_K_M
  • 🎯 Purpose-built — trained for code generation, not chat
  • 🆓 Apache 2.0 — same license as base model

Pepper comparison

📊 Benchmarks

Evaluated against 1–1.5B code models. All runs used temperature=0.0, max_tokens=512, --chat-template none.

Benchmark Pepper 1 Preview Base Qwen 1.5B Llama-3.2-1B-Code Yi-Coder-1.5B
HumanEval@50 82.0% 75.0% 64.0% 24.0%
MBPP@50 26.0% 20.0% 16.0% 38.0%
LiveCodeBench@30 20.0% 30.0% 10.0% 26.7%
BigCodeBench@30 23.3% 23.3% 13.3% 30.0%

📌 Key insight: Pepper 1 Preview outperforms the base model on HumanEval by +7 points — the primary benchmark for instruction-following code generation. On MBPP, both models underperform the official numbers because this evaluation uses zero-shot prompting with --chat-template none. Yi-Coder tested without its native chat template — its HumanEval score reflects format mismatch, not model quality.


🚀 Quick Start

llama.cpp

llama-server -m pepper-1-preview.Q4_K_M.gguf \
  --port 8080 \
  -ngl 99 \
  -c 4096 \
  --chat-template chatml

Requirements: any GPU with ≥2 GB VRAM (full offload), or partial CPU offload as fallback.

Chat request

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      {"role": "user", "content": "Write a Python function that checks if a number is prime."}
    ],
    "max_tokens": 512,
    "temperature": 0.0
  }'

Python

import requests

def ask(prompt, max_tokens=512):
    r = requests.post("http://localhost:8080/v1/chat/completions", json={
        "messages": [{"role": "user", "content": prompt}],
        "max_tokens": max_tokens,
        "temperature": 0.0,
    })
    return r.json()["choices"][0]["message"]["content"]

print(ask("Write a Python function that reverses a string."))
# → def reverse_string(s: str) -> str:
#       return s[::-1]

⚙️ Recommended Settings

Parameter Value Notes
--chat-template chatml Required. Pepper expects ChatML.
n_predict 256–1024 512 is a good default
temperature 0.0 Deterministic; use 0.2 for variation
repeat_penalty 1.1 Prevents repetition
-c 4096 Longer context available but degrades

⚠️ Limitations

  • Python-only — not trained on other languages
  • No FIM support — for fill-in-the-middle tasks, use Quill
  • Not an agent — does not support tool calling or multi-step planning
  • Weak on algorithmic tasks — LiveCodeBench score below base Qwen
  • 4K inference context — long files are truncated
  • Occasional over-explanation — may include comments when only code is requested

🧬 Training Details

Base model Qwen2.5-Coder-1.5B-Instruct
Method QLoRA fine-tuning on curated Python code data
Context 4096 tokens (inference)

📁 Files

File Size Description
pepper-1-preview.Q4_K_M.gguf ~1 GB Ready to use with llama.cpp
model.safetensors ~3 GB Full precision (transformers)
config.json, tokenizer.json — Config for transformers

🗺️ Roadmap

  • Pepper 2 — improve MBPP and LiveCodeBench, add multi-language support
  • Quill 2 — extend FIM to JS/TS/Rust
  • Symphony — flagship code model (3B)

📜 License

Apache 2.0 — same as the base Qwen2.5-Coder-1.5B model.


🙏 Credits


📬 Contact

Minsore — Ukrainian AI lab building open language models.


⭐ If Pepper is useful, star the repo and share your results.

Downloads last month
307
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for minsore/pepper-1-preview

Quantized
(183)
this model