--- title: ByteBot emoji: 🤖 colorFrom: blue colorTo: indigo sdk: gradio sdk_version: "5.34.2" python_version: "3.11" app_file: app.py pinned: false --- # 🤖 Qwen Alpaca GGUF A fine tuned **Qwen2.5-0.5B-Instruct** language model trained on the **tatsu-lab/alpaca** instruction dataset. The model was fine tuned using **Unsloth**, quantized to **GGUF (Q4_K_M)**, and can be run locally with **Ollama** or **llama.cpp**. --- ## 🚀 Model Details | Property | Value | |----------|-------| | Base Model | Qwen2.5-0.5B-Instruct | | Fine Tuning | LoRA | | Framework | Unsloth | | Dataset | tatsu-lab/alpaca | | Format | GGUF | | Quantization | Q4_K_M | | Inference | Ollama, llama.cpp | | Language | English | --- ## ✨ Features - Instruction following chatbot - Lightweight 0.5B parameter model - GGUF format for efficient CPU inference - Optimized using Q4_K_M quantization - Compatible with Ollama - Compatible with llama.cpp - Easy local deployment --- ## 📦 Model File ``` Qwen2.5-0.5B-Instruct.Q4_K_M.gguf ``` --- # 🛠 Fine Tuning Pipeline ``` Base Model │ ▼ Qwen2.5-0.5B-Instruct │ ▼ Alpaca Instruction Dataset │ ▼ LoRA Fine Tuning │ ▼ Unsloth │ ▼ Merge LoRA Adapters │ ▼ GGUF Conversion │ ▼ Q4_K_M Quantization │ ▼ Inference using Ollama / llama.cpp ``` --- # 📚 Dataset This model was fine tuned using the **tatsu-lab/alpaca** instruction dataset. The dataset contains thousands of instruction and response pairs designed to improve instruction following ability. --- # ⚡ Run with Ollama Create a Modelfile ```text FROM ./Qwen2.5-0.5B-Instruct.Q4_K_M.gguf ``` Create the model ```bash ollama create fineqwen -f Modelfile ``` Run the model ```bash ollama run fineqwen ``` --- # 🖥 Example Terminal ```text $ ollama create fineqwen -f Modelfile transferring model... creating new layer... writing manifest... success $ ollama run fineqwen >>> Explain Transformers in simple words. Transformers are neural networks designed to process sequences using self-attention, allowing them to understand relationships between words efficiently. ``` --- # 🦙 Run using llama.cpp ```bash llama-cli \ -hf ciphermosaic/qwen-alpaca-gguf \ --jinja ``` or ```bash llama-cli \ -m Qwen2.5-0.5B-Instruct.Q4_K_M.gguf ``` --- # 🌐 Hugging Face Space Interactive chatbot available on Hugging Face Spaces. --- # 📊 Quantization This model uses ``` Q4_K_M ``` Benefits - Smaller model size - Faster inference - Lower RAM usage - Minimal quality loss --- # 🧰 Tech Stack - Python - Hugging Face Transformers - Unsloth - PEFT (LoRA) - GGUF - llama.cpp - Ollama - Hugging Face Hub - Gradio --- # 📁 Repository Structure ``` . ├── Modelfile ├── Qwen2.5-0.5B-Instruct.Q4_K_M.gguf ├── README.md ├── config.json ``` --- # 👨‍💻 Author **CipherMosaic** GitHub: https://github.com/CipherMosaic Hugging Face: https://huggingface.co/ciphermosaic --- ## ⭐ If you found this project useful, consider giving it a star!