Spaces:
Sleeping
Sleeping
A newer version of the Gradio SDK is available: 6.20.0
metadata
title: ByteBot
emoji: π€
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 5.34.2
python_version: '3.11'
app_file: app.py
pinned: false
π€ Qwen Alpaca GGUF
A fine tuned Qwen2.5-0.5B-Instruct language model trained on the tatsu-lab/alpaca instruction dataset. The model was fine tuned using Unsloth, quantized to GGUF (Q4_K_M), and can be run locally with Ollama or llama.cpp.
π Model Details
| Property | Value |
|---|---|
| Base Model | Qwen2.5-0.5B-Instruct |
| Fine Tuning | LoRA |
| Framework | Unsloth |
| Dataset | tatsu-lab/alpaca |
| Format | GGUF |
| Quantization | Q4_K_M |
| Inference | Ollama, llama.cpp |
| Language | English |
β¨ Features
- Instruction following chatbot
- Lightweight 0.5B parameter model
- GGUF format for efficient CPU inference
- Optimized using Q4_K_M quantization
- Compatible with Ollama
- Compatible with llama.cpp
- Easy local deployment
π¦ Model File
Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
π Fine Tuning Pipeline
Base Model
β
βΌ
Qwen2.5-0.5B-Instruct
β
βΌ
Alpaca Instruction Dataset
β
βΌ
LoRA Fine Tuning
β
βΌ
Unsloth
β
βΌ
Merge LoRA Adapters
β
βΌ
GGUF Conversion
β
βΌ
Q4_K_M Quantization
β
βΌ
Inference using Ollama / llama.cpp
π Dataset
This model was fine tuned using the tatsu-lab/alpaca instruction dataset.
The dataset contains thousands of instruction and response pairs designed to improve instruction following ability.
β‘ Run with Ollama
Create a Modelfile
FROM ./Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
Create the model
ollama create fineqwen -f Modelfile
Run the model
ollama run fineqwen
π₯ Example Terminal
$ ollama create fineqwen -f Modelfile
transferring model...
creating new layer...
writing manifest...
success
$ ollama run fineqwen
>>> Explain Transformers in simple words.
Transformers are neural networks designed to process sequences using
self-attention, allowing them to understand relationships between words
efficiently.
π¦ Run using llama.cpp
llama-cli \
-hf ciphermosaic/qwen-alpaca-gguf \
--jinja
or
llama-cli \
-m Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
π Hugging Face Space
Interactive chatbot available on Hugging Face Spaces.
π Quantization
This model uses
Q4_K_M
Benefits
- Smaller model size
- Faster inference
- Lower RAM usage
- Minimal quality loss
π§° Tech Stack
- Python
- Hugging Face Transformers
- Unsloth
- PEFT (LoRA)
- GGUF
- llama.cpp
- Ollama
- Hugging Face Hub
- Gradio
π Repository Structure
.
βββ Modelfile
βββ Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
βββ README.md
βββ config.json
π¨βπ» Author
CipherMosaic
GitHub: https://github.com/CipherMosaic
Hugging Face: https://huggingface.co/ciphermosaic