Spaces:
Sleeping
Sleeping
File size: 3,159 Bytes
f07454a 937171c a5a431f 937171c a5a431f 937171c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 | ---
title: ByteBot
emoji: π€
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: "5.34.2"
python_version: "3.11"
app_file: app.py
pinned: false
---
# π€ Qwen Alpaca GGUF
A fine tuned **Qwen2.5-0.5B-Instruct** language model trained on the **tatsu-lab/alpaca** instruction dataset. The model was fine tuned using **Unsloth**, quantized to **GGUF (Q4_K_M)**, and can be run locally with **Ollama** or **llama.cpp**.
---
## π Model Details
| Property | Value |
|----------|-------|
| Base Model | Qwen2.5-0.5B-Instruct |
| Fine Tuning | LoRA |
| Framework | Unsloth |
| Dataset | tatsu-lab/alpaca |
| Format | GGUF |
| Quantization | Q4_K_M |
| Inference | Ollama, llama.cpp |
| Language | English |
---
## β¨ Features
- Instruction following chatbot
- Lightweight 0.5B parameter model
- GGUF format for efficient CPU inference
- Optimized using Q4_K_M quantization
- Compatible with Ollama
- Compatible with llama.cpp
- Easy local deployment
---
## π¦ Model File
```
Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
```
---
# π Fine Tuning Pipeline
```
Base Model
β
βΌ
Qwen2.5-0.5B-Instruct
β
βΌ
Alpaca Instruction Dataset
β
βΌ
LoRA Fine Tuning
β
βΌ
Unsloth
β
βΌ
Merge LoRA Adapters
β
βΌ
GGUF Conversion
β
βΌ
Q4_K_M Quantization
β
βΌ
Inference using Ollama / llama.cpp
```
---
# π Dataset
This model was fine tuned using the **tatsu-lab/alpaca** instruction dataset.
The dataset contains thousands of instruction and response pairs designed to improve instruction following ability.
---
# β‘ Run with Ollama
Create a Modelfile
```text
FROM ./Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
```
Create the model
```bash
ollama create fineqwen -f Modelfile
```
Run the model
```bash
ollama run fineqwen
```
---
# π₯ Example Terminal
```text
$ ollama create fineqwen -f Modelfile
transferring model...
creating new layer...
writing manifest...
success
$ ollama run fineqwen
>>> Explain Transformers in simple words.
Transformers are neural networks designed to process sequences using
self-attention, allowing them to understand relationships between words
efficiently.
```
---
# π¦ Run using llama.cpp
```bash
llama-cli \
-hf ciphermosaic/qwen-alpaca-gguf \
--jinja
```
or
```bash
llama-cli \
-m Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
```
---
# π Hugging Face Space
Interactive chatbot available on Hugging Face Spaces.
---
# π Quantization
This model uses
```
Q4_K_M
```
Benefits
- Smaller model size
- Faster inference
- Lower RAM usage
- Minimal quality loss
---
# π§° Tech Stack
- Python
- Hugging Face Transformers
- Unsloth
- PEFT (LoRA)
- GGUF
- llama.cpp
- Ollama
- Hugging Face Hub
- Gradio
---
# π Repository Structure
```
.
βββ Modelfile
βββ Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
βββ README.md
βββ config.json
```
---
# π¨βπ» Author
**CipherMosaic**
GitHub: https://github.com/CipherMosaic
Hugging Face: https://huggingface.co/ciphermosaic
---
## β If you found this project useful, consider giving it a star! |