ByteBot / README.md
ciphermosaic's picture
Update README.md
f07454a verified
|
Raw
History Blame Contribute Delete
3.16 kB
---
title: ByteBot
emoji: πŸ€–
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: "5.34.2"
python_version: "3.11"
app_file: app.py
pinned: false
---
# πŸ€– Qwen Alpaca GGUF
A fine tuned **Qwen2.5-0.5B-Instruct** language model trained on the **tatsu-lab/alpaca** instruction dataset. The model was fine tuned using **Unsloth**, quantized to **GGUF (Q4_K_M)**, and can be run locally with **Ollama** or **llama.cpp**.
---
## πŸš€ Model Details
| Property | Value |
|----------|-------|
| Base Model | Qwen2.5-0.5B-Instruct |
| Fine Tuning | LoRA |
| Framework | Unsloth |
| Dataset | tatsu-lab/alpaca |
| Format | GGUF |
| Quantization | Q4_K_M |
| Inference | Ollama, llama.cpp |
| Language | English |
---
## ✨ Features
- Instruction following chatbot
- Lightweight 0.5B parameter model
- GGUF format for efficient CPU inference
- Optimized using Q4_K_M quantization
- Compatible with Ollama
- Compatible with llama.cpp
- Easy local deployment
---
## πŸ“¦ Model File
```
Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
```
---
# πŸ›  Fine Tuning Pipeline
```
Base Model
β”‚
β–Ό
Qwen2.5-0.5B-Instruct
β”‚
β–Ό
Alpaca Instruction Dataset
β”‚
β–Ό
LoRA Fine Tuning
β”‚
β–Ό
Unsloth
β”‚
β–Ό
Merge LoRA Adapters
β”‚
β–Ό
GGUF Conversion
β”‚
β–Ό
Q4_K_M Quantization
β”‚
β–Ό
Inference using Ollama / llama.cpp
```
---
# πŸ“š Dataset
This model was fine tuned using the **tatsu-lab/alpaca** instruction dataset.
The dataset contains thousands of instruction and response pairs designed to improve instruction following ability.
---
# ⚑ Run with Ollama
Create a Modelfile
```text
FROM ./Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
```
Create the model
```bash
ollama create fineqwen -f Modelfile
```
Run the model
```bash
ollama run fineqwen
```
---
# πŸ–₯ Example Terminal
```text
$ ollama create fineqwen -f Modelfile
transferring model...
creating new layer...
writing manifest...
success
$ ollama run fineqwen
>>> Explain Transformers in simple words.
Transformers are neural networks designed to process sequences using
self-attention, allowing them to understand relationships between words
efficiently.
```
---
# πŸ¦™ Run using llama.cpp
```bash
llama-cli \
-hf ciphermosaic/qwen-alpaca-gguf \
--jinja
```
or
```bash
llama-cli \
-m Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
```
---
# 🌐 Hugging Face Space
Interactive chatbot available on Hugging Face Spaces.
---
# πŸ“Š Quantization
This model uses
```
Q4_K_M
```
Benefits
- Smaller model size
- Faster inference
- Lower RAM usage
- Minimal quality loss
---
# 🧰 Tech Stack
- Python
- Hugging Face Transformers
- Unsloth
- PEFT (LoRA)
- GGUF
- llama.cpp
- Ollama
- Hugging Face Hub
- Gradio
---
# πŸ“ Repository Structure
```
.
β”œβ”€β”€ Modelfile
β”œβ”€β”€ Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
β”œβ”€β”€ README.md
β”œβ”€β”€ config.json
```
---
# πŸ‘¨β€πŸ’» Author
**CipherMosaic**
GitHub: https://github.com/CipherMosaic
Hugging Face: https://huggingface.co/ciphermosaic
---
## ⭐ If you found this project useful, consider giving it a star!