ByteBot / README.md
ciphermosaic's picture
Update README.md
f07454a verified
|
Raw
History Blame Contribute Delete
3.16 kB

A newer version of the Gradio SDK is available: 6.20.0

Upgrade
metadata
title: ByteBot
emoji: πŸ€–
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 5.34.2
python_version: '3.11'
app_file: app.py
pinned: false

πŸ€– Qwen Alpaca GGUF

A fine tuned Qwen2.5-0.5B-Instruct language model trained on the tatsu-lab/alpaca instruction dataset. The model was fine tuned using Unsloth, quantized to GGUF (Q4_K_M), and can be run locally with Ollama or llama.cpp.


πŸš€ Model Details

Property Value
Base Model Qwen2.5-0.5B-Instruct
Fine Tuning LoRA
Framework Unsloth
Dataset tatsu-lab/alpaca
Format GGUF
Quantization Q4_K_M
Inference Ollama, llama.cpp
Language English

✨ Features

  • Instruction following chatbot
  • Lightweight 0.5B parameter model
  • GGUF format for efficient CPU inference
  • Optimized using Q4_K_M quantization
  • Compatible with Ollama
  • Compatible with llama.cpp
  • Easy local deployment

πŸ“¦ Model File

Qwen2.5-0.5B-Instruct.Q4_K_M.gguf

πŸ›  Fine Tuning Pipeline

Base Model
        β”‚
        β–Ό
Qwen2.5-0.5B-Instruct
        β”‚
        β–Ό
Alpaca Instruction Dataset
        β”‚
        β–Ό
LoRA Fine Tuning
        β”‚
        β–Ό
Unsloth
        β”‚
        β–Ό
Merge LoRA Adapters
        β”‚
        β–Ό
GGUF Conversion
        β”‚
        β–Ό
Q4_K_M Quantization
        β”‚
        β–Ό
Inference using Ollama / llama.cpp

πŸ“š Dataset

This model was fine tuned using the tatsu-lab/alpaca instruction dataset.

The dataset contains thousands of instruction and response pairs designed to improve instruction following ability.


⚑ Run with Ollama

Create a Modelfile

FROM ./Qwen2.5-0.5B-Instruct.Q4_K_M.gguf

Create the model

ollama create fineqwen -f Modelfile

Run the model

ollama run fineqwen

πŸ–₯ Example Terminal

$ ollama create fineqwen -f Modelfile

transferring model...
creating new layer...
writing manifest...
success

$ ollama run fineqwen

>>> Explain Transformers in simple words.

Transformers are neural networks designed to process sequences using
self-attention, allowing them to understand relationships between words
efficiently.

πŸ¦™ Run using llama.cpp

llama-cli \
-hf ciphermosaic/qwen-alpaca-gguf \
--jinja

or

llama-cli \
-m Qwen2.5-0.5B-Instruct.Q4_K_M.gguf

🌐 Hugging Face Space

Interactive chatbot available on Hugging Face Spaces.


πŸ“Š Quantization

This model uses

Q4_K_M

Benefits

  • Smaller model size
  • Faster inference
  • Lower RAM usage
  • Minimal quality loss

🧰 Tech Stack

  • Python
  • Hugging Face Transformers
  • Unsloth
  • PEFT (LoRA)
  • GGUF
  • llama.cpp
  • Ollama
  • Hugging Face Hub
  • Gradio

πŸ“ Repository Structure

.
β”œβ”€β”€ Modelfile
β”œβ”€β”€ Qwen2.5-0.5B-Instruct.Q4_K_M.gguf
β”œβ”€β”€ README.md
β”œβ”€β”€ config.json

πŸ‘¨β€πŸ’» Author

CipherMosaic

GitHub: https://github.com/CipherMosaic

Hugging Face: https://huggingface.co/ciphermosaic


⭐ If you found this project useful, consider giving it a star!