CodeDevX's picture
Update README.md
1cf4d1c verified
|
Raw History Blame Contribute Delete
5.31 kB
---
license: apache-2.0
base_model: Qwen/Qwen2.5-1.5B-Instruct
pipeline_tag: text-generation
library_name: transformers
tags:
- qwen2
- quantized
- text-generation
- chat
---
# Qwen2.5-1.5B-Instruct-Quantized
A quantized checkpoint created from **[Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)**.
This repository contains a quantized version of the original instruction-tuned Qwen2.5 1.5B model. The base model was developed by the Qwen team. This repository is a community quantization, not the original Qwen release. The base model is a causal language model designed for instruction following and conversational text generation. See the [official base model card](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) for its architecture and original documentation.
## Model Details
| Field | Details |
|---|---|
| Model name | Qwen2.5-1.5B-Instruct-Quantized |
| Base model | [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) |
| Model family | Qwen2.5 |
| Model size | 1.5B-class model (base model: approximately 1.54B parameters) |
| Task | Text generation, chat, instruction following |
| Quantization | Quantized by the repository maintainer; method and bit-width not specified |
| Maintainer | CodeDevX |
## Intended Use
This model may be used for:
- Conversational assistance and instruction following
- General text generation and question answering
- Summarization, rewriting, and drafting
- Experimentation with quantized language models and local inference, subject to runtime compatibility
## Quick Start
The exact loading method depends on the quantization format used for this checkpoint. Check the repository's **Files and versions** tab for the model file extension and configuration before choosing an inference runtime.
### Transformers (for Transformers-compatible checkpoints)
If the uploaded files are compatible with Transformers, you can try:
```bash
pip install -U transformers torch accelerate safetensors
```
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "CodeDevX/qwen2.5-1.5b-instruct-quantized"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
messages = [
{"role": "user", "content": "Explain quantization in simple terms."}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
temperature=0.7,
top_p=0.8,
)
answer = tokenizer.decode(
output[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
print(answer)
```
**Important:** This example is for Transformers-compatible model files. It will not load a GGUF file directly. For GGUF, use a compatible runtime such as `llama.cpp` and follow the runtime's model-loading instructions.
## Chat Template
The original Qwen2.5-Instruct model uses a chat template. If the tokenizer is included and compatible, use `tokenizer.apply_chat_template()` to format messages.
```python
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarize the benefits of renewable energy."},
]
```
## Quantization Information
The checkpoint has been quantized from the base model listed above. The following technical details can be filled in to make the model card reproducible:
- **Quantization method:** Not specified
- **Bit-width / quantization type:** Not specified
- **Quantization framework or tool:** Not specified
- **Calibration dataset (if applicable):** Not specified
- **Quantization settings:** Not specified
- **Original and quantized file sizes:** Not specified
- **Benchmark comparison with the original model:** Not provided
Quantization can reduce model storage and memory requirements, but the effect on output quality, speed, and hardware compatibility depends on the quantization method and runtime.
## Limitations
- Responses may be inaccurate, incomplete, or biased.
- Quantization may affect output quality and performance.
- Hardware and runtime requirements depend on the checkpoint format and quantization method.
- Do not rely on model outputs as the sole basis for high-stakes decisions.
## Evaluation
No benchmark or quality evaluation results are documented here. If available, add benchmark names, scores, hardware, inference settings, and a comparison against the original model.
## License and Attribution
The official base model is listed under the Apache-2.0 license. Review the [base model license and terms](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) and ensure this derived checkpoint includes any required license and attribution notices before redistribution or commercial use.
## Credits
- **Original model:** [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
- **Quantized checkpoint:** [CodeDevX/qwen2.5-1.5b-instruct-quantized](https://huggingface.co/CodeDevX/qwen2.5-1.5b-instruct-quantized)
---
*Quantization and repository documentation by CodeDevX.*