Instructions to use CodeDevX/qwen2.5-1.5b-instruct-quantized with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="CodeDevX/qwen2.5-1.5b-instruct-quantized") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("CodeDevX/qwen2.5-1.5b-instruct-quantized", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CodeDevX/qwen2.5-1.5b-instruct-quantized" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeDevX/qwen2.5-1.5b-instruct-quantized", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/CodeDevX/qwen2.5-1.5b-instruct-quantized
- SGLang
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "CodeDevX/qwen2.5-1.5b-instruct-quantized" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeDevX/qwen2.5-1.5b-instruct-quantized", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "CodeDevX/qwen2.5-1.5b-instruct-quantized" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeDevX/qwen2.5-1.5b-instruct-quantized", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with Docker Model Runner:
docker model run hf.co/CodeDevX/qwen2.5-1.5b-instruct-quantized
Qwen2.5-1.5B-Instruct-Quantized
A quantized checkpoint created from Qwen/Qwen2.5-1.5B-Instruct.
This repository contains a quantized version of the original instruction-tuned Qwen2.5 1.5B model. The base model was developed by the Qwen team. This repository is a community quantization, not the original Qwen release. The base model is a causal language model designed for instruction following and conversational text generation. See the official base model card for its architecture and original documentation.
Model Details
| Field | Details |
|---|---|
| Model name | Qwen2.5-1.5B-Instruct-Quantized |
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Model family | Qwen2.5 |
| Model size | 1.5B-class model (base model: approximately 1.54B parameters) |
| Task | Text generation, chat, instruction following |
| Quantization | Quantized by the repository maintainer; method and bit-width not specified |
| Maintainer | CodeDevX |
Intended Use
This model may be used for:
- Conversational assistance and instruction following
- General text generation and question answering
- Summarization, rewriting, and drafting
- Experimentation with quantized language models and local inference, subject to runtime compatibility
Quick Start
The exact loading method depends on the quantization format used for this checkpoint. Check the repository's Files and versions tab for the model file extension and configuration before choosing an inference runtime.
Transformers (for Transformers-compatible checkpoints)
If the uploaded files are compatible with Transformers, you can try:
pip install -U transformers torch accelerate safetensors
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "CodeDevX/qwen2.5-1.5b-instruct-quantized"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
messages = [
{"role": "user", "content": "Explain quantization in simple terms."}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
temperature=0.7,
top_p=0.8,
)
answer = tokenizer.decode(
output[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
print(answer)
Important: This example is for Transformers-compatible model files. It will not load a GGUF file directly. For GGUF, use a compatible runtime such as llama.cpp and follow the runtime's model-loading instructions.
Chat Template
The original Qwen2.5-Instruct model uses a chat template. If the tokenizer is included and compatible, use tokenizer.apply_chat_template() to format messages.
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarize the benefits of renewable energy."},
]
Quantization Information
The checkpoint has been quantized from the base model listed above. The following technical details can be filled in to make the model card reproducible:
- Quantization method: Not specified
- Bit-width / quantization type: Not specified
- Quantization framework or tool: Not specified
- Calibration dataset (if applicable): Not specified
- Quantization settings: Not specified
- Original and quantized file sizes: Not specified
- Benchmark comparison with the original model: Not provided
Quantization can reduce model storage and memory requirements, but the effect on output quality, speed, and hardware compatibility depends on the quantization method and runtime.
Limitations
- Responses may be inaccurate, incomplete, or biased.
- Quantization may affect output quality and performance.
- Hardware and runtime requirements depend on the checkpoint format and quantization method.
- Do not rely on model outputs as the sole basis for high-stakes decisions.
Evaluation
No benchmark or quality evaluation results are documented here. If available, add benchmark names, scores, hardware, inference settings, and a comparison against the original model.
License and Attribution
The official base model is listed under the Apache-2.0 license. Review the base model license and terms and ensure this derived checkpoint includes any required license and attribution notices before redistribution or commercial use.
Credits
- Original model: Qwen/Qwen2.5-1.5B-Instruct
- Quantized checkpoint: CodeDevX/qwen2.5-1.5b-instruct-quantized
Quantization and repository documentation by CodeDevX.
docker model run hf.co/CodeDevX/qwen2.5-1.5b-instruct-quantized