Instructions to use CodeDevX/qwen2.5-1.5b-instruct-quantized with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="CodeDevX/qwen2.5-1.5b-instruct-quantized") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("CodeDevX/qwen2.5-1.5b-instruct-quantized", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CodeDevX/qwen2.5-1.5b-instruct-quantized" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeDevX/qwen2.5-1.5b-instruct-quantized", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/CodeDevX/qwen2.5-1.5b-instruct-quantized
- SGLang
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "CodeDevX/qwen2.5-1.5b-instruct-quantized" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeDevX/qwen2.5-1.5b-instruct-quantized", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "CodeDevX/qwen2.5-1.5b-instruct-quantized" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CodeDevX/qwen2.5-1.5b-instruct-quantized", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use CodeDevX/qwen2.5-1.5b-instruct-quantized with Docker Model Runner:
docker model run hf.co/CodeDevX/qwen2.5-1.5b-instruct-quantized
|
Download README.md from CodeDevX/qwen2.5-1.5b-instruct-quantized: direct link, hf CLI and curl.
- Browser
- Download file 5.31 kB
-
https://huggingface.co/CodeDevX/qwen2.5-1.5b-instruct-quantized/resolve/main/README.md
- Command line
-
hf download hf://CodeDevX/qwen2.5-1.5b-instruct-quantized/README.md
-
curl -L -o README.md https://huggingface.co/CodeDevX/qwen2.5-1.5b-instruct-quantized/resolve/main/README.md
5.31 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen2.5-1.5B-Instruct | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| tags: | |
| - qwen2 | |
| - quantized | |
| - text-generation | |
| - chat | |
| # Qwen2.5-1.5B-Instruct-Quantized | |
| A quantized checkpoint created from **[Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)**. | |
| This repository contains a quantized version of the original instruction-tuned Qwen2.5 1.5B model. The base model was developed by the Qwen team. This repository is a community quantization, not the original Qwen release. The base model is a causal language model designed for instruction following and conversational text generation. See the [official base model card](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) for its architecture and original documentation. | |
| ## Model Details | |
| | Field | Details | | |
| |---|---| | |
| | Model name | Qwen2.5-1.5B-Instruct-Quantized | | |
| | Base model | [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) | | |
| | Model family | Qwen2.5 | | |
| | Model size | 1.5B-class model (base model: approximately 1.54B parameters) | | |
| | Task | Text generation, chat, instruction following | | |
| | Quantization | Quantized by the repository maintainer; method and bit-width not specified | | |
| | Maintainer | CodeDevX | | |
| ## Intended Use | |
| This model may be used for: | |
| - Conversational assistance and instruction following | |
| - General text generation and question answering | |
| - Summarization, rewriting, and drafting | |
| - Experimentation with quantized language models and local inference, subject to runtime compatibility | |
| ## Quick Start | |
| The exact loading method depends on the quantization format used for this checkpoint. Check the repository's **Files and versions** tab for the model file extension and configuration before choosing an inference runtime. | |
| ### Transformers (for Transformers-compatible checkpoints) | |
| If the uploaded files are compatible with Transformers, you can try: | |
| ```bash | |
| pip install -U transformers torch accelerate safetensors | |
| ``` | |
| ```python | |
| import torch | |
| from transformers import AutoTokenizer, AutoModelForCausalLM | |
| model_id = "CodeDevX/qwen2.5-1.5b-instruct-quantized" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| torch_dtype="auto", | |
| device_map="auto", | |
| ) | |
| messages = [ | |
| {"role": "user", "content": "Explain quantization in simple terms."} | |
| ] | |
| prompt = tokenizer.apply_chat_template( | |
| messages, | |
| tokenize=False, | |
| add_generation_prompt=True, | |
| ) | |
| inputs = tokenizer(prompt, return_tensors="pt").to(model.device) | |
| with torch.no_grad(): | |
| output = model.generate( | |
| **inputs, | |
| max_new_tokens=256, | |
| do_sample=True, | |
| temperature=0.7, | |
| top_p=0.8, | |
| ) | |
| answer = tokenizer.decode( | |
| output[0][inputs["input_ids"].shape[1]:], | |
| skip_special_tokens=True, | |
| ) | |
| print(answer) | |
| ``` | |
| **Important:** This example is for Transformers-compatible model files. It will not load a GGUF file directly. For GGUF, use a compatible runtime such as `llama.cpp` and follow the runtime's model-loading instructions. | |
| ## Chat Template | |
| The original Qwen2.5-Instruct model uses a chat template. If the tokenizer is included and compatible, use `tokenizer.apply_chat_template()` to format messages. | |
| ```python | |
| messages = [ | |
| {"role": "system", "content": "You are a helpful assistant."}, | |
| {"role": "user", "content": "Summarize the benefits of renewable energy."}, | |
| ] | |
| ``` | |
| ## Quantization Information | |
| The checkpoint has been quantized from the base model listed above. The following technical details can be filled in to make the model card reproducible: | |
| - **Quantization method:** Not specified | |
| - **Bit-width / quantization type:** Not specified | |
| - **Quantization framework or tool:** Not specified | |
| - **Calibration dataset (if applicable):** Not specified | |
| - **Quantization settings:** Not specified | |
| - **Original and quantized file sizes:** Not specified | |
| - **Benchmark comparison with the original model:** Not provided | |
| Quantization can reduce model storage and memory requirements, but the effect on output quality, speed, and hardware compatibility depends on the quantization method and runtime. | |
| ## Limitations | |
| - Responses may be inaccurate, incomplete, or biased. | |
| - Quantization may affect output quality and performance. | |
| - Hardware and runtime requirements depend on the checkpoint format and quantization method. | |
| - Do not rely on model outputs as the sole basis for high-stakes decisions. | |
| ## Evaluation | |
| No benchmark or quality evaluation results are documented here. If available, add benchmark names, scores, hardware, inference settings, and a comparison against the original model. | |
| ## License and Attribution | |
| The official base model is listed under the Apache-2.0 license. Review the [base model license and terms](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) and ensure this derived checkpoint includes any required license and attribution notices before redistribution or commercial use. | |
| ## Credits | |
| - **Original model:** [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) | |
| - **Quantized checkpoint:** [CodeDevX/qwen2.5-1.5b-instruct-quantized](https://huggingface.co/CodeDevX/qwen2.5-1.5b-instruct-quantized) | |
| --- | |
| *Quantization and repository documentation by CodeDevX.* |