--- license: apache-2.0 base_model: Qwen/Qwen2.5-1.5B-Instruct pipeline_tag: text-generation library_name: transformers tags: - qwen2 - quantized - text-generation - chat --- # Qwen2.5-1.5B-Instruct-Quantized A quantized checkpoint created from **[Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)**. This repository contains a quantized version of the original instruction-tuned Qwen2.5 1.5B model. The base model was developed by the Qwen team. This repository is a community quantization, not the original Qwen release. The base model is a causal language model designed for instruction following and conversational text generation. See the [official base model card](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) for its architecture and original documentation. ## Model Details | Field | Details | |---|---| | Model name | Qwen2.5-1.5B-Instruct-Quantized | | Base model | [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) | | Model family | Qwen2.5 | | Model size | 1.5B-class model (base model: approximately 1.54B parameters) | | Task | Text generation, chat, instruction following | | Quantization | Quantized by the repository maintainer; method and bit-width not specified | | Maintainer | CodeDevX | ## Intended Use This model may be used for: - Conversational assistance and instruction following - General text generation and question answering - Summarization, rewriting, and drafting - Experimentation with quantized language models and local inference, subject to runtime compatibility ## Quick Start The exact loading method depends on the quantization format used for this checkpoint. Check the repository's **Files and versions** tab for the model file extension and configuration before choosing an inference runtime. ### Transformers (for Transformers-compatible checkpoints) If the uploaded files are compatible with Transformers, you can try: ```bash pip install -U transformers torch accelerate safetensors ``` ```python import torch from transformers import AutoTokenizer, AutoModelForCausalLM model_id = "CodeDevX/qwen2.5-1.5b-instruct-quantized" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype="auto", device_map="auto", ) messages = [ {"role": "user", "content": "Explain quantization in simple terms."} ] prompt = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, ) inputs = tokenizer(prompt, return_tensors="pt").to(model.device) with torch.no_grad(): output = model.generate( **inputs, max_new_tokens=256, do_sample=True, temperature=0.7, top_p=0.8, ) answer = tokenizer.decode( output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True, ) print(answer) ``` **Important:** This example is for Transformers-compatible model files. It will not load a GGUF file directly. For GGUF, use a compatible runtime such as `llama.cpp` and follow the runtime's model-loading instructions. ## Chat Template The original Qwen2.5-Instruct model uses a chat template. If the tokenizer is included and compatible, use `tokenizer.apply_chat_template()` to format messages. ```python messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Summarize the benefits of renewable energy."}, ] ``` ## Quantization Information The checkpoint has been quantized from the base model listed above. The following technical details can be filled in to make the model card reproducible: - **Quantization method:** Not specified - **Bit-width / quantization type:** Not specified - **Quantization framework or tool:** Not specified - **Calibration dataset (if applicable):** Not specified - **Quantization settings:** Not specified - **Original and quantized file sizes:** Not specified - **Benchmark comparison with the original model:** Not provided Quantization can reduce model storage and memory requirements, but the effect on output quality, speed, and hardware compatibility depends on the quantization method and runtime. ## Limitations - Responses may be inaccurate, incomplete, or biased. - Quantization may affect output quality and performance. - Hardware and runtime requirements depend on the checkpoint format and quantization method. - Do not rely on model outputs as the sole basis for high-stakes decisions. ## Evaluation No benchmark or quality evaluation results are documented here. If available, add benchmark names, scores, hardware, inference settings, and a comparison against the original model. ## License and Attribution The official base model is listed under the Apache-2.0 license. Review the [base model license and terms](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) and ensure this derived checkpoint includes any required license and attribution notices before redistribution or commercial use. ## Credits - **Original model:** [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) - **Quantized checkpoint:** [CodeDevX/qwen2.5-1.5b-instruct-quantized](https://huggingface.co/CodeDevX/qwen2.5-1.5b-instruct-quantized) --- *Quantization and repository documentation by CodeDevX.*