How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="CodeDevX/qwen2.5-1.5b-instruct-quantized")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("CodeDevX/qwen2.5-1.5b-instruct-quantized", device_map="auto")
Quick Links

Qwen2.5-1.5B-Instruct-Quantized

A quantized checkpoint created from Qwen/Qwen2.5-1.5B-Instruct.

This repository contains a quantized version of the original instruction-tuned Qwen2.5 1.5B model. The base model was developed by the Qwen team. This repository is a community quantization, not the original Qwen release. The base model is a causal language model designed for instruction following and conversational text generation. See the official base model card for its architecture and original documentation.

Model Details

Field Details
Model name Qwen2.5-1.5B-Instruct-Quantized
Base model Qwen/Qwen2.5-1.5B-Instruct
Model family Qwen2.5
Model size 1.5B-class model (base model: approximately 1.54B parameters)
Task Text generation, chat, instruction following
Quantization Quantized by the repository maintainer; method and bit-width not specified
Maintainer CodeDevX

Intended Use

This model may be used for:

  • Conversational assistance and instruction following
  • General text generation and question answering
  • Summarization, rewriting, and drafting
  • Experimentation with quantized language models and local inference, subject to runtime compatibility

Quick Start

The exact loading method depends on the quantization format used for this checkpoint. Check the repository's Files and versions tab for the model file extension and configuration before choosing an inference runtime.

Transformers (for Transformers-compatible checkpoints)

If the uploaded files are compatible with Transformers, you can try:

pip install -U transformers torch accelerate safetensors
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "CodeDevX/qwen2.5-1.5b-instruct-quantized"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Explain quantization in simple terms."}
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=True,
        temperature=0.7,
        top_p=0.8,
    )

answer = tokenizer.decode(
    output[0][inputs["input_ids"].shape[1]:],
    skip_special_tokens=True,
)
print(answer)

Important: This example is for Transformers-compatible model files. It will not load a GGUF file directly. For GGUF, use a compatible runtime such as llama.cpp and follow the runtime's model-loading instructions.

Chat Template

The original Qwen2.5-Instruct model uses a chat template. If the tokenizer is included and compatible, use tokenizer.apply_chat_template() to format messages.

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Summarize the benefits of renewable energy."},
]

Quantization Information

The checkpoint has been quantized from the base model listed above. The following technical details can be filled in to make the model card reproducible:

  • Quantization method: Not specified
  • Bit-width / quantization type: Not specified
  • Quantization framework or tool: Not specified
  • Calibration dataset (if applicable): Not specified
  • Quantization settings: Not specified
  • Original and quantized file sizes: Not specified
  • Benchmark comparison with the original model: Not provided

Quantization can reduce model storage and memory requirements, but the effect on output quality, speed, and hardware compatibility depends on the quantization method and runtime.

Limitations

  • Responses may be inaccurate, incomplete, or biased.
  • Quantization may affect output quality and performance.
  • Hardware and runtime requirements depend on the checkpoint format and quantization method.
  • Do not rely on model outputs as the sole basis for high-stakes decisions.

Evaluation

No benchmark or quality evaluation results are documented here. If available, add benchmark names, scores, hardware, inference settings, and a comparison against the original model.

License and Attribution

The official base model is listed under the Apache-2.0 license. Review the base model license and terms and ensure this derived checkpoint includes any required license and attribution notices before redistribution or commercial use.

Credits


Quantization and repository documentation by CodeDevX.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.9B params
Tensor type
F32
路
BF16
路
U8
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for CodeDevX/qwen2.5-1.5b-instruct-quantized

Finetuned
(1922)
this model