CodeDevX commited on
Commit
40d2936
·
verified ·
1 Parent(s): 6d24ddb

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +141 -0
README.md ADDED
@@ -0,0 +1,141 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-1.5B-Instruct
4
+ pipeline_tag: text-generation
5
+ library_name: transformers
6
+ tags:
7
+ - qwen2
8
+ - quantized
9
+ - text-generation
10
+ - chat
11
+ ---
12
+
13
+ # Qwen2.5-1.5B-Instruct-Quantized
14
+
15
+ A quantized checkpoint created from **[Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)**.
16
+
17
+ This repository contains a quantized version of the original instruction-tuned Qwen2.5 1.5B model. The base model was developed by the Qwen team. This repository is a community quantization, not the original Qwen release. The base model is a causal language model designed for instruction following and conversational text generation. See the [official base model card](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) for its architecture and original documentation. citeturn953640search0
18
+
19
+ ## Model Details
20
+
21
+ | Field | Details |
22
+ |---|---|
23
+ | Model name | Qwen2.5-1.5B-Instruct-Quantized |
24
+ | Base model | [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) |
25
+ | Model family | Qwen2.5 |
26
+ | Model size | 1.5B-class model (base model: approximately 1.54B parameters) |
27
+ | Task | Text generation, chat, instruction following |
28
+ | Quantization | Quantized by the repository maintainer; method and bit-width not specified |
29
+ | Maintainer | CodeDevX |
30
+
31
+ ## Intended Use
32
+
33
+ This model may be used for:
34
+
35
+ - Conversational assistance and instruction following
36
+ - General text generation and question answering
37
+ - Summarization, rewriting, and drafting
38
+ - Experimentation with quantized language models and local inference, subject to runtime compatibility
39
+
40
+ ## Quick Start
41
+
42
+ The exact loading method depends on the quantization format used for this checkpoint. Check the repository's **Files and versions** tab for the model file extension and configuration before choosing an inference runtime.
43
+
44
+ ### Transformers (for Transformers-compatible checkpoints)
45
+
46
+ If the uploaded files are compatible with Transformers, you can try:
47
+
48
+ ```bash
49
+ pip install -U transformers torch accelerate safetensors
50
+ ```
51
+
52
+ ```python
53
+ import torch
54
+ from transformers import AutoTokenizer, AutoModelForCausalLM
55
+
56
+ model_id = "CodeDevX/qwen2.5-1.5b-instruct-quantized"
57
+
58
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
59
+ model = AutoModelForCausalLM.from_pretrained(
60
+ model_id,
61
+ torch_dtype="auto",
62
+ device_map="auto",
63
+ )
64
+
65
+ messages = [
66
+ {"role": "user", "content": "Explain quantization in simple terms."}
67
+ ]
68
+
69
+ prompt = tokenizer.apply_chat_template(
70
+ messages,
71
+ tokenize=False,
72
+ add_generation_prompt=True,
73
+ )
74
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
75
+
76
+ with torch.no_grad():
77
+ output = model.generate(
78
+ **inputs,
79
+ max_new_tokens=256,
80
+ do_sample=True,
81
+ temperature=0.7,
82
+ top_p=0.8,
83
+ )
84
+
85
+ answer = tokenizer.decode(
86
+ output[0][inputs["input_ids"].shape[1]:],
87
+ skip_special_tokens=True,
88
+ )
89
+ print(answer)
90
+ ```
91
+
92
+ **Important:** This example is for Transformers-compatible model files. It will not load a GGUF file directly. For GGUF, use a compatible runtime such as `llama.cpp` and follow the runtime's model-loading instructions.
93
+
94
+ ## Chat Template
95
+
96
+ The original Qwen2.5-Instruct model uses a chat template. If the tokenizer is included and compatible, use `tokenizer.apply_chat_template()` to format messages.
97
+
98
+ ```python
99
+ messages = [
100
+ {"role": "system", "content": "You are a helpful assistant."},
101
+ {"role": "user", "content": "Summarize the benefits of renewable energy."},
102
+ ]
103
+ ```
104
+
105
+ ## Quantization Information
106
+
107
+ The checkpoint has been quantized from the base model listed above. The following technical details can be filled in to make the model card reproducible:
108
+
109
+ - **Quantization method:** Not specified
110
+ - **Bit-width / quantization type:** Not specified
111
+ - **Quantization framework or tool:** Not specified
112
+ - **Calibration dataset (if applicable):** Not specified
113
+ - **Quantization settings:** Not specified
114
+ - **Original and quantized file sizes:** Not specified
115
+ - **Benchmark comparison with the original model:** Not provided
116
+
117
+ Quantization can reduce model storage and memory requirements, but the effect on output quality, speed, and hardware compatibility depends on the quantization method and runtime.
118
+
119
+ ## Limitations
120
+
121
+ - Responses may be inaccurate, incomplete, or biased.
122
+ - Quantization may affect output quality and performance.
123
+ - Hardware and runtime requirements depend on the checkpoint format and quantization method.
124
+ - Do not rely on model outputs as the sole basis for high-stakes decisions.
125
+
126
+ ## Evaluation
127
+
128
+ No benchmark or quality evaluation results are documented here. If available, add benchmark names, scores, hardware, inference settings, and a comparison against the original model.
129
+
130
+ ## License and Attribution
131
+
132
+ The official base model is listed under the Apache-2.0 license. Review the [base model license and terms](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) and ensure this derived checkpoint includes any required license and attribution notices before redistribution or commercial use. citeturn953640search0
133
+
134
+ ## Credits
135
+
136
+ - **Original model:** [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
137
+ - **Quantized checkpoint:** [CodeDevX/qwen2.5-1.5b-instruct-quantized](https://huggingface.co/CodeDevX/qwen2.5-1.5b-instruct-quantized)
138
+
139
+ ---
140
+
141
+ *Quantization and repository documentation by CodeDevX.*