Text Generation
Transformers
Safetensors
qwen3
code
reasoning
lora-merged
livecodebench
conversational
text-generation-inference
Instructions to use modrill/code-think-q4b-20260908 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use modrill/code-think-q4b-20260908 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="modrill/code-think-q4b-20260908") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("modrill/code-think-q4b-20260908") model = AutoModelForCausalLM.from_pretrained("modrill/code-think-q4b-20260908", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use modrill/code-think-q4b-20260908 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "modrill/code-think-q4b-20260908" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modrill/code-think-q4b-20260908", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/modrill/code-think-q4b-20260908
- SGLang
How to use modrill/code-think-q4b-20260908 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "modrill/code-think-q4b-20260908" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modrill/code-think-q4b-20260908", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "modrill/code-think-q4b-20260908" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "modrill/code-think-q4b-20260908", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use modrill/code-think-q4b-20260908 with Docker Model Runner:
docker model run hf.co/modrill/code-think-q4b-20260908
Add MANIFEST.sha256
Browse files- MANIFEST.sha256 +20 -0
MANIFEST.sha256
ADDED
|
@@ -0,0 +1,20 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
832dd9e00a68dd83b3c3fb9f5588dad7dcf337a0db50f7d9483f310cd292e92e LICENSE
|
| 2 |
+
5e66b33663c05d41fe3758efd5dde9aaa83e587e032e6f4d458a6fc74bc881f9 OFFICIAL_MERGE_RECEIPT.json
|
| 3 |
+
1592a771cd3fe8b9e6a976c36a2f8edbcd8c373b4ff5645c69911f0944802f33 README.md
|
| 4 |
+
490367bf1e2bd55f5a7381dc4083bcffc18430bcbc26457680c5c17268225bcc adapter/MANIFEST.json
|
| 5 |
+
16a5f02dafeca4c7804f7306e353375ac567d8158ed37baa20b80dcfe9188f33 adapter/TOKEN_ROWS_META.json
|
| 6 |
+
4e04ab4bb3a660f837b6f57d3964cebfdf5b5a58cebc0c5cd9a30372a396212c adapter/adapter_config.json
|
| 7 |
+
619dcee8c5c04416fed0cee928375583332ce2986bf3f51c0caf8408433831ed adapter/adapter_model.safetensors
|
| 8 |
+
407274bf3280d79ed8efabaf3c3b3e75debaeb6b8efdb4433ac0bae934144688 adapter/token_rows_both_sides.safetensors
|
| 9 |
+
87a2728cb8dc9fe424d624542f6060ec05a1d285ebbec578bb078900e33396b5 chat_template.jinja
|
| 10 |
+
40bcf4e0a6a7b5eb8aa582ee23f0c9700d76f147d17159419893bbad0f762176 config.json
|
| 11 |
+
a2ac0d07aac37502207a212b122abe32bd9564266594a317c0ff490ebce1f226 generation_config.json
|
| 12 |
+
0e43b8415e406761c7ab73c8225fde51abdb7f51a4b80a9e90c616f53c3ea419 model.safetensors
|
| 13 |
+
5ebf618c3d547252a8078665415fa4fb3b080d071c7e0f2750f6af44e231bc3e provenance/DEV256_COMPLETE.json
|
| 14 |
+
24fc75c1ce08f0e4b7dcc31684c36c1f95b7032999c0e40308426781728f939c provenance/MQ0_NEAR_DUP.json
|
| 15 |
+
c9f6e3b33a4693354da05fd6f18b3d05932430ca7c8dc9f7f008d3dcc49ac592 provenance/POLICY.json
|
| 16 |
+
5c45fca93b73d0293c39913fc74160cc24f3ed5efa4bfd51582428cc029e1cd2 provenance/RUN_IDENTITY.json
|
| 17 |
+
0304abe092d416ca7974b8f98f39783c52e71467bea8382ed000315ed917ec8a provenance/TRAINING_CONFIG.json
|
| 18 |
+
2ff0f61e200d18dd55b1301f9573e923abd6fe87608220f7126cf709d4cb7288 provenance/build_mix_distill_payload.py
|
| 19 |
+
be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506 tokenizer.json
|
| 20 |
+
1749ac66e3e7ce3862337c23cca5895a2ee131929563e18a944f8c1c80363712 tokenizer_config.json
|