Text Generation
Transformers
Safetensors
GGUF
English
granite
granite-4.2
formal-logic
reasoning
lora
model-merging
wise-ft
reinforcement-learning
grpo
conversational
Instructions to use webAI-Official/TwIL-LM3-Pro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use webAI-Official/TwIL-LM3-Pro with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="webAI-Official/TwIL-LM3-Pro") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("webAI-Official/TwIL-LM3-Pro") model = AutoModelForCausalLM.from_pretrained("webAI-Official/TwIL-LM3-Pro", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webAI-Official/TwIL-LM3-Pro with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Use Docker
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use webAI-Official/TwIL-LM3-Pro with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webAI-Official/TwIL-LM3-Pro" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- SGLang
How to use webAI-Official/TwIL-LM3-Pro with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM3-Pro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM3-Pro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use webAI-Official/TwIL-LM3-Pro with Ollama:
ollama run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- Unsloth Desktop
- Pi
How to use webAI-Official/TwIL-LM3-Pro with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webAI-Official/TwIL-LM3-Pro:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webAI-Official/TwIL-LM3-Pro with Docker Model Runner:
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- Lemonade
How to use webAI-Official/TwIL-LM3-Pro with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webAI-Official/TwIL-LM3-Pro:Q4_K_M
Run and chat with the model
lemonade run user.TwIL-LM3-Pro-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use webAI-Official/TwIL-LM3-Pro with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webAI-Official/TwIL-LM3-Pro:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webAI-Official/TwIL-LM3-Pro with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webAI-Official/TwIL-LM3-Pro:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
VibeThinker-webAI-trained: add estimated tok/s
Browse files
README.md
CHANGED
|
@@ -129,7 +129,7 @@ Cells marked † need the engine note below.
|
|
| 129 |
| **macro gate** | **0.5539** | 0.4313 | 0.4118 | 0.508 | 0.4466 | 0.4218 | 0.2925 | 0.3473 | 0.3757 | 0.5336 | — |
|
| 130 |
| **strict-7** | **0.2879** | 0.1821 | 0.2021 | — | 0.1121 | 0.1971 | 0.1229 | 0.1579 | 0.1714 | 0.2093 | — |
|
| 131 |
| macro_primary | **0.5875** | 0.4825 | 0.4637 | 0.579 | 0.4313 | 0.4475 | 0.3450 | 0.4188 | 0.4213 | 0.5750 | — |
|
| 132 |
-
| tok/s | 21169 † | 21916 † | 28010 † |
|
| 133 |
| mean gen length | 1902 | 2951 | 2688 | — | 2486 | **564** | 696 | 2296 | 1830 | 2094 | 1005 |
|
| 134 |
| **ans/s** | 11.1 † | 7.4 † | 10.4 † | — | 4.5 † | **28.1** | 23.2 | 10.9 | 12.0 | 4.5 | 3.4 |
|
| 135 |
|
|
@@ -153,7 +153,12 @@ family tables score `mcq_answer` with loose-match credit (0.695), so the strict
|
|
| 153 |
six-lane average and strict-7 — which all need strict scoring — are left as —. Its perplexities
|
| 154 |
come from that separate run; the family table records the untuned VibeThinker-3B at 18.70 and
|
| 155 |
21.80 there, against 16.93 and 27.03 in the columns above, so read them within their own source.
|
| 156 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 157 |
|
| 158 |
† **Throughput for TwIL-LM3-Pro, its base and VibeThinker-3B** was measured with the same
|
| 159 |
dedicated protocol and prompt file as the peer columns (128 prompts × 512 generated tokens, EOS
|
|
@@ -241,7 +246,7 @@ spots in absolute terms are `procedural` (strict 0.1200, loose 0.2350) and FOL t
|
|
| 241 |
| math500 | 0.7467 | 0.6567 | 0.7900 | 0.790 | 0.3600 ¶¶¶ | 0.6900 | 0.4233 | 0.7133 | 0.7800 | 0.6100 | **0.8433** |
|
| 242 |
| **macro (10 CoT datasets)** | 0.7901 | 0.7942 | 0.8097 | 0.802 | 0.7683 | 0.7339 | 0.6997 | 0.7523 | 0.7884 | 0.8493 | **0.8689** |
|
| 243 |
| **macro (all 14)** | 0.7425 | 0.7332 | 0.7262 | 0.728 | 0.6611 | 0.6694 | 0.6245 | 0.6814 | 0.7378 | 0.7591 | **0.8086** |
|
| 244 |
-
| tok/s | 21169 † | 21916 † | 28010 † |
|
| 245 |
| mean gen length | ≈792 | ≈1282 | ≈1789 | — | ≈2787 | **482** | 510 | ≈796 | ≈1327 | ≈1931 | 801 |
|
| 246 |
| **ans/s** | 26.7 † | 17.1 † | ≈15.7 † | — | ≈4.0 † | **32.9** | 31.7 | ≈31.7 | ≈16.9 | 4.9 | 4.2 |
|
| 247 |
|
|
|
|
| 129 |
| **macro gate** | **0.5539** | 0.4313 | 0.4118 | 0.508 | 0.4466 | 0.4218 | 0.2925 | 0.3473 | 0.3757 | 0.5336 | — |
|
| 130 |
| **strict-7** | **0.2879** | 0.1821 | 0.2021 | — | 0.1121 | 0.1971 | 0.1229 | 0.1579 | 0.1714 | 0.2093 | — |
|
| 131 |
| macro_primary | **0.5875** | 0.4825 | 0.4637 | 0.579 | 0.4313 | 0.4475 | 0.3450 | 0.4188 | 0.4213 | 0.5750 | — |
|
| 132 |
+
| tok/s | 21169 † | 21916 † | 28010 † | ≈28010 ★ † | 11097 † | 15880 | 16160 | 25230 | 22480 | 9420 | 3374 |
|
| 133 |
| mean gen length | 1902 | 2951 | 2688 | — | 2486 | **564** | 696 | 2296 | 1830 | 2094 | 1005 |
|
| 134 |
| **ans/s** | 11.1 † | 7.4 † | 10.4 † | — | 4.5 † | **28.1** | 23.2 | 10.9 | 12.0 | 4.5 | 3.4 |
|
| 135 |
|
|
|
|
| 153 |
six-lane average and strict-7 — which all need strict scoring — are left as —. Its perplexities
|
| 154 |
come from that separate run; the family table records the untuned VibeThinker-3B at 18.70 and
|
| 155 |
21.80 there, against 16.93 and 27.03 in the columns above, so read them within their own source.
|
| 156 |
+
Its `tok/s` is an **estimate, not a measurement**: merging changes weights but not architecture,
|
| 157 |
+
parameter count or tokenizer, and the throughput protocol fixes the output at 512 generated tokens
|
| 158 |
+
with EOS ignored, so the decode rate does not depend on what the weights say. It is therefore set
|
| 159 |
+
equal to the 28,010 tok/s measured for the base VibeThinker-3B in the same session (a ≈ and both
|
| 160 |
+
marks in the cell). Mean generation length, and so `ans/s`, does depend on the weights and was not
|
| 161 |
+
measured, so those rows stay —.
|
| 162 |
|
| 163 |
† **Throughput for TwIL-LM3-Pro, its base and VibeThinker-3B** was measured with the same
|
| 164 |
dedicated protocol and prompt file as the peer columns (128 prompts × 512 generated tokens, EOS
|
|
|
|
| 246 |
| math500 | 0.7467 | 0.6567 | 0.7900 | 0.790 | 0.3600 ¶¶¶ | 0.6900 | 0.4233 | 0.7133 | 0.7800 | 0.6100 | **0.8433** |
|
| 247 |
| **macro (10 CoT datasets)** | 0.7901 | 0.7942 | 0.8097 | 0.802 | 0.7683 | 0.7339 | 0.6997 | 0.7523 | 0.7884 | 0.8493 | **0.8689** |
|
| 248 |
| **macro (all 14)** | 0.7425 | 0.7332 | 0.7262 | 0.728 | 0.6611 | 0.6694 | 0.6245 | 0.6814 | 0.7378 | 0.7591 | **0.8086** |
|
| 249 |
+
| tok/s | 21169 † | 21916 † | 28010 † | ≈28010 ★ † | 11097 † | 15880 | 16160 | 25230 | 22480 | 9420 | 3374 |
|
| 250 |
| mean gen length | ≈792 | ≈1282 | ≈1789 | — | ≈2787 | **482** | 510 | ≈796 | ≈1327 | ≈1931 | 801 |
|
| 251 |
| **ans/s** | 26.7 † | 17.1 † | ≈15.7 † | — | ≈4.0 † | **32.9** | 31.7 | ≈31.7 | ≈16.9 | 4.9 | 4.2 |
|
| 252 |
|