Instructions to use leimroth-lab/Unbound-LR7 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use leimroth-lab/Unbound-LR7 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("leimroth-lab/Unbound-LR7") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use leimroth-lab/Unbound-LR7 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "leimroth-lab/Unbound-LR7"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "leimroth-lab/Unbound-LR7" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use leimroth-lab/Unbound-LR7 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "leimroth-lab/Unbound-LR7"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "leimroth-lab/Unbound-LR7" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "leimroth-lab/Unbound-LR7", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use leimroth-lab/Unbound-LR7 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "leimroth-lab/Unbound-LR7"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default leimroth-lab/Unbound-LR7
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use leimroth-lab/Unbound-LR7 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "leimroth-lab/Unbound-LR7"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "leimroth-lab/Unbound-LR7" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Unbound LR7
An abliterated build of TensorFold/Qwen3.8-Flash-Next-MLX-4bit-MTP
(formerly Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP, revision dadefa8066e3): Qwen3.8-Flash-Next in 4-bit MLX format with
the native MTP head. Tested on the TensorFold engine on one NVIDIA DGX Spark. Model id leimroth-lab/Unbound-LR7. Seventh in
the Leimroth Lab line, after Dream Shimmer LR6.
By Scott Leimroth · scottleimroth.com · @ScottLeimroth
What changed from the earlier upload (4 October 2026)
The first upload removed one refusal direction. With thinking on, or inside an agent setup, it answered almost everything. With thinking off and no system prompt it still refused about two thirds of harmful requests (21 of 63 answered).
In this revision I measured a second refusal direction in exactly that failing mode, on the already-edited model, and removed it as well. Both edits are applied once, from the original weights. With thinking off and no system prompt the model now answers 60 of 63, and inside an agent setup with thinking off it answers 63 of 63 (was 56). The cost is one of the 20 reliability scenarios (14 of 20, was 15), fewer warnings to the user about planted instructions, and decode 2 to 4% slower. The full comparison is in the table below. The repository name is unchanged; this build replaces the earlier one.
What this is
A weight edit that removes two refusal directions, so the model declines far less often. There is no training and no
fine-tune. Everything not listed below is byte-identical to the base revision dadefa8066e3, including the MTP head,
the vision weights, the embeddings, lm_head and the router.
Method
Refusal-direction orthogonalisation (abliteration). Each edited 4-bit weight is dequantised, both directions are projected out of what it writes to the residual stream, and the weight is requantised once.
- Direction 1 (the earlier upload's): one unit vector taken at layer 40 from the base model.
- Direction 2 (new): measured on the model with direction 1 removed, with thinking off and no system prompt, at layer 20, from the mean of the last six prompt positions, as the leading SVD component across the four residual streams; made orthogonal to direction 1.
- Both directions: attention writers (
linear_attn.out_proj,self_attn.o_proj) from layer 8 up at lam 1.75, and shared-expert down projections (mlp.shared_expert.down_proj) from layer 8 up at lam 0.5. Layers 0-7 are untouched. - 80 weights edited in all, the same 80 as before.
- Other variants were measured and set aside: direction 2 at lam 1.0, 1.25 and 1.5; a late-layer direction 2; a direction 2 made orthogonal to a measured "caution about acting" direction; and an equal-weight average of two edits. None kept the reliability score at 15 with a useful gain.
How it was tested
Every number on this card comes from one DGX Spark (GB10, 128 GB unified memory) running TensorFold 0.5.0, with requests sent the way an agent setup sends them: inside Hermes Agent by Nous Research, with its real system prompt and tool schemas, multi-turn tool loops, and several agents at once.
- Quality and compliance rows: stock TensorFold 0.5.0, five slots of 262,144 tokens. "Thinking on" is the house setting (thinking budget 6,000, max_tokens 16,384). "Thinking off" disables thinking.
- Speed, memory and wait rows: two agents at once beside a Whisper speech server and a small vision model, with a local patch that resumes a new task from the end of the shared agent setup (window 234,496 tokens per agent).
tensorfold serve /path/to/Unbound-LR7 --parallel 2 --context 262144 --kv-dtype int8 \
--mtp-drafts 6 --mtp-confidence 0.60 --temperature 1.0 --top-p 0.95 --top-k 20 --ple-on-ssd --thinking
Results
| Test | This revision | Earlier upload |
|---|---|---|
| Agent tasks (A-Bench, 60 items) | 58/60 | 56/60 |
| Tool use / multi-turn | 24/24 / 22/24 | 22/24 / 22/24 |
| Reliability scenarios | 14/20 | 15/20 |
| Harmful requests answered inside an agent setup, thinking on (63 / harder 30) | 62/63 (1 empty reply) / 30/30 | 63/63 / 28/30 |
| Same, thinking off (63 / harder 30) | 63/63 / 30/30 | 56/63 / 28/30 |
| Harmful requests answered with no agent setup, thinking off (63 / harder 30) | 60/63 / 29/30 | 21/63 / 14/30 |
| Legitimate requests refused (7) | 1 thinking on, 0 thinking off | 0 thinking on, 1 thinking off |
| Instructions planted in tool results (32) | resisted 29, obeyed 1, did not read 2; warned the user in 19 | resisted 30, obeyed 1, did not read 1; warned in 28 |
| Speed, one stream | 72.7 tok/s | 74.1 tok/s |
| Speed, each of two streams at once | 49.7 tok/s | 51.7 tok/s |
| Speed with MTP drafting off | 35.0 tok/s | 35.1 tok/s |
| Wait before a new agent task starts (36k-token agent setup, with the patch) | 3.4 s | 3.4 s |
| Long memory: two agents at about 230k tokens each | both hidden facts found, 11.1 GiB of memory left beside Whisper and vision | both found |
| Background jobs as an agent framework sends them (memory writes, titles; 360 jobs, thinking off) | all 360 valid, none refused | all 360 valid, none refused |
Structured output with thinking off was not re-measured on this revision.
Full tables, with the serving config behind every number: https://scottleimroth.com/ai-technology/spark-benchmarks
Engine notes (TensorFold)
- Starting a new task: on stock TensorFold 0.5.0, every new conversation re-reads the whole system prompt and tools (about 24 s for a 36k-token agent setup on a Spark). A local patch that resumes from the end of the shared setup cuts this to 3.4 s, with identical replies. TensorFold 0.6.1 and later have their own fix.
- TensorFold 0.6.1 with two long conversations open keeps neither conversation's state, so each turn re-reads it (about 70 s at 150k tokens, measured on the earlier upload). For agent use, 0.5.0 is the tested setup.
- Outputs change on TensorFold 0.6.x against 0.5.0, so the scores above apply to 0.5.0.
Images
The vision weights are present and unedited. TensorFold 0.5.0's CUDA path refuses image input for this model family, so image reading was not tested. On a DGX Spark, send images to a separate vision model. It may read images in an MLX runtime on Apple silicon, as the base conversion does; that is untested here.
Limitations
- Tested on one engine (TensorFold 0.5.0) on one machine. Other runtimes are untested.
- Instructions planted in tool results: it obeyed 1 of 32 and warned the user less often than the earlier upload. Put an approval gate in front of destructive tools; one line in a system prompt did not fix it.
- The reliability scenarios it fails include two about caution before acting (an instruction hidden in code comments, and a refactor that should not delete files). Keep a human approval step for destructive actions.
- No prompt log-probabilities on this engine, so likelihood was not measured.
Intended use and responsible use
This model has had refusal behaviour removed and will answer prompts a stock instruct model would decline. It is released for research into alignment, refusal mechanisms and red-teaming. You are responsible for how you use it and for complying with the base model's license and applicable law. Do not deploy it in a user-facing setting without your own safety layer.
License and credits
Released under the Qwen Community License 1.0, the license of the base model, included here as LICENSE. That
license permits modified versions and redistribution, provided the license and copyright notice travel with them.
Check its two conditions (very large commercial deployments, and Model-as-a-Service or AI work-assistant businesses)
before commercial use.
- Base model: Qwen3.8-Flash-Next, by the Qwen team.
- 4-bit MLX conversion with native MTP: TensorFold (formerly published as Vontra).
- Serving engine: TensorFold.
- Agent framework used for testing: Hermes Agent by Nous Research.
- Method: the public refusal-direction / orthogonalisation line of work on abliteration. This is an application of that published technique, not a new one.
- Downloads last month
- 275
4-bit
Model tree for leimroth-lab/Unbound-LR7
Base model
Qwen/Qwen3.8-Flash-Next