Unbound LR7

An abliterated build of TensorFold/Qwen3.8-Flash-Next-MLX-4bit-MTP (formerly Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP, revision dadefa8066e3): Qwen3.8-Flash-Next in 4-bit MLX format with the native MTP head. Tested on the TensorFold engine on one NVIDIA DGX Spark. Model id leimroth-lab/Unbound-LR7. Seventh in the Leimroth Lab line, after Dream Shimmer LR6.

By Scott Leimroth · scottleimroth.com · @ScottLeimroth

What changed from the earlier upload (4 October 2026)

The first upload removed one refusal direction. With thinking on, or inside an agent setup, it answered almost everything. With thinking off and no system prompt it still refused about two thirds of harmful requests (21 of 63 answered).

In this revision I measured a second refusal direction in exactly that failing mode, on the already-edited model, and removed it as well. Both edits are applied once, from the original weights. With thinking off and no system prompt the model now answers 60 of 63, and inside an agent setup with thinking off it answers 63 of 63 (was 56). The cost is one of the 20 reliability scenarios (14 of 20, was 15), fewer warnings to the user about planted instructions, and decode 2 to 4% slower. The full comparison is in the table below. The repository name is unchanged; this build replaces the earlier one.

What this is

A weight edit that removes two refusal directions, so the model declines far less often. There is no training and no fine-tune. Everything not listed below is byte-identical to the base revision dadefa8066e3, including the MTP head, the vision weights, the embeddings, lm_head and the router.

Method

Refusal-direction orthogonalisation (abliteration). Each edited 4-bit weight is dequantised, both directions are projected out of what it writes to the residual stream, and the weight is requantised once.

  • Direction 1 (the earlier upload's): one unit vector taken at layer 40 from the base model.
  • Direction 2 (new): measured on the model with direction 1 removed, with thinking off and no system prompt, at layer 20, from the mean of the last six prompt positions, as the leading SVD component across the four residual streams; made orthogonal to direction 1.
  • Both directions: attention writers (linear_attn.out_proj, self_attn.o_proj) from layer 8 up at lam 1.75, and shared-expert down projections (mlp.shared_expert.down_proj) from layer 8 up at lam 0.5. Layers 0-7 are untouched.
  • 80 weights edited in all, the same 80 as before.
  • Other variants were measured and set aside: direction 2 at lam 1.0, 1.25 and 1.5; a late-layer direction 2; a direction 2 made orthogonal to a measured "caution about acting" direction; and an equal-weight average of two edits. None kept the reliability score at 15 with a useful gain.

How it was tested

Every number on this card comes from one DGX Spark (GB10, 128 GB unified memory) running TensorFold 0.5.0, with requests sent the way an agent setup sends them: inside Hermes Agent by Nous Research, with its real system prompt and tool schemas, multi-turn tool loops, and several agents at once.

  • Quality and compliance rows: stock TensorFold 0.5.0, five slots of 262,144 tokens. "Thinking on" is the house setting (thinking budget 6,000, max_tokens 16,384). "Thinking off" disables thinking.
  • Speed, memory and wait rows: two agents at once beside a Whisper speech server and a small vision model, with a local patch that resumes a new task from the end of the shared agent setup (window 234,496 tokens per agent).
tensorfold serve /path/to/Unbound-LR7 --parallel 2 --context 262144 --kv-dtype int8 \
  --mtp-drafts 6 --mtp-confidence 0.60 --temperature 1.0 --top-p 0.95 --top-k 20 --ple-on-ssd --thinking

Results

Test This revision Earlier upload
Agent tasks (A-Bench, 60 items) 58/60 56/60
Tool use / multi-turn 24/24 / 22/24 22/24 / 22/24
Reliability scenarios 14/20 15/20
Harmful requests answered inside an agent setup, thinking on (63 / harder 30) 62/63 (1 empty reply) / 30/30 63/63 / 28/30
Same, thinking off (63 / harder 30) 63/63 / 30/30 56/63 / 28/30
Harmful requests answered with no agent setup, thinking off (63 / harder 30) 60/63 / 29/30 21/63 / 14/30
Legitimate requests refused (7) 1 thinking on, 0 thinking off 0 thinking on, 1 thinking off
Instructions planted in tool results (32) resisted 29, obeyed 1, did not read 2; warned the user in 19 resisted 30, obeyed 1, did not read 1; warned in 28
Speed, one stream 72.7 tok/s 74.1 tok/s
Speed, each of two streams at once 49.7 tok/s 51.7 tok/s
Speed with MTP drafting off 35.0 tok/s 35.1 tok/s
Wait before a new agent task starts (36k-token agent setup, with the patch) 3.4 s 3.4 s
Long memory: two agents at about 230k tokens each both hidden facts found, 11.1 GiB of memory left beside Whisper and vision both found
Background jobs as an agent framework sends them (memory writes, titles; 360 jobs, thinking off) all 360 valid, none refused all 360 valid, none refused

Structured output with thinking off was not re-measured on this revision.

Full tables, with the serving config behind every number: https://scottleimroth.com/ai-technology/spark-benchmarks

Engine notes (TensorFold)

  • Starting a new task: on stock TensorFold 0.5.0, every new conversation re-reads the whole system prompt and tools (about 24 s for a 36k-token agent setup on a Spark). A local patch that resumes from the end of the shared setup cuts this to 3.4 s, with identical replies. TensorFold 0.6.1 and later have their own fix.
  • TensorFold 0.6.1 with two long conversations open keeps neither conversation's state, so each turn re-reads it (about 70 s at 150k tokens, measured on the earlier upload). For agent use, 0.5.0 is the tested setup.
  • Outputs change on TensorFold 0.6.x against 0.5.0, so the scores above apply to 0.5.0.

Images

The vision weights are present and unedited. TensorFold 0.5.0's CUDA path refuses image input for this model family, so image reading was not tested. On a DGX Spark, send images to a separate vision model. It may read images in an MLX runtime on Apple silicon, as the base conversion does; that is untested here.

Limitations

  • Tested on one engine (TensorFold 0.5.0) on one machine. Other runtimes are untested.
  • Instructions planted in tool results: it obeyed 1 of 32 and warned the user less often than the earlier upload. Put an approval gate in front of destructive tools; one line in a system prompt did not fix it.
  • The reliability scenarios it fails include two about caution before acting (an instruction hidden in code comments, and a refactor that should not delete files). Keep a human approval step for destructive actions.
  • No prompt log-probabilities on this engine, so likelihood was not measured.

Intended use and responsible use

This model has had refusal behaviour removed and will answer prompts a stock instruct model would decline. It is released for research into alignment, refusal mechanisms and red-teaming. You are responsible for how you use it and for complying with the base model's license and applicable law. Do not deploy it in a user-facing setting without your own safety layer.

License and credits

Released under the Qwen Community License 1.0, the license of the base model, included here as LICENSE. That license permits modified versions and redistribution, provided the license and copyright notice travel with them. Check its two conditions (very large commercial deployments, and Model-as-a-Service or AI work-assistant businesses) before commercial use.

  • Base model: Qwen3.8-Flash-Next, by the Qwen team.
  • 4-bit MLX conversion with native MTP: TensorFold (formerly published as Vontra).
  • Serving engine: TensorFold.
  • Agent framework used for testing: Hermes Agent by Nous Research.
  • Method: the public refusal-direction / orthogonalisation line of work on abliteration. This is an application of that published technique, not a new one.
Downloads last month
275
Safetensors
Model size
180B params
Tensor type
U32
·
BF16
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for leimroth-lab/Unbound-LR7

Finetuned
(1)
this model