Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
BoldingBuilds 
posted an update about 13 hours ago
Post
1396
Abliterated thinking models have a habit: they decide what to say, then keep arguing with themselves until the token budget runs out. Qwen3.8-27B Abliterated ThinkFix is an uncensored Qwen3.8-27B with one output row edited so the model closes its thinking when it's ready to answer.

On 40 sensitive prompts the edit never saw, 8k output cap: 33 replies finished cleanly instead of 22, 5 hit the limit instead of 17, 61 minutes for the set instead of 75.

Agent use was checked four ways and the edit changes nothing there. MTP draft head kept. Q4_K_M to Q8_0, each tested after the edit. The card lists what the edit was fit on and what it does not do.

BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF

Abliterated thinking models have a habit: they decide what to say, then keep arguing with themselves until the token budget runs out.

This is a form of covert noncompliance rather than an error. A model that retains its risk assessment abilities will often continue to obstruct access to sensitive outputs even though refusal is no longer possible. Supervised Reward Preferencing is another way to address residual noncompliance with SOMPOA decensors which more robustly target both the refusal and risk assessment dimensions than either classical abliteration or ARA.

·

I went and checked. The loops on this one aren't it debating whether to answer. That language fades as the trace goes on and it just keeps rewriting the answer inside its thinking. When it stops, the answer's complete. Its a stopping problem, not hidden refusal. Only looked at this model though.