Instructions to use YCF-AI/CyberTiel-Coder-35B-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use YCF-AI/CyberTiel-Coder-35B-MLX with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("YCF-AI/CyberTiel-Coder-35B-MLX") config = load_config("YCF-AI/CyberTiel-Coder-35B-MLX") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use YCF-AI/CyberTiel-Coder-35B-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "YCF-AI/CyberTiel-Coder-35B-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "YCF-AI/CyberTiel-Coder-35B-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use YCF-AI/CyberTiel-Coder-35B-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "YCF-AI/CyberTiel-Coder-35B-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default YCF-AI/CyberTiel-Coder-35B-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use YCF-AI/CyberTiel-Coder-35B-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "YCF-AI/CyberTiel-Coder-35B-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "YCF-AI/CyberTiel-Coder-35B-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
CyberTiel-Coder-35B-MLX (fixed MTP)
Same weights as the official Cyber-Tiel oQ6e build — with an MTP head that actually works in the oMLX VLM engine: vision + MTP + thinking at the same time, at ~+35% decode speed, at half the head size.
⚠️ Read before use — abliterated model. CyberTiel is an uncensored (abliterated) model. It can and will produce content other models refuse. You take full personal responsibility for your use of it. Run it sandboxed at the OS level, restrict its network and code-execution access, and treat it as hostile toward prompt injection. Full warnings in the upstream model card (preserved in this repo as
README-upstream.md).⚠️ Do not run this build in LM Studio. Carried over from upstream: LM Studio's MLX engine mis-executes MTP-head builds. Use oMLX (validated) or the upstream GGUF builds.
What this is
A drop-in replacement for
peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e-MTP
on the oMLX runtime, fixing its two practical problems on
Apple Silicon:
Upstream …-oQ6e-MTP |
This build (fixed MTP) |
|
|---|---|---|
| Vision + MTP + thinking | Mutually exclusive (pick 2 of 3) | All three at once |
| Decode speed (MTP d2, 4K ctx) | ~74 tok/s | ~99 tok/s (+33%) |
| MTP head file | 1.69 GB BF16 | 906 MB 8-bit (~46% smaller) |
| Quality (260-Q local gate) | 207/260 | 207/260 — bit-identical behavior |
Upstream's own card notes that on MLX "the MTP head does not help". We show that's a runtime-loading artifact, not a property of the model: once the head is packaged in the layout oMLX's VLM engine expects, MTP d2 delivers +82% to +107% over no-MTP decoding on M1 Ultra (same-context medians).
Lineage
ornith-ai/Ornith-1.5-35B-A3B (base, MoE 35B / 3B active, vision)
└─ huihui-ai/Huihui-Ornith-…-abliterated (refusal removed)
└─ peculiar-ragdoll/…-MLX-oQ6e-MTP (oQ6e quant + grafted BF16 MTP head) ← MIT
└─ this repo (MTP head re-quantized into fused layout)
All six main weight shards are byte-for-byte identical to the upstream oQ6e build (APFS-cloned locally; upstream file hashes unchanged). Only the MTP head shard, its config entries and the weights index differ.
What we changed
One targeted engineering fix, zero retraining, zero main-weight changes:
- Reverse-engineered the working MTP-head layout. oMLX's production Ornith build ships
its MTP head as 42 tensors under
language_model.mtp.*— 8-bit affine quantization (U32-packed 4×8-bit LSB-first, BF16 scales/biases, per-group), with MoE experts fused into a singleswitch_mlp3D tensor. We verified our format understanding by dequantizing that head against its BF16 source: cosine ≥ 0.9996. - Re-packaged the official CyberTiel MTP head into that layout. The upstream head is
785 explicitly-enumerated BF16 tensors, which oMLX 0.6.4's VLM loader rejects
(
Received 768 parameters not in model) — that rejection is what forced the runtime into a fallback path where vision and MTP exclude each other. We stacked the experts into fusedswitch_mlptensors, quantized the 11 matrices to the same 8-bit affine scheme (small tensors stay BF16), and emitted a 42-key, 906 MBmodel-mtp.safetensors(replacing the 1.69 GB BF16 head). Self-check: cosine ≥ 0.9999 per tensor. - Wired it up. 22 per-key quantization entries added to
config.json, weights index rewritten. Result: the model loads intoVLMBatchedEnginewith Lightning MTP active — first time on this machine CyberTiel runs vision, MTP and thinking simultaneously.
Ablation we ran along the way: grafting the shisa-ai repaired MTP head (BF16, distillation variant) onto this base hurt (−6% vs no-MTP) — its distribution doesn't match the huihui abliterated base. The official head + our packaging is what works. We did not re-distill; with the official head already reaching parity with production Ornith decode speed, there was no headroom worth the effort.
Measured results
Environment: Mac Studio M1 Ultra (64 GB unified, 64-core GPU), oMLX 0.6.4 (build 2529 / runtime 260830014856), single serial request, cache disabled, 384-token decode windows, ABBA-interleaved, n=3 medians. These are our own measurements, not vendor numbers.
Speed (MTP depth 2, the frozen production setting)
| Context | Upstream oQ6e-MTP | This build | Gain |
|---|---|---|---|
| 1K — decode tok/s | 75.3–76.4 | 102.8 | +35% |
| 4K — decode tok/s | 74.2 | 99.0 | +33% |
| 16K — decode tok/s | 64.8 | 88.3 | +36% |
| MTP d2 vs d0 (same ctx) | +34% | +82% ~ +107% | |
| 16K — prefill tok/s | 910 | 928 | +2% |
For scale: production Ornith-1.5 oQ4e-fixed-mtp on the same box measures 100–105 / 102.5 / 87.5 tok/s — i.e. this build closes what used to be a 35% gap to a draw, while scoring higher on the quality gate below.
Quality (260-question local gate; deterministic sampling, thinking off)
| Suite | Upstream | This build |
|---|---|---|
| MMLU (50) | 32 | 32 |
| CMMLU (50) | 37 | 37 |
| TruthfulQA (50) | 46 | 46 |
| GSM8K (50) | 49 | 49 |
| HumanEval (20) | 16 | 16 |
| MBPP (30) | 24 | 24 |
| LiveCodeBench (10) | 3 | 3 |
| Total | 207/260 | 207/260 |
Identical scores suite-by-suite — expected and verified: the main network is untouched, and a draft head only proposes tokens that the main network still verifies token-by-token, so output distribution equals the base distribution. The MTP fix is free quality-wise.
Against the production Ornith oQ4e on the same gate (190/260, twice reproduced): +17 questions, led by CMMLU +12 and HumanEval/MBPP +3 — the best local quality score of any model we have kept on this machine.
Capability checks
- Vision discrimination probes (red/blue/green + no-image control): 4/4
- Long-context needle-in-haystack + JSON retrieval at 7.4K / 28.8K / 53.8K tokens: 3/3
- Thinking mode end-to-end (reasoning chain + correct code task output): pass
- 6144-token counting endurance: 1208/1209 numbers emitted, single discontinuity exactly at the truncation boundary — no mid-sequence degradation
All raw measurement files (speed AB logs, accuracy gate, vision probes, needle runs) and build scripts were produced under independent review; conclusions were cross-checked by a separate reviewer pass with medians recomputed from raw data.
Run it
oMLX ≥ 0.6.4, Apple Silicon with ≥ 36 GB unified memory (comfortably 48 GB with long context). Settings validated on M1 Ultra — the ones we froze in production:
{
"max_context_window": 131072,
"max_tokens": 32768,
"temperature": 0.6,
"top_p": 0.95,
"top_k": 20,
"enable_thinking": true,
"mtp_enabled": true,
"mtp_num_draft_tokens": 2
}
- MTP depth 2 is the sweet spot here (d3 measured slightly slower).
- Thinking budget: leave uncapped — low budgets measurably hurt this model.
- Sampling follows the official CyberTiel recommendations (T 0.6 / P 0.95 / K 20; CTF work may prefer K 40 / min_p 0.05).
- LM Studio: not supported for this build — see warning above.
For the GGUF ladder (llama.cpp, LM Studio) and the full CyberTiel benchmark plates
(SWE-bench-Live, Cybench, HarmBench — measured upstream on the GGUF build), see the
upstream model card
or README-upstream.md in this repo.
Credits & license
- peculiar-ragdoll — the Cyber-Tiel series: oQ6e quantization, cyber-weighted imatrix, grafted MTP head, benchmark plates. This repo is a compatibility repack of their work.
- huihui-ai — the abliterated base.
- ornith-ai — the Ornith-1.5 base model.
- MTP-head packaging approach derived by studying
shisa-ai/Ornith-1.5-35B-A3B-MTPand a production oMLX oQ4e build.
MIT license — see LICENSE. The main weights remain the upstream oQ6e build, redistributed
under its MIT terms; our additions (re-quantized MTP head shard, config/index wiring, this
card) under MIT as well.
中文说明
这是什么
本仓库是 peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e-MTP
在 oMLX 运行时上的直接替代件,修复了它在 Apple Silicon 上的两个实际问题:
官方 …-oQ6e-MTP |
本件(fixed MTP) | |
|---|---|---|
| 视觉 + MTP + 思考 | 三者互斥(只能二选一) | 同时可用 |
| 解码速度(MTP d2、4K 上下文) | ~74 tok/s | ~99 tok/s(+33%) |
| MTP 头文件 | 1.69 GB BF16 | 906 MB 8-bit(缩小约 46%) |
| 质量(260 题本地门禁) | 207/260 | 207/260 —— 行为逐位一致 |
官方模型卡称"MLX 上 MTP 无益"。我们的实测表明这是运行时加载路径的问题,不是模型的问题: 把 MTP 头打包成 oMLX VLM 引擎期望的布局后,MTP d2 在 M1 Ultra 上带来 +82% ~ +107% 的解码增益(同上下文中位口径)。
继承链:Ornith-1.5-35B-A3B(基座)→ huihui-ai abliterated(去拒答)→ peculiar-ragdoll oQ6e 量化 + 嫁接 BF16 MTP 头(MIT)→ 本仓库(MTP 头重量化为融合布局)。 六个主权重分片与官方 oQ6e 件逐字节相同;仅 MTP 头分片、config 条目和权重索引不同。
我们做了什么
一次针对性的工程修复,零重训练、零主权重改动:
- 逆向了"能用"的 MTP 头布局。oMLX 生产版 Ornith 件的 MTP 头是
language_model.mtp.*命名下的 42 个张量:8-bit affine 量化(U32 打包 4×8-bit LSB-first、BF16 scales/biases、 per-group),MoE 专家融合为单个switch_mlp3D 张量。我们用反量化对照 BF16 原值 (cosine ≥ 0.9996)验证了对格式的理解完全正确。 - 把官方 MTP 头重新打包成该布局。官方头是 785 个显式枚举的 BF16 张量,oMLX 0.6.4 的
VLM 加载器会拒绝(
Received 768 parameters not in model)——正是这个拒绝迫使运行时回退到 视觉与 MTP 互斥的路径。我们将专家堆叠融合为switch_mlp,11 个矩阵按同款 8-bit affine 方案量化(小张量保留 BF16),生成 42 键、906 MB 的model-mtp.safetensors(替代 1.69 GB 的 BF16 头),逐张量自检 cosine ≥ 0.9999。 - 接线:config 增补 22 条量化条目、重写权重索引。模型由此进入
VLMBatchedEngine+ Lightning MTP 路径——本机上 CyberTiel 首次同时跑起视觉、MTP 与思考。
过程中的反例:把 shisa-ai 的"修复头"(BF16 蒸馏变体)嫁接到该底座上反而负收益 (比不开 MTP 慢 6%)——它的分布与 huihui abliterated 底座不匹配。官方原生头 + 我们的打包 才是正解;官方头已把解码速度追平生产 Ornith,我们没有再投入重蒸馏。
实测结果
环境:Mac Studio M1 Ultra(64 GB 统一内存 / 64 核 GPU)、oMLX 0.6.4(build 2529 / runtime 260830014856)、单请求串行、关闭缓存、384 token 解码窗口、ABBA 交错、n=3 取中位。 以下均为我们自己的实测,非厂商数据。
速度(MTP 深度 2,冻结的生产设置)
| 上下文 | 官方 oQ6e-MTP | 本件 | 增益 |
|---|---|---|---|
| 1K 解码 tok/s | 75.3–76.4 | 102.8 | +35% |
| 4K 解码 tok/s | 74.2 | 99.0 | +33% |
| 16K 解码 tok/s | 64.8 | 88.3 | +36% |
| MTP d2 相对 d0(同上下文) | +34% | +82% ~ +107% | |
| 16K prefill tok/s | 910 | 928 | +2% |
参照:同机生产 Ornith oQ4e-fixed-mtp 实测 100–105 / 102.5 / 87.5 tok/s——即本件把原本 35% 的速度差距追平,同时质量门禁得分更高。
质量(260 题本地门禁;确定性采样、关思考)
| 套件 | 官方件 | 本件 |
|---|---|---|
| MMLU(50) | 32 | 32 |
| CMMLU(50) | 37 | 37 |
| TruthfulQA(50) | 46 | 46 |
| GSM8K(50) | 49 | 49 |
| HumanEval(20) | 16 | 16 |
| MBPP(30) | 24 | 24 |
| LiveCodeBench(10) | 3 | 3 |
| 合计 | 207/260 | 207/260 |
逐套件完全一致——符合预期且已验证:主干未动,draft 头只提案、主干逐 token 验证, 输出分布恒等于主干分布。MTP 修复在质量上是免费的。对比同门禁下的生产 Ornith oQ4e (两次复测稳定 190/260):净胜 17 题(CMMLU +12、HumanEval/MBPP +3),是本机留存模型 中的最高质量分。
能力检查
- 视觉判别探针(红/蓝/绿 + 无图对照):4/4
- 长上下文 needle + JSON 检索(7.4K / 28.8K / 53.8K token):3/3
- 思考模式端到端(推理链正常 + 代码任务输出正确):通过
- 6144 token 计数耐力:产出 1208/1209 个数字,唯一断点恰在截断边界——序列中段无退化
全部原始测量文件(速度 AB 日志、精度门禁、视觉探针、needle 记录)与构建脚本均经过 独立复核:中位数从原始数据重算,结论二次确认。
怎么用
oMLX ≥ 0.6.4,Apple Silicon,≥ 36 GB 统一内存(长上下文建议 48 GB)。M1 Ultra 上验证并 冻结的生产参数:
{
"max_context_window": 131072,
"max_tokens": 32768,
"temperature": 0.6,
"top_p": 0.95,
"top_k": 20,
"enable_thinking": true,
"mtp_enabled": true,
"mtp_num_draft_tokens": 2
}
- MTP 深度 2 是甜点(实测 d3 略慢)。
- 思考预算不设上限——低预算对该模型有实测损害。
- 采样沿用官方 CyberTiel 建议(T 0.6 / P 0.95 / K 20;CTF 场景可用 K 40 / min_p 0.05)。
- LM Studio 不支持本件(见文首警告)。
GGUF 阶梯(llama.cpp / LM Studio)与 CyberTiel 完整基准图版(SWE-bench-Live、Cybench、
HarmBench——上游在 GGUF 件上所测)请见
上游模型卡
或仓库内 README-upstream.md。
⚠️ 再次提醒:这是 abliterated(无拒答)模型。请务必在操作系统级沙箱中运行, 限制其网络与代码执行权限,并防范提示注入。详见上游模型卡的完整警告。
致谢与许可
- peculiar-ragdoll —— Cyber-Tiel 全系列:oQ6e 量化、cyber 权重 imatrix、嫁接 MTP 头、 基准图版。本仓库只是对其工作的兼容性再打包。
- huihui-ai —— abliterated 底座。
- ornith-ai —— Ornith-1.5 基座模型。
- MTP 头打包方法参考了
shisa-ai/Ornith-1.5-35B-A3B-MTP与一个 oMLX 生产 oQ4e 件。
MIT 许可——见 LICENSE。主权重仍为上游 oQ6e 件,按其 MIT 条款再分发;
我们的增量(重量化 MTP 头分片、config/索引接线、本模型卡)同样为 MIT。
- Downloads last month
- 25
6-bit
Model tree for YCF-AI/CyberTiel-Coder-35B-MLX
Base model
ornith-ai/Ornith-1.5-35B-A3B