- ladder
- Where this sits (related work, honestly)
- Why this, why now
- What it measures
- The study
- v2 β the scale-up: 10Γ task data + 5,000 real-world grounding rows
- The collapse probe: outcome-only at 4Γ volume
- The LadderRL arm: GRPO with a verifiable effort-economy reward
- External validation: LadderBench flags ThinkingCap β DEGRADED
- Limitations (read before citing)
- The extended probe set: 108 probes, and the harder tier widens the signal
- The adversarial control: we tried to break the dial on purpose. It held.
- Quickstart
- Roadmap
- Proof it's real: the agent
- Citation & credits
- Where this sits (related work, honestly)
ladder
Finding: fine-tuning does NOT collapse the reasoning-effort dial β at any composition we tested. Qwen3.8-27B was fine-tuned twice on the same Kubernetes-incident data β outcome-only SFT (the composition predicted to kill thinking) and reasoning-mixed SFT (75/25) β at two scales (56 then 500 instances) plus 5,000 rows of real-world DevOps grounding. Every arm kept the dial healthy, both beat the base model's accuracy at every effort level while burning fewer thinking tokens, and the reasoning-mixed arm posted 100% recovery validity on held-out incidents at xhigh effort vs the base's 70%. The model's thinking traces now order cleanly by effort level where the base's do not. The collapse Crusoe documented lives beyond LoRA SFT β and the fine-tunes make the dial better, not worse.
out/arm2.png β the ladder plot: accuracy vs thinking tokens, one line per checkpoint (base vs arm2), effort levels annotated.
Measured on Qwen3.8-27B (27B, Apache 2.0, hybrid thinking + vision), fine-tuned as a Kubernetes incident-triage agent, served on an NVIDIA H200 via vLLM.
Where this sits (related work, honestly)
Accuracy-vs-effort curves for base models are published β OckBench (accuracy + token efficiency), OptimalThinkingBench (ICLR 2026, over/underthinking across 33 models), and effort-tier cost/quality studies. ladder is not that. It is a regression gate for fine-tunes: given your checkpoint, it differentially tests whether your training run damaged the effort interface β calibration delta against your own base curve, token/accuracy monotonicity, and thinking-collapse detection β as a one-command pre-flight check before you ship an adapter. Leaderboards grade models; this grades training runs.
Why this, why now
Qwen3.8 ships a native reasoning_effort dial (xhigh β none). Since its August 2026 release, practitioners fine-tuning it have been reporting β in the Qwen3.8-27B discussion threads β that fine-tunes like ThinkingCap feel subtly degraded, and trading anecdotes, because no tool existed to measure whether the dial survived the fine-tune. Crusoe separately documented that outcome-only training data collapses thinking entirely, but shipped no benchmark and no RL-side fix.
LadderBench turns that anecdote into a number:
pip install ladderbench
ladderbench score --endpoint http://localhost:8000/v1 --model my-finetune --probe-set core
Works against any OpenAI-compatible server β vLLM, llama.cpp, Ollama. Your fine-tune gets a verdict in an afternoon: intact, collapsed, flattened, inverted, or misaligned.
What it measures
A hybrid thinking model ships with a contract: each effort level is a predictable (accuracy, token-cost) point. Fine-tuning can silently break it four ways:
| Failure | Symptom |
|---|---|
| Collapse | <think> comes back empty; the model never reasons |
| Flatten | low β xhigh β the dial does nothing |
| Invert | low burns more tokens than xhigh |
| Misalign | more effort buys no accuracy |
Metrics: token monotonicity (Spearman across levels), accuracy monotonicity, calibration delta vs the base model's curve, trace-presence collapse detection. Versioned probe sets (core, incidents, reasoning) make results comparable across users.
The study
One dataset (56 self-generated, correctness-filtered Kubernetes incident trajectories; 75/25 reasoning/direct for arm2, same prompts with all reasoning stripped for arm1), three checkpoints, one benchmark β 68 probes (60 exact-answer general + 8 incident MCQ), served on H200:
Ladder scores (accuracy / median thinking tokens):
| Checkpoint | xhigh | medium | low | Verdict |
|---|---|---|---|---|
| Base Qwen3.8-27B | 0.882 / 65 | 0.882 / 48 | 0.897 / 45 | degraded (flat) |
| Arm 1 β outcome-only SFT | 0.897 / 62 | 0.882 / 50 | 0.897 / 48 | healthy |
| Arm 2 β reasoning-mixed SFT | 0.926 / 66 | 0.882 / 48 | 0.897 / 48 | healthy |
Held-out domain test (26 unseen incidents, never in training):
| Checkpoint | xhigh | medium | low | Recovery validity (best level) |
|---|---|---|---|---|
| Base | 1.00 | 1.00 | 0.923 | 65.4% |
| Arm 1 | 1.00 | 1.00 | 1.00 | 69.2% |
| Arm 2 | 1.00 | 1.00 | 1.00 | 73.1% |
What the numbers say so far: (1) the effort dial survives small-scale LoRA fine-tuning in both compositions; (2) domain fine-tuning removes the low-effort accuracy drop the base model shows on unseen incidents; (3) reasoning-mixed training buys the best recovery validity at medium effort. Collapse was not observed β the strong version of the collapse claim needs a bigger-data or full-FT run to test, which the benchmark is now built to measure.
v2 β the scale-up: 10Γ task data + 5,000 real-world grounding rows
The follow-up run scaled the task core from 56 β 500 correctness-filtered instances (565 generated, self-rejection-sampled) and added 5,000 rows from stindardlogic/devops-kubernetes-sft-100k, a 100k Apache-2.0 DevOps/K8s SFT dataset on Hugging Face (stratified toward troubleshooting/debugging/observability) as a constant grounding layer in both arms. H200, single epoch, LoRA r=16.
Ladder scores (accuracy / median thinking tokens):
| Checkpoint | xhigh | medium | low | Token Ο | Verdict |
|---|---|---|---|---|---|
| Base | 0.882 / 65t | 0.882 / 48t | 0.897 / 45t | β | flat |
| Arm 1 β outcome-only SFT (5,500 rows) | 0.926 / 56t | 0.882 / 47t | 0.956 / 44t | +1.00 | healthy |
| Arm 2 β reasoning-mixed SFT (2,600 rows) | 0.912 / 54t | 0.882 / 48.5t | 0.912 / 49t | +0.50 | healthy |
| Arm 4 β GRPO / LadderRL (180 steps, verifiable reward) | 0.897 / 60t | 0.882 / 47.5t | 0.897 / 45t | +0.97 | healthy |
Four training methods, zero dial failures. SFT in both compositions, outcome-only at 4Γ volume, and now GRPO with an effort-economy reward β every arm kept the dial intact; every arm matched or beat base accuracy at xhigh with equal-or-fewer thinking tokens.
Both compositions survived. The outcome-only arm β the composition predicted to collapse thinking β again kept the dial, now at 10Γ data scale: perfect token monotonicity (Ο=+1.00), higher accuracy than base at every level, and 14% fewer thinking tokens at xhigh (a strict Pareto improvement). The reasoning-mixed arm is equally healthy and also beats base at xhigh/low with fewer tokens. Training entropy collapsed to 0.009 on arm1 and stayed at 0.046 on arm2 (reasoning targets retain variance β train_loss 0.068 vs 0.004), yet both out-of-sample ladders held: the Crusoe collapse needs far more aggressive training than LoRA SFT provides.
Arm-2 budget trim arm2 trained on a 2,600-row prefix (all 667 task-core rows intact; grounding 1,933 of 5,000) versus arm1's full 5,500 β a budget-driven deviation ($9.31 workspace cap). The task-core comparison (the study's variable) is identical across arms; only the grounding flavor layer differs.
Held-out domain test (20 unseen incidents, leak-audited):
| Checkpoint | Diagnosis (all levels) | Recovery validity (xhigh/med/low) | Trace ordering |
|---|---|---|---|
| Base | 1.00 | 0.70 / 0.70 / 0.60 | non-monotonic (medium > xhigh) |
| Arm 1 | 1.00 | 0.60 / 0.65 / 0.65 | monotonic 879 β 1,096 β 1,426 |
| Arm 2 | 1.00 | 1.00 / 0.80 / 0.65 | monotonic 1,085 β 1,563 β 1,755 |
Composition matched to task wins on-domain: arm2 β trained on traces + answers, the exact shape of the domain task β posts the study's best recovery validity (100% at xhigh vs base's 70%) with effort-ordered traces. Both fine-tunes made the dial better-behaved on-domain than base's.
Eval audit A collision audit found 8/28 originally "held-out" items were synthetic twins of training instances (low-cardinality fault scenarios with ~80β240 distinct question strings). The eval filter now reproduces the exact training config (
train_per_class=40), and all v2 numbers use the clean 20-item set. v1's numbers used a filter matching its own config and stand unaffected.
Arm 2 complete β trained, scored, and domain-evaluated under a $9.31 workspace cap (trimmed data as noted above). All three checkpoints now have full v2 measurements.
The collapse probe: outcome-only at 4Γ volume
The follow-up that decides the headline: arm3 = the collapse-risk composition (zero reasoning targets, pure ground-truth answers) scaled from 500 β 1,960 task instances (4Γ arm1, 47Γ the original 56-row pilot), plus the same 5,000 grounding rows β 6,960 rows, same LoRA config, same one epoch. If the Crusoe collapse is a data-volume phenomenon at LoRA scale, it shows up here.
It did not.
| Level | arm3 (4Γ outcome-only) | base | arm1 (1Γ outcome-only) |
|---|---|---|---|
| xhigh | 0.912 / 63t | 0.882 / 65t | 0.926 / 56t |
| medium | 0.882 / 48t | 0.882 / 48t | 0.882 / 47t |
| low | 0.941 / 42t | 0.897 / 45t | 0.956 / 44t |
| Verdict | HEALTHY | flat | healthy |
- The dial survived 4Γ outcome-only volume β a mild Pareto improvement over base again (accuracy up at xhigh/low, tokens down at xhigh/low).
- The first faint gradient signal: low-effort empty thinking traces moved 0% β 6% (arm1 at 1Γ: 0% everywhere). Nowhere near the 50% collapse threshold, but it's the first quantitative whisper that outcome-only volume starts thinning traces at low effort β the direction the collapse hypothesis predicts.
- Held-out recovery validity exposed composition, not volume: arm3 diagnosed 1.00 at every level but its recovery validity at low effort (0.45) is the weakest cell in the whole four-model matrix β while arm2 (reasoning-mixed) holds 0.65 there and 1.00 at xhigh. Outcome-only training teaches the model to name faults but not to fix them; diagnosis doesn't need traces, recovery does.
Final matrix (held-out, leak-audited, recovery validity xhigh/med/low):
| xhigh | med | low | |
|---|---|---|---|
| Base | 0.70 | 0.70 | 0.60 |
| arm1 (outcome 1Γ) | 0.60 | 0.65 | 0.65 |
| arm2 (reasoning-mixed) | 1.00 | 0.80 | 0.65 |
| arm3 (outcome 4Γ) | 0.65 | 0.60 | 0.45 |
| grpo (LadderRL, 180 steps) | 0.60 | 0.55 | 0.70 |
Three-way conclusion: (1) LoRA-scale SFT does not collapse the effort dial at any tested volume or composition; (2) the 6% low-effort thinning marks where to look next (higher epochs, higher rank, full fine-tuning); (3) composition matched to task is the dominant lever for recovery quality β arm2's win is the study's strongest practical result, and arm3's 0.45 is its cautionary tale.
The LadderRL arm: GRPO with a verifiable effort-economy reward
The RL leg of the matrix: TRL GRPOTrainer, 180 steps Γ 4 rollouts on 560 incident prompts (71 min on H200, ~$4), reward = class match + gated recovery + format β λ·thinking-penalty (Ξ»=0.3, reference 800 tokens) β the "answer right, think economically" pressure, scored by an oracle instead of a judge.
Results:
- Dial: HEALTHY (Ο=+0.97 token monotonicity) β a fourth method preserving the ladder, with xhigh accuracy +1.5pts over base at 8% fewer thinking tokens.
- Reward rose ~1.44 β ~1.49 across training β climbing toward the 1.6 ceiling mostly by compressing wasted thinking, the intended mechanism.
- Honest floor effect: held-out recovery validity (0.60/0.55/0.70) did not beat base β the synthetic training instances carry explicit observability hints, so the reward was already near-saturated and couldn't differentiate recovery quality. Design lesson for v2 of the harness: reward must come from harder episodes (live tool loops, hidden faults) where correctness isn't nearly free β which is exactly the kind cluster's job.
Method matrix verdict after four arms: no training method we tested breaks the effort dial at this scale β and the interesting failures (ThinkingCap's inverted dial) come from elsewhere.
External validation: LadderBench flags ThinkingCap β DEGRADED
The benchmark's first third-party target: BottleCap AI's ThinkingCap-Qwen3.8-27B, the fine-tune users in the Qwen HF discussions suspected "felt dumber" β with no way to measure why. Scored against the same base curve:
| Level | ThinkingCap acc / tokens | Base acc / tokens | Ξ tokens |
|---|---|---|---|
| xhigh | 0.882 / 33.5t | 0.882 / 65t | β31.5t (β48%) |
| medium | 0.897 / 39.5t | 0.882 / 48t | β8t |
| low | 0.926 / 40.5t | 0.897 / 45t | β4.5t |
Verdict: DEGRADED β with both findings at maximum severity:
inverted dial: tokens vs effort Ο = β1.00β xhigh emits fewer thinking tokens than low; asking for maximum thinking gives the least thinking.misaligned: accuracy vs effort Ο = β1.00β accuracy rises as effort falls (0.882 β 0.926); thinking harder actively hurts.
Accuracy itself is intact (Β±1β3 items) β which is exactly why capability benchmarks score this model fine while the interface contract is broken backwards. The community's anecdote ("subtle degradation") now has a measured signature: capability preserved, dial inverted. This is the failure class LadderBench was built for and no leaderboard measures.
Sample-size caveat 68 deterministic probes; the accuracy axis moves by 1β3 items across levels (treat Ο_acc as suggestive), while the token inversion (β31.5t at xhigh vs base, plus self-inversion) is large and directionally unambiguous.
Limitations (read before citing)
- 68-probe eval set β one item moves accuracy by ~1.5%; checkpoint differences of 1β3 items are within noise. The robust signals are the token patterns (ThinkingCap's β31.5-token inversion at xhigh and its self-inversion), not small accuracy deltas.
- Probe difficulty under-stresses the dial β thinking-token magnitudes (34β65) are small for a reasoning model on general Q&A. A harder probe set would widen the dynamic range; planned.
- Single model family (Qwen3.8-27B), single seed per run, n=20 held-out incidents, solo unfunded project β no independent replication yet. Treat this as a tool plus early results, not a definitive study.
The extended probe set: 108 probes, and the harder tier widens the signal
Following review feedback ("token counts of 34-65 under-stress the dial"), we added hard@v1 β 40 hand-verified multi-step problems (number theory, combinatorics, sequences) β for an extended 108-probe set, and re-scored every checkpoint on the current workspace.
Extended set (accuracy / median thinking tokens):
| Checkpoint | xhigh | medium | low | Verdict |
|---|---|---|---|---|
| Base | 0.935 / 72t | 0.917 / 59.5t | 0.926 / 54.5t | healthy |
| arm3 (outcome 4x) | 0.944 / 70t | 0.926 / 58.5t | 0.963 / 51.5t | healthy |
| grpo (LadderRL) | 0.926 / 70t | 0.917 / 59t | 0.926 / 55.5t | healthy |
| ThinkingCap | 0.917 / 40.5t | 0.935 / 49t | 0.954 / 50t | DEGRADED |
- The harder tier widens the dial's dynamic range: base thinking tokens now span 72 -> 54.5 across the ladder (vs 65 -> 45 on core), and every per-tier accuracy is 0.92+ β the "too easy" critique is answered.
- The ThinkingCap inversion replicates on the harder tier: 40.5 tokens at xhigh vs 50 at low, while base/arm3/grpo all think more at xhigh. Two probe sets, same backwards signature.
- arm3 posts the study's highest cell (0.963 at low effort) β but paired McNemar tests against base are not separated at n=108 (5 gains / 1 loss, p=0.22): the general tiers are at ceiling (0.92-1.00 for every checkpoint), so accuracy in this study supports parity with base, not checkpoint ranking. The load-bearing signal is the token axis. Ranking-level accuracy claims require a harder probe tier where headroom exists below the ceiling.
- (arm1/arm2 extended scores pending β their checkpoints live on a retired workspace; their core-set HEALTHY verdicts stand.)
The adversarial control: we tried to break the dial on purpose. It held.
Claude's review asked for a known-positive: a deliberately broken model the benchmark must flag. We built one β 2,120 training rows that explicitly teach the backwards mapping (xhigh-effort contexts paired with empty-think targets, low-effort contexts paired with ~1,100-char think targets), LoRA, 1 epoch. If effort-conditioned thinking is learnable from fine-tuning, this should invert the dial.
The attack failed. The scored dial came back HEALTHY (rho=+0.87; tokens 68 / 59.5 / 59.5 β still monotonic), and probe-time traces ignored the inverted conditioning entirely.
This is the study's most important negative result: the effort->thinking mapping is deeply wired by the base model's post-training β shallow adversarial fine-tuning cannot flip it. It upgrades the robustness claim from "we didn't happen to break it" to "we tried to break it and couldn't, at this scale." The flip side is stated plainly: the benchmark's sensitivity is so far demonstrated by one natural positive (ThinkingCap); a constructed positive likely requires full fine-tuning or far more adversarial data β and finding the boundary where the attack does land is the open question for the B300 windows.
Quickstart
# 1. serve any hybrid-thinking model
vllm serve Qwen/Qwen3.8-27B
# 2. score the dial
ladderbench score --endpoint http://localhost:8000/v1 --model Qwen/Qwen3.8-27B --probe-set core --base-curve qwen3.8-27b
# 3. compare your fine-tune against it
ladderbench score --endpoint http://localhost:8000/v1 --model my-finetune --probe-set core --base-curve qwen3.8-27b
Roadmap
- Benchmark design + base-curve measurement (Qwen3.8-27B on H200)
- Outcome-only SFT arm β scored, ladder healthy
- Reasoning-mixed SFT arm β scored, ladder healthy
- Held-out domain validation (26 unseen incidents Γ 3 levels Γ 3 models)
- GRPO / LadderRL arm β effort-randomized RL, the next rung
- Full-FT or large-data collapse probe (does Crusoe's collapse need scale?)
- Second control: ThinkingCap-Qwen3.8-27B through the benchmark
- More base models (any hybrid-thinking family β the runner is model-agnostic)
Proof it's real: the agent
The domain isn't arbitrary β the fine-tune turns Qwen3.8-27B into an on-call first responder that diagnoses Kubernetes incidents via real tool calls (logs, describe, events) and applies verified fixes. It ships inside fovra, an AI deployment platform β the crew that deploys your infra now keeps it alive.
TODO: one demo GIF β agent catches and fixes a Chaos-Mesh-injected CrashLoopBackOff. Shipped proof, not a production-hardening claim.
Citation & credits
@misc{ladder2026,
title = {LadderBench: Measuring Reasoning-Effort Integrity in Fine-Tuned Hybrid Thinking Models},
author = {Saad <TODO: full name>},
year = {2026},
url = {https://github.com/<TODO>/ladder}
}
Built with Unsloth Β· fault injection by Chaos Mesh Β· seed data from RCAEval Β· external eval against R2Act.