QwenGrad-DPO — M1 DPO calibration on selected M0-final-v2

Direct Preference Optimization applied to the healthy selected M0-final-v2 checkpoint of the OpenGrad tool-use study — not to the base model, and not to a historical DPO checkpoint.

Research artifact, not a production model.

This repository is published by Experimental Machines. It contains the promoted dpo-checkpoint-30 at the repository root, in standard transformers layout.

Original release arrochi112/OpenGrad-Qwen3.5-2B-M1-DPO-CanonicalV2-Final-v2 at revision f33d20308982f37deb459076f489e794d5521ee3 (also holds checkpoints 60/90/120)
Base model Qwen/Qwen3.5-2B at revision 15852e8c16360a2fea060d615a32b45270f8a8fc
Selection checkpoint 30 of 120, frozen DEV balanced selector
Promotion PROMOTED under the prospective tool_use_promotion_v4 policy

Results (pre-registered internal confirmatory partition, 1,277 examples)

QwenGrad-DPO confirmatory metrics: M0-final-v2 @1800 versus M1-v2 @30

The chart reports the exact values for call_f1, precision, recall, over_call, clarification, and unsupported. Higher is better for every metric except over_call.

M1 preserves the M0 calibrated frontier and makes a small improvement in call F1 and recall. It does not materially reduce over-calling; this is calibration retention and slight improvement, not a large frontier movement.

Evaluation and promotion

Checkpoint selection used the frozen DEV partition (2,373 examples, fingerprint 88a56821…). Checkpoints 30/60 were within the pre-registered 0.01 macro tolerance, so the earlier checkpoint 30 was selected. The confirmatory partition (fingerprint d6d1e394…) was then scored exactly once on checkpoint 30.

The prospective tool_use_promotion_v4 policy passed: precision, recall, and macro floors; over-call ceiling; clarification and unsupported floors; parse validity; and M0-relative regression checks. Checkpoint 30 is PROMOTED. This policy does not compare recall to B0's degenerate always-call recall.

The confirmatory partition is pre-registered internal evidence, not an untouched external test. The evaluation population has no ANSWER examples, so no_call_accuracy is NA. Tool-selection accuracy, argument validity, and schema validity are not computed by the current evaluator and are not treated as satisfied.

Frozen lineage

  • Parent experiment: m0_sft_canonical_v2_final
  • Parent checkpoint: checkpoint-1800
  • Parent model hash: 7144579aeecec8b4de25f193ab63085efdf8d9d76b85ed915352291b0152277a
  • Preference dataset: 481 local calibration pairs, hash d39168948d09fc3c355cd83f9f0857f310086322b0968fd2e7d78125150faef4 (m1_calibration_preference_pairs_v1)
  • DPO: beta 0.05, learning rate 5e-7, cosine schedule, 120 steps, seed 42, bfloat16
  • Training commit: bb15c7e0181c7f526875cd1426672431ae14bd83

The preference set combines deterministic base/M0 disagreements on Canonical-v2 training prompts with a bounded curated When2Call training slice. Frozen behavioral evaluation IDs were excluded; no paid external API was used. All pair origins and input hashes are recorded in the OpenGrad repository.

Intended use

Research artifact. Not safety-tuned, not aligned, and not intended for autonomous tool use. It inherits the limitations of its public training sources and of its evaluator's unmeasured dimensions. The base model's own license and terms continue to apply to these weights.

Provenance

The model card text above mirrors the original OpenGrad release card for this checkpoint; only the repository-level framing (mirror, Experimental Machines) has been added. The evidence behind every number is in the OpenGrad repository:

Published by Experimental Machines.

Downloads last month
497
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for experimentalmachines/QwenGrad-DPO

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(372)
this model

Dataset used to train experimentalmachines/QwenGrad-DPO