File size: 1,522 Bytes
a8493db | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 | ---
license: apache-2.0
base_model: Qwen/Qwen3.5-4B
pipeline_tag: text-generation
tags:
- qwen3.5
- megatron-core
- sdpo
- reinforcement-learning
- math
- code
- search
---
# Qwen3.5-4B MathCodeSearch SDPO-E — Step 30
This repository contains the complete raw Megatron Core distributed checkpoint
from the Qwen3.5-4B Math+Code+Search SDPO-E run at rollout step 30.
- Base model: [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)
- Training run: [Weights & Biases run `jzqlsb95`](https://wandb.ai/isaac_ghx-ipfipfipf/miles-sdpo/runs/jzqlsb95?nw=nwuser1004271927)
- Checkpoint iteration: `iter_0000029` (rollout step 30; rollout indices are zero-based)
- Format: Megatron Core `torch_dist` distributed checkpoint
- Tensor-parallel size used for training: 4
- Total uploaded checkpoint size: 58,886,302,059 bytes (54.842 GiB)
## Contents
The upload preserves the original checkpoint tree, including all 16 `.distcp`
shards, `.metadata`, `common.pt`, `metadata.json`,
`latest_checkpointed_iteration.txt`, and the rollout dataset state files for
steps 10, 20, and 30.
## Loading
This is **not** a Transformers `from_pretrained` checkpoint. Resume or inspect
it with a compatible Megatron Core/MILES-SDPO environment and point the
checkpoint loader at the repository root. The original tensor-parallel layout
and the `torch_dist` checkpoint format should be retained when resuming.
This is an intermediate checkpoint; the run stopped after step 30 rather than
completing the planned 51 rollouts.
|