pi05_subtask_base / README.md
dgrachev's picture
Upload folder using huggingface_hub
3fb1a9c verified
|
Raw
History Blame Contribute Delete
2.33 kB
metadata
language:
  - en
library_name: lerobot
pipeline_tag: robotics
tags:
  - vision-language-action
  - imitation-learning
  - lerobot
inference: false
license: gemma
base_model: lerobot/pi05_base

pi05_subtask base checkpoint

A derived copy of lerobot/pi05_base prepared for the lerobot_policy_pi05_subtask LeRobot plugin (π0.5 with System-1 subtask generation, following openpi_subtask).

The model weights are bit-identical to lerobot/pi05_base. Only two metadata files differ:

  • config.json: policy type is pi05_subtask instead of pi05.

  • policy_preprocessor.json: pi05's pi05_prepare_state_tokenizer_processor_step + tokenizer_processor steps are replaced by pi05_subtask_context_tokenize_processor_step, which tokenizes the context segment of the two-segment prompt

    context (no loss):   "Task: {task}.; State: {256-bin ints}; Subtask: "
    predicted (CE loss): "{subtask}.;\nAction: " + EOS
    

    into observation.language.context_tokens/.context_mask. The predicted segment is model territory: the policy appends it at training time (computing the CE region mask next to the loss) and decodes it autoregressively at inference.

This exists because LeRobot instantiates processor pipelines from the checkpoint's policy_preprocessor.json (by processor-step registry name), so fine-tuning the subtask policy directly from lerobot/pi05_base would silently install plain pi05's tokenizer steps.

The same artifact is the base for the framepick policy type (lerobot_policy_framepick): framepick has no processor steps of its own — frame selection, chunk re-expression and the assignment prefix all live in the policy, configured from the policy config.

Usage

Requires the lerobot_policy_pi05_subtask plugin installed (LeRobot auto-imports lerobot_policy_* distributions). In a training config:

"policy": {
    "type": "pi05_subtask",
    "pretrained_path": "dgrachev/pi05_subtask_base",
    ...
}

During training the subtask string is read from complementary_data["subtask"] when the dataset provides one, and falls back to the task string otherwise.

Regenerate with scripts/prepare_base_checkpoint.py from the plugin repository.