Release initial Double-DQN consolidation checkpoint
Browse filesAdd the initial 71-128-64-1 consolidation policy, configuration, model card, and SHA256 checksums.
- README.md +70 -0
- SHA256SUMS +3 -0
- config.json +97 -0
- consolidation_policy.pt +3 -0
README.md
ADDED
|
@@ -0,0 +1,70 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
tags:
|
| 3 |
+
- robotics
|
| 4 |
+
- humanoid
|
| 5 |
+
- reinforcement-learning
|
| 6 |
+
- double-dqn
|
| 7 |
+
- pytorch
|
| 8 |
+
---
|
| 9 |
+
|
| 10 |
+
# HumanoidTTT — Consolidation Policy
|
| 11 |
+
|
| 12 |
+
Initial Double-DQN consolidation weights for
|
| 13 |
+
[HumaniodTTT: Test-Time Capability Reuse for Efficient Humanoid Control](https://github.com/AIGeeksGroup/HumaniodTTT).
|
| 14 |
+
|
| 15 |
+
## Model
|
| 16 |
+
|
| 17 |
+
The policy scores retention actions for a finite-capacity capability store. Each
|
| 18 |
+
action has a 71-dimensional consolidation feature and is scored by a shared
|
| 19 |
+
`71 → 128 → 64 → 1` ReLU network with **17,537 parameters**. When the ten-slot store
|
| 20 |
+
is full, the actions are `SKIP` or `REPLACE(j)`. Online Double-DQN updates use the
|
| 21 |
+
fraction of subsequent requests with successful reuse between consecutive
|
| 22 |
+
full-store qualified-miss decisions.
|
| 23 |
+
|
| 24 |
+
## Release files
|
| 25 |
+
|
| 26 |
+
| File | Contents |
|
| 27 |
+
| --- | --- |
|
| 28 |
+
| `consolidation_policy.pt` | Initial FP32 PyTorch scorer state dictionary, seed 83001. |
|
| 29 |
+
| `config.json` | Architecture, feature layout, action mapping, online settings, and weight checksum. |
|
| 30 |
+
| `SHA256SUMS` | File checksums. |
|
| 31 |
+
|
| 32 |
+
This checkpoint is the starting point before online adaptation. It does not
|
| 33 |
+
contain a prefilled capability store or optimizer/replay state.
|
| 34 |
+
|
| 35 |
+
## Usage
|
| 36 |
+
|
| 37 |
+
Install the code and download dependencies:
|
| 38 |
+
|
| 39 |
+
```bash
|
| 40 |
+
git clone https://github.com/AIGeeksGroup/HumaniodTTT.git
|
| 41 |
+
cd HumaniodTTT
|
| 42 |
+
pip install -e .
|
| 43 |
+
pip install huggingface_hub
|
| 44 |
+
```
|
| 45 |
+
|
| 46 |
+
```python
|
| 47 |
+
import torch
|
| 48 |
+
from huggingface_hub import hf_hub_download
|
| 49 |
+
from humanoid_ttt import ConsolidationPolicy
|
| 50 |
+
|
| 51 |
+
path = hf_hub_download("AIGeeksGroup/HumanoidTTT", "consolidation_policy.pt")
|
| 52 |
+
weights = torch.load(path, map_location="cpu", weights_only=True)
|
| 53 |
+
policy = ConsolidationPolicy(weights, capacity=10, online_learning_rate=3e-5, gamma=0.95)
|
| 54 |
+
```
|
| 55 |
+
|
| 56 |
+
See the [GitHub README](https://github.com/AIGeeksGroup/HumaniodTTT#using-the-components)
|
| 57 |
+
for entry applicability, execution feedback, and update interfaces. Pin the model
|
| 58 |
+
revision and code commit when reproducing experiments.
|
| 59 |
+
|
| 60 |
+
## Dependencies and scope
|
| 61 |
+
|
| 62 |
+
The 45-dimensional A2 entry-applicability module uses state features and geometric
|
| 63 |
+
certificates; it has no separate neural-network checkpoint. Applications supply
|
| 64 |
+
motion-specific certificates, qualification evidence, and the execution runtime.
|
| 65 |
+
Download frozen generation and tracking models from
|
| 66 |
+
[OMG](https://github.com/Tsinghua-MARS-Lab/OMG) and
|
| 67 |
+
[HoloMotion](https://github.com/HorizonRobotics/HoloMotion).
|
| 68 |
+
|
| 69 |
+
The released scorer weights alone do not constitute an end-to-end robot controller.
|
| 70 |
+
Successful loading verifies network compatibility, not complete experimental reproduction.
|
SHA256SUMS
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2b9232271b1c1e64393a687104570bd04508f39c6abc96baaae4c0fd528072fc config.json
|
| 2 |
+
cdb1b5f4f9c671da5c6c26f9d9cc752aee3c81fc87b40dad67d0efa02f0e60b5 consolidation_policy.pt
|
| 3 |
+
8093b0bb8d7cfbd284b7ab89fcb89453f8842a0ed9332910d30da39da819ef93 README.md
|
config.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"format_version": 1,
|
| 3 |
+
"algorithm": "Double-DQN",
|
| 4 |
+
"checkpoint": "consolidation_policy.pt",
|
| 5 |
+
"checkpoint_sha256": "cdb1b5f4f9c671da5c6c26f9d9cc752aee3c81fc87b40dad67d0efa02f0e60b5",
|
| 6 |
+
"checkpoint_stage": "before_online_adaptation",
|
| 7 |
+
"initialization_seed": 83001,
|
| 8 |
+
"architecture": [
|
| 9 |
+
71,
|
| 10 |
+
128,
|
| 11 |
+
64,
|
| 12 |
+
1
|
| 13 |
+
],
|
| 14 |
+
"hidden_activation": "ReLU",
|
| 15 |
+
"parameter_count": 17537,
|
| 16 |
+
"dtype": "float32",
|
| 17 |
+
"state_dict_format": "raw_pytorch_state_dict",
|
| 18 |
+
"capacity": 10,
|
| 19 |
+
"recent_request_window": 16,
|
| 20 |
+
"action_count_when_full": 11,
|
| 21 |
+
"actions": {
|
| 22 |
+
"0": "SKIP",
|
| 23 |
+
"1..10": "REPLACE(entries[action_index - 1])"
|
| 24 |
+
},
|
| 25 |
+
"api_action_labels": {
|
| 26 |
+
"SKIP": "DROP",
|
| 27 |
+
"REPLACE": "REPLACE",
|
| 28 |
+
"free_slot": "INSERT"
|
| 29 |
+
},
|
| 30 |
+
"feature_layout": [
|
| 31 |
+
{
|
| 32 |
+
"slice": [
|
| 33 |
+
0,
|
| 34 |
+
16
|
| 35 |
+
],
|
| 36 |
+
"name": "candidate_motion_descriptor"
|
| 37 |
+
},
|
| 38 |
+
{
|
| 39 |
+
"slice": [
|
| 40 |
+
16,
|
| 41 |
+
32
|
| 42 |
+
],
|
| 43 |
+
"name": "store_descriptor_mean"
|
| 44 |
+
},
|
| 45 |
+
{
|
| 46 |
+
"slice": [
|
| 47 |
+
32,
|
| 48 |
+
48
|
| 49 |
+
],
|
| 50 |
+
"name": "store_descriptor_max"
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"slice": [
|
| 54 |
+
48,
|
| 55 |
+
64
|
| 56 |
+
],
|
| 57 |
+
"name": "replacement_target_descriptor_or_zeros_for_skip"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"slice": [
|
| 61 |
+
64,
|
| 62 |
+
69
|
| 63 |
+
],
|
| 64 |
+
"name": "global_features",
|
| 65 |
+
"fields": [
|
| 66 |
+
"store_size_div_20",
|
| 67 |
+
"recent_unique_count_div_16",
|
| 68 |
+
"candidate_recent_frequency",
|
| 69 |
+
"mean_candidate_store_distance_div_20",
|
| 70 |
+
"min_full_store_decisions_20_div_20"
|
| 71 |
+
]
|
| 72 |
+
},
|
| 73 |
+
{
|
| 74 |
+
"slice": [
|
| 75 |
+
69,
|
| 76 |
+
71
|
| 77 |
+
],
|
| 78 |
+
"name": "action_features",
|
| 79 |
+
"fields": [
|
| 80 |
+
"is_skip",
|
| 81 |
+
"target_recent_frequency"
|
| 82 |
+
]
|
| 83 |
+
}
|
| 84 |
+
],
|
| 85 |
+
"online_learning_rate": 3e-05,
|
| 86 |
+
"gamma": 0.95,
|
| 87 |
+
"optimizer": "Adam",
|
| 88 |
+
"loss": "smooth_l1",
|
| 89 |
+
"gradient_norm_clip": 5.0,
|
| 90 |
+
"target_sync_optimizer_steps": 50,
|
| 91 |
+
"replay_min_transitions": 8,
|
| 92 |
+
"replay_draws_per_update_call": 8,
|
| 93 |
+
"reward": "successful reuse count / request count between consecutive full-store qualified-miss decisions",
|
| 94 |
+
"reward_input": "observe_request(capability_id, reused=successful_reuse)",
|
| 95 |
+
"code_repository": "https://github.com/AIGeeksGroup/HumaniodTTT",
|
| 96 |
+
"verified_source_commit": "39f33575fff8378c0cabf4571b595646d833667c"
|
| 97 |
+
}
|
consolidation_policy.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:cdb1b5f4f9c671da5c6c26f9d9cc752aee3c81fc87b40dad67d0efa02f0e60b5
|
| 3 |
+
size 73154
|