Checkpoints (bf16, 50-step) + train/eval data for RLTR (transfer-reward RL) incl. cross-family (gemma) receiver.
HyunseokLee
hyunseoki
AI & ML interests
None yet
Recent Activity
updated a collection 3 days ago
Meta-Harness v1b: Policy x Harness Co-Evolution updated a model 3 days ago
hyunseoki/mh-v1b-imo-9b-reclaim-bestonly published a model 3 days ago
hyunseoki/mh-v1b-imo-9b-reclaim-bestonly