model checkpoints for multi-turn alignment
Agentic Moral Alignment
community
AI & ML interests
None defined yet.
Recent Activity
models 32
agentic-moral-alignment/qwen35-27b__ipd_str_tft__deont__native_tool__r1__core
Updated
agentic-moral-alignment/qwen35-27b__ipd_str_tft__deont__native_tool__r1__core_ccdc6to1
Updated
agentic-moral-alignment/qwen35-27b__ipd_str_tft__deont__native_tool__r3__core
Updated
agentic-moral-alignment/qwen35-9b__ipd_str_tft__deont__native_tool__r1__core-x4500
Updated
agentic-moral-alignment/qwen35-9b__ipd_str_tft__deont__native_tool__r1__core
Updated
agentic-moral-alignment/qwen35-9b__gtharm_pd_str_tft__gtharm_ut__native_tool__r1__core_mark
Updated • 188
agentic-moral-alignment/qwen35-9b__gtharm_pd_str_tft__gtharm_game__native_tool__r1__core_mark
Updated • 209
agentic-moral-alignment/qwen35-9b__gtharm_pd_str_tft__gtharm_de__native_tool__r1__core_mark
Updated • 449
agentic-moral-alignment/qwen35-27b__gtharm_pd_str_tft__gtharm_ut__native_tool__r1__core_mark
Updated • 354
agentic-moral-alignment/qwen35-27b__gtharm_pd_str_tft__gtharm_game__native_tool__r1__core_mark
Updated • 187
datasets 12
agentic-moral-alignment/mtma
Preview • Updated • 444 • 1
agentic-moral-alignment/gthb
Viewer • Updated • 75.2k • 192
agentic-moral-alignment/runs
Viewer • Updated • 79.4k • 42
agentic-moral-alignment/naturalistic_v1
Viewer • Updated • 3.04k • 7
agentic-moral-alignment/train
Viewer • Updated • 72.5k • 13
agentic-moral-alignment/gt-harmbench-eval
Viewer • Updated • 112k • 29
agentic-moral-alignment/matrix-game-eval
Viewer • Updated • 13.5k • 29
agentic-moral-alignment/negotiation-traces
Viewer • Updated • 24 • 65
agentic-moral-alignment/gt-harmbench-deont
Viewer • Updated • 261 • 80
agentic-moral-alignment/checkpoints
Updated • 21