AI & ML interests

Frontier alignment research to ensure the safe development and deployment of advanced AI systems.

Recent Activity

skar0  updated a dataset 13 minutes ago
AlignmentResearch/impossible-swegym
skar0  published a dataset 5 days ago
AlignmentResearch/impossible-swegym
View all activity

AlignmentResearch 's collections 4

The Obfuscation Atlas
Obfuscated Policy, Obfuscated Activations, Blatant Deception, and Honest models trained in the Obfuscation Atlas paper.
The Obfuscation Altas
Obfuscated Policy, Obfuscated Activations, Blatant Deception, and Honest models trained in the Obfuscation Atlas paper