Leo Park
leopark42
ยท
AI & ML interests
AI alignment, jailbreak detection, red teaming, model robustness, safety evaluation
Recent Activity
upvoted a paper about 17 hours ago
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures upvoted a paper 5 days ago
On the Diffusibility of High-Dimensional LatentsOrganizations
None yet