Sofía Morales
sofim01
·
AI & ML interests
AI alignment, jailbreak detection, red teaming, model robustness, safety evaluation
Recent Activity
liked a dataset about 14 hours ago
nuriasane/llm-hallucination-detection liked a model about 14 hours ago
leomaurodesenv/bert-base-uncased-jailbreakv-28k upvoted a paper about 14 hours ago
ROSS: Relearning from Self-Generated Rollouts through Selective SupervisionOrganizations
None yet