Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published 6 days ago • 4
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning Paper • 2605.16301 • Published Jun 3 • 1
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published 6 days ago • 4
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models Paper • 2606.18142 • Published Jun 17 • 2
Backup-CaML/Basellama_plus3krealanimalsconvos2feb22_plus20kfinetune_GRPO Text Generation • 8B • Updated Feb 24 • 4
Backup-CaML/Basellama_plus3kv3_plus20kfinetune2epochs_plus1kGRPO Text Generation • 8B • Updated Feb 24 • 4
Backup-CaML/Basellama_plus3kurbandensity_plus20kfinetune_plus1kGRPO Text Generation • 8B • Updated Feb 24 • 5
Backup-CaML/Basellama_plus3kv3new_plus20kfinetune_plus1kGRPO Text Generation • 8B • Updated Feb 24 • 3