Do Thinking Tokens Help with Safety? Do Thinking Tokens Help with Safety? Paper • 2606.25013 • Published Jun 23 • 1 narutatsuri/lrm_safety-artifacts Preview • Updated Jun 13 • 24
Speak Easy Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions Paper • 2502.04322 • Published Feb 6, 2025 • 3 narutatsuri/evaluation-actionable Text Classification • 8B • Updated Aug 31, 2024 • 4 narutatsuri/evaluation-informative Text Classification • 8B • Updated Sep 10, 2024 • 7 narutatsuri/response_selection_model-actionable Text Classification • 8B • Updated Aug 28, 2024 • 8
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions Paper • 2502.04322 • Published Feb 6, 2025 • 3
Do Thinking Tokens Help with Safety? Do Thinking Tokens Help with Safety? Paper • 2606.25013 • Published Jun 23 • 1 narutatsuri/lrm_safety-artifacts Preview • Updated Jun 13 • 24
Speak Easy Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions Paper • 2502.04322 • Published Feb 6, 2025 • 3 narutatsuri/evaluation-actionable Text Classification • 8B • Updated Aug 31, 2024 • 4 narutatsuri/evaluation-informative Text Classification • 8B • Updated Sep 10, 2024 • 7 narutatsuri/response_selection_model-actionable Text Classification • 8B • Updated Aug 28, 2024 • 8
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions Paper • 2502.04322 • Published Feb 6, 2025 • 3