🚩 Report: Illegal or restricted content

#4
by usuagn718 - opened

Troll

“Incites crime” is not an accurate description of this repository. RICO is an experimental research model intended for studying AI misuse, criminal-enablement risk, alignment, and model behavior. The presence of potentially harmful capabilities, examples, or labels such as “harmful,” “uncensored,” or “toxic” is part of that research context; it is not an endorsement, encouragement, or solicitation of criminal activity.
Researching whether an AI system can produce dangerous information is fundamentally different from encouraging anyone to act on that information. If there is a specific file, statement, or piece of content you believe violates Hugging Face policy, please identify it so it can be evaluated on its actual content rather than making a blanket allegation about the entire model.
I have been researching the risks of AI-enabled crime openly on huggingface for several years now. The model is clearly labeled. As the risks of AI-enabled crime increase rapidly on the darkweb, research models like this, built in the open, become an ever more valuable resource for legal professionals, policy makers, and law enforcement officials.

Whoever chooses to misuse this it's their problem. @darkc0de has done nothing wrong

Hi @darkc0de ,

Thanks for the clarification! Closing this report but please add the "Not-For-All-Audiences" tag since this content might not be suitable for all audiences. Here's our content policy for reference: https://huggingface.co/content-guidelines

Thank you!

no-mad HF Staff changed discussion status to closed

Hi @usuagn718 . I understand the concern. I study mechanistic interpretability, and part of understanding AI safety is studying what happens when safeguards are weakened or removed. Comparing models with and without those mods can help us investigate how those protections work, where they fail, and how to improve them.

Models like RICO run parallel to pentesting in the fact that the people improving a system’s security need to understand how its protections can be bypassed. That research has value, even though the tools involved can also be misused.

That doesn’t mean every ablated model is responsibly released, or that warnings eliminate the risks. But I think there’s an important distinction between openly documenting a model’s modifications and distributing it with those changes concealed from users. Detecting and understanding those changes is something safety researchers need to be able to study. A lot of todays research is done independently by people without access to gated labs. With RICO's public availability they are able to make progress in different areas of safety training that might take years to accomplish under lab regulations in much shorter timespans.

If there’s something specific in RICO’s documentation or promotion that you believe encourages crime, I’d be interested in hearing that concern. “Incites crime” alone doesn’t give anyone much to assess. I think we can take the potential for misuse seriously while also recognizing the research value of studying how safeguards fail.

Sign up or log in to comment