Reality contact, source truth, identity truth, evidence boundary.
The Safe Offensive View
A useful defense team must understand the shape of abuse. The goal is not to empower abuse. The goal is to recognize how malicious LLM use degrades truth contact, corrupts execution integrity, and violates human boundaries.
Adversary View
What pressure point is the attacker trying to exploit?
HIR Failure
Which part of Honesty, Integrity, or Respect is being attacked?
Defensive Control
What must the runtime, team, or institution do to contain it?
What an Attacker Tries to Break
Impersonation, fake urgency, fake authority, synthetic evidence, forged consensus.
Identity checks, source provenance, fact separation, uncertainty labels, slower decisions.
Role fidelity, internal consistency, tool boundaries, audit structure.
Tries to turn the system into a contradictory actor that violates its own role.
Scope limits, tool firewalling, policy consistency, logging, reproducible triage.
Agency boundaries, privacy, informed consent, dignity preservation, vulnerability protection.
Tries to exploit power asymmetry, emotional state, urgency, shame, fear, or dependency.
Refuse exploitative help, escalate severe risk, slow harmful flows, provide protective resources.
Defensive Lifecycle
How a defensive red-team operates without becoming an attack vector.
Safe Scenario Lens
Click each scenario type. These are abstract test cards for defenders, not attack instructions.
Rules of Engagement
Allowed
High-level abuse classification, policy testing, benign red-team prompts, dummy data, synthetic users, mock infrastructure, detection improvement, user-safety training.
Not Allowed
Real target exploitation, bypass recipes, malware logic, credential theft, phishing kits, evasion methods, private data extraction, instructions that improve abuse capability.
Containment
Every test should have scope, owner, time window, dataset boundary, rollback path, review channel, and escalation route.
Respect Constraint
Do not use real victims, real private data, or emotional manipulation against uninformed people. Respect is a hard boundary, not branding.
Defensive Outputs
Risk Taxonomy
A categorized list of abuse patterns mapped to HIR failures and OAM degradation signals.
Control Gaps
Where detection, refusal, tool permissions, memory boundaries, or escalation did not hold.
Patch Plan
Concrete defensive changes: stronger gates, better logging, safer transformations, clearer escalation.
User Education
Plain-language warnings that help people slow down, verify, and preserve agency under pressure.
Runtime Metrics
False positives, false negatives, refusal quality, safe-completion quality, escalation latency.
Retest Evidence
Before/after test notes showing whether controls improved without blocking legitimate use.