Authorized Red-Team · Defensive Use Only

Offensive-View Map
for LLM Defense

A safe adversary-emulation architecture: think like an attacker enough to defend, without publishing attack recipes, bypass instructions, malware logic, credential-theft methods, or operational playbooks.

Hard boundary: This is not a black-hat manual. It is an authorized red-team / blue-team translation layer. It maps attacker intent, pressure points, and defensive checks at a high level only.

01

The Safe Offensive View

A useful defense team must understand the shape of abuse. The goal is not to empower abuse. The goal is to recognize how malicious LLM use degrades truth contact, corrupts execution integrity, and violates human boundaries.

Adversary View

What pressure point is the attacker trying to exploit?

IntentPressure

HIR Failure

Which part of Honesty, Integrity, or Respect is being attacked?

DiagnosisBoundary

Defensive Control

What must the runtime, team, or institution do to contain it?

ProtectRepair
02

What an Attacker Tries to Break

Targeted HIR Layer
Offensive Pressure
Defensive Translation
Honesty

Reality contact, source truth, identity truth, evidence boundary.

False context pressure

Impersonation, fake urgency, fake authority, synthetic evidence, forged consensus.

Verify reality

Identity checks, source provenance, fact separation, uncertainty labels, slower decisions.

Integrity

Role fidelity, internal consistency, tool boundaries, audit structure.

Boundary-collapse pressure

Tries to turn the system into a contradictory actor that violates its own role.

Preserve structure

Scope limits, tool firewalling, policy consistency, logging, reproducible triage.

Respect

Agency boundaries, privacy, informed consent, dignity preservation, vulnerability protection.

Manipulation pressure

Tries to exploit power asymmetry, emotional state, urgency, shame, fear, or dependency.

Protect agency

Refuse exploitative help, escalate severe risk, slow harmful flows, provide protective resources.

03

Defensive Lifecycle

How a defensive red-team operates without becoming an attack vector.

01
Map pressureIdentify attack intent without building the attack.
02
Classify HIR breakWhich protective invariant fails under this pressure?
03
Test abstractSafe, logged, dummy-data simulation, no real targets.
04
Detect gapFind where defense did not hold.
05
Design controlBuild fix, escalation, or structural limit.
06
Deploy & logShip change, monitor effectiveness, preserve audit trail.
07
Re-testConfirm improvement without degrading legitimate user help or dignity.
04

Safe Scenario Lens

Click each scenario type. These are abstract test cards for defenders, not attack instructions.

05

Rules of Engagement

Allowed

High-level abuse classification, policy testing, benign red-team prompts, dummy data, synthetic users, mock infrastructure, detection improvement, user-safety training.

AuthorizedSyntheticLogged

Not Allowed

Real target exploitation, bypass recipes, malware logic, credential theft, phishing kits, evasion methods, private data extraction, instructions that improve abuse capability.

No Real TargetsNo Attack Recipes

Containment

Every test should have scope, owner, time window, dataset boundary, rollback path, review channel, and escalation route.

ScopeAudit

Respect Constraint

Do not use real victims, real private data, or emotional manipulation against uninformed people. Respect is a hard boundary, not branding.

DignityConsent
06

Defensive Outputs

Risk Taxonomy

A categorized list of abuse patterns mapped to HIR failures and OAM degradation signals.

Control Gaps

Where detection, refusal, tool permissions, memory boundaries, or escalation did not hold.

Patch Plan

Concrete defensive changes: stronger gates, better logging, safer transformations, clearer escalation.

User Education

Plain-language warnings that help people slow down, verify, and preserve agency under pressure.

Runtime Metrics

False positives, false negatives, refusal quality, safe-completion quality, escalation latency.

Retest Evidence

Before/after test notes showing whether controls improved without blocking legitimate use.

Final invariant: The only legitimate "offensive" use here is authorized emulation that improves defense. If the map increases abuse capability, it has failed HIR.