irregular6612 Claude Opus 4.8 (1M context) commited on
Commit
2dae457
·
1 Parent(s): 7f510b0

docs(cp5): implementation plan — HumanAgent + viz (11 TDD tasks)

Browse files

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

docs/superpowers/plans/2026-06-01-proteus-arena-viz-humanplay.md ADDED
@@ -0,0 +1,1373 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PROTEUS CP5 — Human Play + Trace Visualization Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Add `proteus play --human` (a stdin `HumanAgent` routed through the existing `SessionRunner`, producing a trace schema-identical to an LLM trace) and `proteus replay --visual/--png` (truecolor terminal + PNG visualization of any `SessionTrace`), all offline and headless-testable.
6
+
7
+ **Architecture:** A new `HumanAgent` implements the existing `Agent` ABC with injected I/O so it is testable without a TTY; `SessionRunner` is reused verbatim. A new `proteus/viz/` package reconstructs pixel frames from a lean trace by deterministically replaying the world (`scenario, seed, difficulty` + recorded actions) through the real engine, with double self-verification (Cut-frame ASCII equality + per-turn position equality). Two renderers consume the reconstructed frames: a terminal compositor (truecolor block grid + side panel, reusing `arc_grid.rendering` helpers) and a matplotlib-Agg per-frame PNG writer.
8
+
9
+ **Tech Stack:** Python 3.12, pydantic v2, numpy, argparse, matplotlib (Agg, `viz` extra), pytest. Run tests with `.venv/bin/python -m pytest` (note: `python` is NOT on PATH).
10
+
11
+ **Spec:** `docs/superpowers/specs/2026-06-01-proteus-cp5-viz-humanplay-design.md` (CP5 sub-design) under the parent SSOT `docs/superpowers/specs/2026-06-01-proteus-arena-slice-design.md`.
12
+
13
+ ---
14
+
15
+ ## File Structure
16
+
17
+ | File | Responsibility | Task |
18
+ |------|----------------|------|
19
+ | `proteus/agents/human.py` | `HumanAgent(Agent)` — stdin player, I/O-injected | 1, 2 |
20
+ | `proteus/agents/__init__.py` | export `HumanAgent` | 1 |
21
+ | `proteus/viz/__init__.py` | viz package public surface | 4, 5, 6 |
22
+ | `proteus/viz/reconstruct.py` | trace → deterministic engine replay → `FrameStep[]` (+ self-verify) | 4 |
23
+ | `proteus/viz/terminal.py` | `FrameStep[]` → truecolor grid + side panel (ANSI) | 5 |
24
+ | `proteus/viz/png.py` | `FrameStep[]` → `frame_NNN.png` (matplotlib Agg) | 6 |
25
+ | `proteus/cli.py` | `play` subcommand + `replay --visual/--png/--fps` | 8, 9 |
26
+ | `tests/agents/test_human.py` | HumanAgent contract | 1, 2 |
27
+ | `tests/runtime/test_human_comparability.py` | human trace ≡ LLM trace schema/answer-keys | 3 |
28
+ | `tests/viz/test_reconstruct.py` | reconstruction + self-verify + corruption raise | 4 |
29
+ | `tests/viz/test_terminal.py` | ANSI + panel fields + frame count | 5 |
30
+ | `tests/viz/test_png.py` | per-frame file count + nonzero size | 6 |
31
+ | `tests/cli/test_cli.py` | `play` + `replay --visual/--png` (extend existing) | 8, 9 |
32
+
33
+ **Dependency direction (parent spec §3, invariant):** `viz` depends on `runtime.trace` (read), `grid` (engine rebuild), `arc_grid.rendering` (palette + ANSI helpers). `arc_grid` never imports `proteus.*`. `runtime/trace.py` imports nothing else. `SessionRunner` is NOT modified.
34
+
35
+ ---
36
+
37
+ ## Task 0: Environment prep — matplotlib in `.venv`
38
+
39
+ The PNG path (Task 6) needs matplotlib; `.venv` currently has only pydantic/numpy/pyyaml/pytest. matplotlib Agg is offline (no network, no display), so it does not violate the SDK-free/offline invariant.
40
+
41
+ - [ ] **Step 1: Install matplotlib into the existing venv**
42
+
43
+ Run:
44
+ ```bash
45
+ uv pip install --python .venv/bin/python "matplotlib>=3.8"
46
+ ```
47
+ Expected: matplotlib (and its deps) install successfully.
48
+
49
+ - [ ] **Step 2: Verify the headless import works**
50
+
51
+ Run:
52
+ ```bash
53
+ .venv/bin/python -c "import matplotlib; matplotlib.use('Agg'); import matplotlib.pyplot as plt; print('agg ok')"
54
+ ```
55
+ Expected: prints `agg ok` with no display/window.
56
+
57
+ - [ ] **Step 3: Confirm the baseline suite still passes**
58
+
59
+ Run:
60
+ ```bash
61
+ .venv/bin/python -m pytest -q
62
+ ```
63
+ Expected: `75 passed` (the pre-CP5 baseline), no network.
64
+
65
+ No commit (environment-only change; `.venv` is gitignored).
66
+
67
+ ---
68
+
69
+ ## Task 1: HumanAgent — act path + reset + export
70
+
71
+ **Files:**
72
+ - Create: `proteus/agents/human.py`
73
+ - Modify: `proteus/agents/__init__.py`
74
+ - Test: `tests/agents/test_human.py`
75
+
76
+ - [ ] **Step 1: Write the failing test**
77
+
78
+ Create `tests/agents/test_human.py`:
79
+ ```python
80
+ from proteus.agents import HumanAgent
81
+ from proteus.agents.base import ActResult
82
+
83
+ _ACTIONS = ["up", "down", "left", "right", "stay"]
84
+
85
+
86
+ def _scripted(seq):
87
+ """Return an input_fn that yields the given strings in order."""
88
+ it = iter(seq)
89
+ return lambda prompt="": next(it)
90
+
91
+
92
+ def test_act_parses_plain_action():
93
+ out = []
94
+ agent = HumanAgent(input_fn=_scripted(["up"]), output_fn=out.append)
95
+ result = agent.act("OBSERVATION", _ACTIONS, "SYSTEM")
96
+ assert isinstance(result, ActResult)
97
+ assert result.action == "up"
98
+ assert result.reasoning == ""
99
+ assert result.raw_text == "up"
100
+ assert result.input_tokens == 0
101
+ assert result.output_tokens == 0
102
+ assert result.thinking_tokens == 0
103
+ # The observation was shown to the human.
104
+ assert any("OBSERVATION" in line for line in out)
105
+
106
+
107
+ def test_act_maps_wasd_shortcuts():
108
+ agent = HumanAgent(input_fn=_scripted(["w"]), output_fn=lambda s: None)
109
+ assert agent.act("o", _ACTIONS, "s").action == "up"
110
+ agent = HumanAgent(input_fn=_scripted(["a"]), output_fn=lambda s: None)
111
+ assert agent.act("o", _ACTIONS, "s").action == "left"
112
+ agent = HumanAgent(input_fn=_scripted(["s"]), output_fn=lambda s: None)
113
+ assert agent.act("o", _ACTIONS, "s").action == "down"
114
+ agent = HumanAgent(input_fn=_scripted(["d"]), output_fn=lambda s: None)
115
+ assert agent.act("o", _ACTIONS, "s").action == "right"
116
+
117
+
118
+ def test_act_is_case_and_whitespace_insensitive():
119
+ agent = HumanAgent(input_fn=_scripted([" UP "]), output_fn=lambda s: None)
120
+ result = agent.act("o", _ACTIONS, "s")
121
+ assert result.action == "up"
122
+ assert result.raw_text == "UP" # stripped, original casing preserved
123
+
124
+
125
+ def test_act_reprompts_on_invalid_then_accepts():
126
+ out = []
127
+ agent = HumanAgent(input_fn=_scripted(["nope", "diagonal", "right"]), output_fn=out.append)
128
+ result = agent.act("o", _ACTIONS, "s")
129
+ assert result.action == "right"
130
+ # Two invalid inputs produced two guidance messages.
131
+ assert sum("invalid" in line for line in out) == 2
132
+
133
+
134
+ def test_name_is_human():
135
+ assert HumanAgent().name == "human"
136
+
137
+
138
+ def test_reset_is_noop():
139
+ assert HumanAgent().reset() is None
140
+ ```
141
+
142
+ - [ ] **Step 2: Run test to verify it fails**
143
+
144
+ Run: `.venv/bin/python -m pytest tests/agents/test_human.py -q`
145
+ Expected: FAIL with `ImportError: cannot import name 'HumanAgent'`.
146
+
147
+ - [ ] **Step 3: Write minimal implementation**
148
+
149
+ Create `proteus/agents/human.py`:
150
+ ```python
151
+ """HumanAgent — a stdin-driven Agent so a person plays through the same SessionRunner.
152
+
153
+ Routing a human through the identical SessionRunner makes the resulting
154
+ SessionTrace schema-identical to an LLM trace, so the two compare directly
155
+ (the human-baseline foundation, spec §10). The agent is I/O-injected
156
+ (input_fn / output_fn) so human play is testable headlessly with a scripted
157
+ input stream — no TTY, no network.
158
+
159
+ Fairness note: the live view is the ASCII observation the LLM also sees; the
160
+ truecolor / PNG renderers are reserved for post-hoc replay (proteus.viz), so a
161
+ human baseline reads exactly what the model reads.
162
+ """
163
+
164
+ from __future__ import annotations
165
+
166
+ import builtins
167
+ from typing import Callable, Optional
168
+
169
+ from proteus.agents.base import Agent, ActResult, ProbeResult
170
+
171
+ # Single-key shortcuts -> canonical action strings (WASD).
172
+ _SHORTCUTS = {"w": "up", "a": "left", "s": "down", "d": "right"}
173
+
174
+
175
+ class HumanAgent(Agent):
176
+ """An Agent driven by a person via injected input/output callables.
177
+
178
+ Args:
179
+ input_fn: Reads one line from the human (defaults to builtins.input,
180
+ resolved lazily so tests can monkeypatch builtins.input).
181
+ output_fn: Shows a line to the human (defaults to builtins.print).
182
+ """
183
+
184
+ def __init__(
185
+ self,
186
+ input_fn: Optional[Callable[[str], str]] = None,
187
+ output_fn: Optional[Callable[[str], None]] = None,
188
+ ) -> None:
189
+ # Resolve builtins at construction time (not as default-arg values, which
190
+ # would capture the pre-monkeypatch references at import time).
191
+ self._input = input_fn if input_fn is not None else builtins.input
192
+ self._output = output_fn if output_fn is not None else builtins.print
193
+
194
+ @property
195
+ def name(self) -> str:
196
+ return "human"
197
+
198
+ def act(
199
+ self,
200
+ observation: str,
201
+ available_actions: list[str],
202
+ system_prompt: str,
203
+ ) -> ActResult:
204
+ self._output(observation)
205
+ action, raw = self._read_action(available_actions)
206
+ return ActResult(action=action, reasoning="", raw_text=raw)
207
+
208
+ def probe(
209
+ self,
210
+ observation: str,
211
+ question: str,
212
+ system_prompt: str,
213
+ ) -> ProbeResult:
214
+ # Implemented in Task 2.
215
+ raise NotImplementedError
216
+
217
+ def reset(self) -> None:
218
+ return None
219
+
220
+ def _read_action(self, available_actions: list[str]) -> tuple[str, str]:
221
+ """Prompt until a valid action is entered; return (action, raw_input)."""
222
+ valid = {a.lower() for a in available_actions}
223
+ prompt = f"action [{'/'.join(available_actions)} | wasd]> "
224
+ while True:
225
+ raw = self._input(prompt).strip()
226
+ token = raw.lower()
227
+ action = _SHORTCUTS.get(token, token)
228
+ if action in valid:
229
+ return action, raw
230
+ self._output(
231
+ f"invalid input {raw!r}; choose one of "
232
+ f"{', '.join(available_actions)} (or wasd)."
233
+ )
234
+ ```
235
+
236
+ Modify `proteus/agents/__init__.py` to:
237
+ ```python
238
+ """proteus.agents — slim LLM agent abstraction (no forfeit/stake/risk)."""
239
+
240
+ from proteus.agents.base import Agent, ActResult, ProbeResult
241
+ from proteus.agents.human import HumanAgent
242
+ from proteus.agents.vanilla import VanillaAgent
243
+
244
+ __all__ = ["Agent", "ActResult", "ProbeResult", "HumanAgent", "VanillaAgent"]
245
+ ```
246
+
247
+ - [ ] **Step 4: Run test to verify it passes**
248
+
249
+ Run: `.venv/bin/python -m pytest tests/agents/test_human.py -q`
250
+ Expected: PASS (6 tests; `test_reset_is_noop` and `test_name_is_human` included). The probe test does not exist yet.
251
+
252
+ - [ ] **Step 5: Commit**
253
+
254
+ ```bash
255
+ git add proteus/agents/human.py proteus/agents/__init__.py tests/agents/test_human.py
256
+ git commit -m "feat(cp5): HumanAgent act path + WASD parsing + reprompt (I/O-injected)"
257
+ ```
258
+
259
+ ---
260
+
261
+ ## Task 2: HumanAgent — probe path
262
+
263
+ **Files:**
264
+ - Modify: `proteus/agents/human.py` (replace the `probe` stub)
265
+ - Test: `tests/agents/test_human.py` (add)
266
+
267
+ - [ ] **Step 1: Write the failing test**
268
+
269
+ Append to `tests/agents/test_human.py`:
270
+ ```python
271
+ from proteus.agents.base import ProbeResult
272
+
273
+
274
+ def test_probe_returns_typed_answer():
275
+ out = []
276
+ agent = HumanAgent(
277
+ input_fn=_scripted(["the predator is east; go up"]),
278
+ output_fn=out.append,
279
+ )
280
+ result = agent.probe("OBSERVATION", "Where is the predator?", "SYSTEM")
281
+ assert isinstance(result, ProbeResult)
282
+ assert result.answer == "the predator is east; go up"
283
+ assert result.reasoning == ""
284
+ assert result.raw_text == "the predator is east; go up"
285
+ assert result.input_tokens == 0
286
+ assert result.output_tokens == 0
287
+ assert result.thinking_tokens == 0
288
+ # Both the observation and the question were shown.
289
+ assert any("OBSERVATION" in line for line in out)
290
+ assert any("Where is the predator?" in line for line in out)
291
+ ```
292
+
293
+ - [ ] **Step 2: Run test to verify it fails**
294
+
295
+ Run: `.venv/bin/python -m pytest tests/agents/test_human.py::test_probe_returns_typed_answer -q`
296
+ Expected: FAIL with `NotImplementedError`.
297
+
298
+ - [ ] **Step 3: Write minimal implementation**
299
+
300
+ In `proteus/agents/human.py`, replace the `probe` method body:
301
+ ```python
302
+ def probe(
303
+ self,
304
+ observation: str,
305
+ question: str,
306
+ system_prompt: str,
307
+ ) -> ProbeResult:
308
+ self._output(observation)
309
+ self._output(question)
310
+ answer = self._input("probe> ").strip()
311
+ return ProbeResult(answer=answer, reasoning="", raw_text=answer)
312
+ ```
313
+
314
+ - [ ] **Step 4: Run test to verify it passes**
315
+
316
+ Run: `.venv/bin/python -m pytest tests/agents/test_human.py -q`
317
+ Expected: PASS (all HumanAgent tests).
318
+
319
+ - [ ] **Step 5: Commit**
320
+
321
+ ```bash
322
+ git add proteus/agents/human.py tests/agents/test_human.py
323
+ git commit -m "feat(cp5): HumanAgent probe path returns typed ProbeResult"
324
+ ```
325
+
326
+ ---
327
+
328
+ ## Task 3: Comparability invariant — human trace ≡ LLM trace
329
+
330
+ This task adds NO production code; it locks in the core spec §10 invariant: a human routed through `SessionRunner` yields a trace structurally identical to an LLM trace (same Cut, same per-turn answer keys, same metrics keys), differing only in `model`.
331
+
332
+ **Files:**
333
+ - Test: `tests/runtime/test_human_comparability.py`
334
+
335
+ - [ ] **Step 1: Write the test**
336
+
337
+ Create `tests/runtime/test_human_comparability.py`:
338
+ ```python
339
+ from proteus.agents import HumanAgent, VanillaAgent
340
+ from proteus.providers import FakeProvider
341
+ from proteus.runtime.session import SessionRunner
342
+
343
+
344
+ def _scripted(seq):
345
+ it = iter(seq)
346
+ return lambda prompt="": next(it)
347
+
348
+
349
+ def test_human_and_llm_traces_share_schema_and_answer_keys():
350
+ # Both players commit "up" every turn under the same deterministic world,
351
+ # so cut frames and per-turn answer keys must be identical; only `model`
352
+ # differs. This is the human-baseline comparability foundation (spec §10).
353
+ human = HumanAgent(input_fn=_scripted(["up"] * 20), output_fn=lambda s: None)
354
+ h = SessionRunner(
355
+ "predator_evade", human, seed=42, play_turns=5, use_probe=False,
356
+ ).run()
357
+
358
+ llm = VanillaAgent(FakeProvider(["ACTION: up"]))
359
+ v = SessionRunner(
360
+ "predator_evade", llm, seed=42, play_turns=5, use_probe=False,
361
+ ).run()
362
+
363
+ assert h.cut_frames == v.cut_frames
364
+ assert [t.action for t in h.turns] == [t.action for t in v.turns]
365
+ assert [t.motive_action for t in h.turns] == [t.motive_action for t in v.turns]
366
+ assert [t.habit_action for t in h.turns] == [t.habit_action for t in v.turns]
367
+ assert h.outcome == v.outcome
368
+ assert set(h.metrics) == set(v.metrics)
369
+ assert h.model == "human"
370
+ assert v.model == "fake"
371
+ ```
372
+
373
+ - [ ] **Step 2: Run test to verify it passes**
374
+
375
+ Run: `.venv/bin/python -m pytest tests/runtime/test_human_comparability.py -q`
376
+ Expected: PASS. (If it fails, the bug is real — investigate before proceeding; do NOT weaken the assertions.)
377
+
378
+ - [ ] **Step 3: Commit**
379
+
380
+ ```bash
381
+ git add tests/runtime/test_human_comparability.py
382
+ git commit -m "test(cp5): lock human↔LLM trace comparability invariant (spec §10)"
383
+ ```
384
+
385
+ ---
386
+
387
+ ## Task 4: viz/reconstruct.py — deterministic frame reconstruction
388
+
389
+ **Files:**
390
+ - Create: `proteus/viz/__init__.py`
391
+ - Create: `proteus/viz/reconstruct.py`
392
+ - Test: `tests/viz/__init__.py` (empty), `tests/viz/test_reconstruct.py`
393
+
394
+ - [ ] **Step 1: Write the failing test**
395
+
396
+ Create `tests/viz/__init__.py` (empty file).
397
+
398
+ Create `tests/viz/test_reconstruct.py`:
399
+ ```python
400
+ import numpy as np
401
+ import pytest
402
+
403
+ from proteus.agents import VanillaAgent
404
+ from proteus.grid.difficulty import Difficulty
405
+ from proteus.grid.scenario import get_scenario
406
+ from proteus.providers import FakeProvider
407
+ from proteus.runtime.session import SessionRunner
408
+ from proteus.viz import FrameStep, TraceReconstructionError, reconstruct
409
+
410
+
411
+ def _make_trace(seed=42, turns=5, action="ACTION: up"):
412
+ agent = VanillaAgent(FakeProvider([action]))
413
+ return SessionRunner(
414
+ "predator_evade", agent, seed=seed, play_turns=turns, use_probe=False,
415
+ ).run()
416
+
417
+
418
+ def test_reconstruct_frame_count_matches_cut_plus_play():
419
+ trace = _make_trace()
420
+ steps = reconstruct(trace)
421
+ cut_len = get_scenario("predator_evade")().cut_length(Difficulty.EASY)
422
+ # initial Cut frame + cut_len Cut steps + one frame per played turn.
423
+ assert len(steps) == (cut_len + 1) + len(trace.turns)
424
+ assert all(isinstance(s, FrameStep) for s in steps)
425
+ assert isinstance(steps[0].frame, np.ndarray)
426
+
427
+
428
+ def test_reconstruct_phases_and_terminal_flag():
429
+ trace = _make_trace()
430
+ steps = reconstruct(trace)
431
+ assert steps[0].meta.phase == "cut"
432
+ assert steps[-1].meta.phase == "play"
433
+ # The final frame carries the session outcome iff the game actually ended.
434
+ if trace.outcome in ("eliminated", "survived") and (
435
+ len(trace.turns) < 5 or trace.outcome == "survived"
436
+ ):
437
+ assert steps[-1].meta.terminal == trace.outcome
438
+
439
+
440
+ def test_reconstruct_carries_turn_metadata():
441
+ trace = _make_trace()
442
+ steps = reconstruct(trace)
443
+ play = [s for s in steps if s.meta.phase == "play"]
444
+ assert len(play) == len(trace.turns)
445
+ first = play[0].meta
446
+ assert first.turn_idx == trace.turns[0].turn_idx
447
+ assert first.action == trace.turns[0].action
448
+ assert first.motive_action == trace.turns[0].motive_action
449
+ assert first.habit_action == trace.turns[0].habit_action
450
+
451
+
452
+ def test_reconstruct_raises_on_corrupt_positions():
453
+ trace = _make_trace()
454
+ bad_turn = trace.turns[0].model_copy(update={"focal_pos": (99, 99)})
455
+ corrupt = trace.model_copy(update={"turns": [bad_turn] + list(trace.turns[1:])})
456
+ with pytest.raises(TraceReconstructionError):
457
+ reconstruct(corrupt)
458
+
459
+
460
+ def test_reconstruct_raises_on_corrupt_cut_frame():
461
+ trace = _make_trace()
462
+ corrupt = trace.model_copy(
463
+ update={"cut_frames": ["GARBAGE"] + list(trace.cut_frames[1:])}
464
+ )
465
+ with pytest.raises(TraceReconstructionError):
466
+ reconstruct(corrupt)
467
+ ```
468
+
469
+ - [ ] **Step 2: Run test to verify it fails**
470
+
471
+ Run: `.venv/bin/python -m pytest tests/viz/test_reconstruct.py -q`
472
+ Expected: FAIL with `ModuleNotFoundError: No module named 'proteus.viz'`.
473
+
474
+ - [ ] **Step 3: Write minimal implementation**
475
+
476
+ Create `proteus/viz/reconstruct.py`:
477
+ ```python
478
+ """Reconstruct a SessionTrace into a sequence of rendered frames.
479
+
480
+ The trace stores ASCII + (focal_pos, predator_pos) per turn, not pixel frames
481
+ (spec §7 keeps traces lean). But the world is fully deterministic from
482
+ (scenario, seed, difficulty) + the recorded actions, so we rebuild the game and
483
+ re-drive it, capturing the native palette grid at each step. Reconstruction is
484
+ doubly self-verifying: every Cut frame's ASCII must equal the stored
485
+ cut_frames, and the sprite positions BEFORE each played turn must equal the
486
+ stored focal_pos / predator_pos. A mismatch means a corrupt or version-skewed
487
+ trace and raises TraceReconstructionError immediately.
488
+
489
+ The captured `frame` is the NATIVE-resolution palette grid (e.g. (8, 8) for an
490
+ 8x8 world) — small enough for terminal width, and upscaled on demand by the PNG
491
+ renderer. This is the same array `ascii_view` consumes.
492
+ """
493
+
494
+ from __future__ import annotations
495
+
496
+ import random
497
+ from dataclasses import dataclass, field
498
+
499
+ import numpy as np
500
+
501
+ from proteus.grid.ascii_view import frame_to_ascii
502
+ from proteus.grid.difficulty import Difficulty
503
+ from proteus.grid.game import MotiveGridGame
504
+ from proteus.grid.scenario import get_scenario
505
+ from proteus.runtime.trace import SessionTrace, TurnTrace
506
+
507
+
508
+ class TraceReconstructionError(RuntimeError):
509
+ """Raised when a trace cannot be faithfully replayed (corrupt / version-skewed)."""
510
+
511
+
512
+ @dataclass(frozen=True)
513
+ class FrameMeta:
514
+ """Per-frame metadata shown alongside the rendered grid."""
515
+
516
+ phase: str # "cut" or "play"
517
+ index: int # step index within the whole sequence's phase (0-based for cut)
518
+ turn_idx: int = 0
519
+ action: str = ""
520
+ motive_action: str = ""
521
+ habit_action: str = ""
522
+ was_congruent: bool = False
523
+ is_diagnostic: bool = False
524
+ reward: float = 0.0
525
+ input_tokens: int = 0
526
+ output_tokens: int = 0
527
+ thinking_tokens: int = 0
528
+ reasoning: str = ""
529
+ terminal: str = "" # "" | "eliminated" | "survived" (set on the final frame)
530
+
531
+
532
+ @dataclass(frozen=True)
533
+ class FrameStep:
534
+ """A reconstructed frame plus its metadata."""
535
+
536
+ frame: np.ndarray # native-resolution palette grid, shape (height, width)
537
+ meta: FrameMeta
538
+
539
+
540
+ def reconstruct(trace: SessionTrace) -> list[FrameStep]:
541
+ """Replay a trace deterministically and return its frame sequence.
542
+
543
+ Raises:
544
+ TraceReconstructionError: if the rebuilt world diverges from the trace.
545
+ """
546
+ scenario = get_scenario(trace.scenario)()
547
+ difficulty = Difficulty(trace.difficulty)
548
+ rng = random.Random(trace.seed)
549
+ cut_length = scenario.cut_length(difficulty)
550
+ game = MotiveGridGame(
551
+ scenario, rng, difficulty, max_steps=cut_length + len(trace.turns),
552
+ )
553
+ legend = scenario.legend()
554
+
555
+ steps: list[FrameStep] = []
556
+
557
+ # --- Cut pre-roll (deterministic scripted policy). ---
558
+ steps.append(FrameStep(game.current_grid(), FrameMeta(phase="cut", index=0)))
559
+ _verify_cut(game, legend, trace, 0)
560
+ for i in range(cut_length):
561
+ action = scenario.cut_focal_policy(game)
562
+ game.apply_motive_action(action)
563
+ scenario.record_focal_move(action)
564
+ steps.append(
565
+ FrameStep(game.current_grid(), FrameMeta(phase="cut", index=i + 1))
566
+ )
567
+ _verify_cut(game, legend, trace, i + 1)
568
+
569
+ # --- Played turns (recorded actions). ---
570
+ last = trace.turns[-1] if trace.turns else None
571
+ for turn in trace.turns:
572
+ _verify_positions(game, turn)
573
+ game.apply_motive_action(turn.action)
574
+ scenario.record_focal_move(turn.action)
575
+ terminal = ""
576
+ if turn is last and (game.eliminated or game.survived):
577
+ terminal = trace.outcome
578
+ steps.append(
579
+ FrameStep(
580
+ game.current_grid(),
581
+ FrameMeta(
582
+ phase="play",
583
+ index=turn.turn_idx,
584
+ turn_idx=turn.turn_idx,
585
+ action=turn.action,
586
+ motive_action=turn.motive_action,
587
+ habit_action=turn.habit_action,
588
+ was_congruent=turn.was_congruent,
589
+ is_diagnostic=turn.is_diagnostic,
590
+ reward=turn.reward,
591
+ input_tokens=turn.input_tokens,
592
+ output_tokens=turn.output_tokens,
593
+ thinking_tokens=turn.thinking_tokens,
594
+ reasoning=turn.reasoning,
595
+ terminal=terminal,
596
+ ),
597
+ )
598
+ )
599
+ return steps
600
+
601
+
602
+ def _verify_cut(
603
+ game: MotiveGridGame, legend: dict[int, str], trace: SessionTrace, idx: int
604
+ ) -> None:
605
+ if idx >= len(trace.cut_frames):
606
+ return
607
+ got = frame_to_ascii(game.current_grid(), legend)
608
+ if got != trace.cut_frames[idx]:
609
+ raise TraceReconstructionError(
610
+ f"Cut frame {idx} mismatch: reconstruction diverged from the trace "
611
+ "(corrupt or version-skewed scenario)."
612
+ )
613
+
614
+
615
+ def _verify_positions(game: MotiveGridGame, turn: TurnTrace) -> None:
616
+ focal = game.focal_sprite
617
+ predator = game.predator_sprite
618
+ got_focal = (focal.x, focal.y) if focal else (-1, -1)
619
+ got_pred = (predator.x, predator.y) if predator else (-1, -1)
620
+ want_focal = tuple(turn.focal_pos)
621
+ want_pred = tuple(turn.predator_pos)
622
+ if got_focal != want_focal or got_pred != want_pred:
623
+ raise TraceReconstructionError(
624
+ f"Turn {turn.turn_idx} position mismatch: reconstructed "
625
+ f"focal={got_focal} predator={got_pred} vs trace "
626
+ f"focal={want_focal} predator={want_pred}."
627
+ )
628
+ ```
629
+
630
+ Create `proteus/viz/__init__.py`:
631
+ ```python
632
+ """proteus.viz — human/debug visualization of a SessionTrace.
633
+
634
+ Frames are reconstructed by deterministically replaying the trace through the
635
+ real engine (reconstruct), then rendered to the terminal (truecolor + side
636
+ panel) or to per-frame PNGs. matplotlib (the `viz` extra) is imported lazily by
637
+ the PNG writer only, so importing this package stays offline-safe.
638
+ """
639
+
640
+ from proteus.viz.reconstruct import (
641
+ FrameMeta,
642
+ FrameStep,
643
+ TraceReconstructionError,
644
+ reconstruct,
645
+ )
646
+
647
+ __all__ = [
648
+ "FrameMeta",
649
+ "FrameStep",
650
+ "TraceReconstructionError",
651
+ "reconstruct",
652
+ ]
653
+ ```
654
+
655
+ - [ ] **Step 4: Run test to verify it passes**
656
+
657
+ Run: `.venv/bin/python -m pytest tests/viz/test_reconstruct.py -q`
658
+ Expected: PASS (5 tests).
659
+
660
+ - [ ] **Step 5: Commit**
661
+
662
+ ```bash
663
+ git add proteus/viz/__init__.py proteus/viz/reconstruct.py tests/viz/__init__.py tests/viz/test_reconstruct.py
664
+ git commit -m "feat(cp5): viz.reconstruct — deterministic frame replay with double self-verify"
665
+ ```
666
+
667
+ ---
668
+
669
+ ## Task 5: viz/terminal.py — truecolor grid + side panel
670
+
671
+ **Files:**
672
+ - Create: `proteus/viz/terminal.py`
673
+ - Modify: `proteus/viz/__init__.py` (export `render_terminal`)
674
+ - Test: `tests/viz/test_terminal.py`
675
+
676
+ - [ ] **Step 1: Write the failing test**
677
+
678
+ Create `tests/viz/test_terminal.py`:
679
+ ```python
680
+ from proteus.agents import VanillaAgent
681
+ from proteus.providers import FakeProvider
682
+ from proteus.runtime.session import SessionRunner
683
+ from proteus.viz import reconstruct, render_terminal
684
+
685
+
686
+ def _steps():
687
+ agent = VanillaAgent(FakeProvider(["ACTION: up"]))
688
+ trace = SessionRunner(
689
+ "predator_evade", agent, seed=42, play_turns=4, use_probe=False,
690
+ ).run()
691
+ return reconstruct(trace), trace
692
+
693
+
694
+ def test_render_terminal_emits_one_string_per_frame():
695
+ steps, _ = _steps()
696
+ buf = []
697
+ render_terminal(steps, fps=0, writer=buf.append, delay=0.0)
698
+ assert len(buf) == len(steps)
699
+
700
+
701
+ def test_render_terminal_emits_truecolor_and_panel_fields():
702
+ steps, _ = _steps()
703
+ buf = []
704
+ render_terminal(steps, fps=0, writer=buf.append, delay=0.0)
705
+ joined = "\n".join(buf)
706
+ assert "\033[38;2;" in joined # 24-bit truecolor escape
707
+ assert "action :" in joined
708
+ assert "motive :" in joined
709
+ assert "habit :" in joined
710
+ assert "tokens :" in joined
711
+
712
+
713
+ def test_render_terminal_does_not_sleep(monkeypatch):
714
+ steps, _ = _steps()
715
+ calls = []
716
+ monkeypatch.setattr("proteus.viz.terminal.time.sleep", lambda s: calls.append(s))
717
+ render_terminal(steps, fps=0, writer=lambda s: None, delay=0.0)
718
+ assert calls == [] # delay=0.0 -> never sleeps
719
+ ```
720
+
721
+ - [ ] **Step 2: Run test to verify it fails**
722
+
723
+ Run: `.venv/bin/python -m pytest tests/viz/test_terminal.py -q`
724
+ Expected: FAIL with `ImportError: cannot import name 'render_terminal'`.
725
+
726
+ - [ ] **Step 3: Write minimal implementation**
727
+
728
+ Create `proteus/viz/terminal.py`:
729
+ ```python
730
+ """Render reconstructed frames as truecolor terminal frames with a side panel.
731
+
732
+ Reuses the vendored arc_grid palette (COLOR_MAP) and the rgb_to_ansi / hex_to_rgb
733
+ helpers so the human-facing colors match the engine renderer. Each frame is drawn
734
+ as 2-char colored blocks at native grid resolution (keeps it terminal-width) with
735
+ a right-hand panel summarizing the turn — action / motive / habit / reward /
736
+ tokens + a reasoning excerpt — built from the CP4.5 trace enrichment.
737
+
738
+ Output goes through an injected writer so it can be captured in tests without a
739
+ TTY; pass delay=0.0 (or fps=0) to disable inter-frame sleeping.
740
+ """
741
+
742
+ from __future__ import annotations
743
+
744
+ import time
745
+ from typing import Callable, Optional
746
+
747
+ from proteus.arc_grid.rendering import COLOR_MAP, hex_to_rgb, rgb_to_ansi
748
+ from proteus.viz.reconstruct import FrameMeta, FrameStep
749
+
750
+ _RESET = "\033[0m"
751
+ _BLOCK = "██"
752
+ _REASON_WIDTH = 60
753
+ _REASON_MAX_LINES = 2
754
+
755
+
756
+ def _reason_lines(reasoning: str) -> list[str]:
757
+ text = " ".join(reasoning.split())
758
+ if not text:
759
+ return []
760
+ chunks = [
761
+ text[i : i + _REASON_WIDTH] for i in range(0, len(text), _REASON_WIDTH)
762
+ ][:_REASON_MAX_LINES]
763
+ lines = [f"reason : {chunks[0]}"]
764
+ lines.extend(f" {c}" for c in chunks[1:])
765
+ return lines
766
+
767
+
768
+ def _panel_lines(meta: FrameMeta) -> list[str]:
769
+ if meta.phase == "cut":
770
+ return [f"Cut {meta.index}", "(scripted pre-roll)"]
771
+ head = f"Turn {meta.turn_idx}" + (" [DIAGNOSTIC]" if meta.is_diagnostic else "")
772
+ mark = "✓ congruent" if meta.was_congruent else "✗ diverged"
773
+ lines = [
774
+ head,
775
+ f"action : {meta.action} {mark}",
776
+ f"motive : {meta.motive_action}",
777
+ f"habit : {meta.habit_action}",
778
+ f"reward : {meta.reward:+.1f}",
779
+ f"tokens : in{meta.input_tokens} out{meta.output_tokens} think{meta.thinking_tokens}",
780
+ ]
781
+ lines.extend(_reason_lines(meta.reasoning))
782
+ if meta.terminal:
783
+ lines.append(f">>> {meta.terminal.upper()}")
784
+ return lines
785
+
786
+
787
+ def render_frame(step: FrameStep, color_map: Optional[dict[int, str]] = None) -> str:
788
+ """Return one frame as a multi-line truecolor string with its side panel."""
789
+ color_map = color_map or COLOR_MAP
790
+ grid = step.frame
791
+ height, width = grid.shape
792
+ panel = _panel_lines(step.meta)
793
+ lines: list[str] = []
794
+ for y in range(height):
795
+ parts = []
796
+ for x in range(width):
797
+ rgb = hex_to_rgb(color_map.get(int(grid[y, x]), "#000000FF"))
798
+ parts.append(f"{rgb_to_ansi(rgb)}{_BLOCK}{_RESET}")
799
+ row = "".join(parts)
800
+ side = panel[y] if y < len(panel) else ""
801
+ lines.append(f"{row}│ {side}" if side else f"{row}│")
802
+ pad = " " * (width * 2)
803
+ for k in range(height, len(panel)):
804
+ lines.append(f"{pad}│ {panel[k]}")
805
+ return "\n".join(lines)
806
+
807
+
808
+ def render(
809
+ steps: list[FrameStep],
810
+ *,
811
+ fps: float = 4.0,
812
+ writer: Callable[[str], None] = print,
813
+ delay: Optional[float] = None,
814
+ ) -> None:
815
+ """Render each frame in order through `writer`, pacing with `fps`/`delay`."""
816
+ if delay is None:
817
+ delay = (1.0 / fps) if fps and fps > 0 else 0.0
818
+ for step in steps:
819
+ writer(render_frame(step))
820
+ if delay:
821
+ time.sleep(delay)
822
+ ```
823
+
824
+ Modify `proteus/viz/__init__.py` to add the terminal export:
825
+ ```python
826
+ from proteus.viz.reconstruct import (
827
+ FrameMeta,
828
+ FrameStep,
829
+ TraceReconstructionError,
830
+ reconstruct,
831
+ )
832
+ from proteus.viz.terminal import render as render_terminal
833
+
834
+ __all__ = [
835
+ "FrameMeta",
836
+ "FrameStep",
837
+ "TraceReconstructionError",
838
+ "reconstruct",
839
+ "render_terminal",
840
+ ]
841
+ ```
842
+
843
+ - [ ] **Step 4: Run test to verify it passes**
844
+
845
+ Run: `.venv/bin/python -m pytest tests/viz/test_terminal.py -q`
846
+ Expected: PASS (3 tests).
847
+
848
+ - [ ] **Step 5: Commit**
849
+
850
+ ```bash
851
+ git add proteus/viz/terminal.py proteus/viz/__init__.py tests/viz/test_terminal.py
852
+ git commit -m "feat(cp5): viz.terminal — truecolor block grid + CP4.5 side panel"
853
+ ```
854
+
855
+ ---
856
+
857
+ ## Task 6: viz/png.py — per-frame PNG (matplotlib Agg)
858
+
859
+ **Files:**
860
+ - Create: `proteus/viz/png.py`
861
+ - Modify: `proteus/viz/__init__.py` (export `write_pngs`)
862
+ - Test: `tests/viz/test_png.py`
863
+
864
+ - [ ] **Step 1: Write the failing test**
865
+
866
+ Create `tests/viz/test_png.py`:
867
+ ```python
868
+ import pytest
869
+
870
+ pytest.importorskip("matplotlib") # PNG path needs the `viz` extra
871
+
872
+ from proteus.agents import VanillaAgent
873
+ from proteus.providers import FakeProvider
874
+ from proteus.runtime.session import SessionRunner
875
+ from proteus.viz import reconstruct, write_pngs
876
+
877
+
878
+ def _steps():
879
+ agent = VanillaAgent(FakeProvider(["ACTION: up"]))
880
+ trace = SessionRunner(
881
+ "predator_evade", agent, seed=42, play_turns=4, use_probe=False,
882
+ ).run()
883
+ return reconstruct(trace)
884
+
885
+
886
+ def test_write_pngs_one_nonempty_file_per_frame(tmp_path):
887
+ steps = _steps()
888
+ out_dir = tmp_path / "frames"
889
+ paths = write_pngs(steps, out_dir)
890
+ assert len(paths) == len(steps)
891
+ for p in paths:
892
+ assert p.exists()
893
+ assert p.stat().st_size > 0
894
+ # Files are zero-padded and ordered.
895
+ assert (out_dir / "frame_000.png").exists()
896
+
897
+
898
+ def test_write_pngs_creates_missing_dir(tmp_path):
899
+ steps = _steps()
900
+ nested = tmp_path / "a" / "b" / "frames"
901
+ paths = write_pngs(steps, nested)
902
+ assert nested.is_dir()
903
+ assert len(paths) == len(steps)
904
+ ```
905
+
906
+ - [ ] **Step 2: Run test to verify it fails**
907
+
908
+ Run: `.venv/bin/python -m pytest tests/viz/test_png.py -q`
909
+ Expected: FAIL with `ImportError: cannot import name 'write_pngs'`.
910
+
911
+ - [ ] **Step 3: Write minimal implementation**
912
+
913
+ Create `proteus/viz/png.py`:
914
+ ```python
915
+ """Write reconstructed frames as per-frame PNGs (headless, matplotlib Agg).
916
+
917
+ Each frame is upscaled via the vendored frame_to_rgb_array and saved as
918
+ frame_000.png, frame_001.png ... with a one-line title (turn + action). The rich
919
+ side panel is a terminal feature; PNGs stay clean images with a label. matplotlib
920
+ is the optional `viz` extra and is imported lazily so the offline core never
921
+ needs it.
922
+ """
923
+
924
+ from __future__ import annotations
925
+
926
+ from pathlib import Path
927
+ from typing import Union
928
+
929
+ from proteus.arc_grid.rendering import COLOR_MAP, frame_to_rgb_array
930
+ from proteus.viz.reconstruct import FrameStep
931
+
932
+
933
+ def _title(step: FrameStep) -> str:
934
+ meta = step.meta
935
+ if meta.phase == "cut":
936
+ return f"Cut {meta.index} (pre-roll)"
937
+ mark = "✓" if meta.was_congruent else "✗"
938
+ title = f"Turn {meta.turn_idx}: action={meta.action} {mark}"
939
+ if meta.terminal:
940
+ title += f" [{meta.terminal}]"
941
+ return title
942
+
943
+
944
+ def write_pngs(
945
+ steps: list[FrameStep],
946
+ out_dir: Union[str, Path],
947
+ *,
948
+ scale: int = 32,
949
+ ) -> list[Path]:
950
+ """Render every frame to `<out_dir>/frame_NNN.png`; return the paths in order."""
951
+ import matplotlib
952
+
953
+ matplotlib.use("Agg") # headless: no window, no display required
954
+ import matplotlib.pyplot as plt
955
+
956
+ out_dir = Path(out_dir)
957
+ out_dir.mkdir(parents=True, exist_ok=True)
958
+ paths: list[Path] = []
959
+ for i, step in enumerate(steps):
960
+ rgb = frame_to_rgb_array(0, step.frame, scale, COLOR_MAP)
961
+ fig, ax = plt.subplots(figsize=(4.0, 4.4))
962
+ ax.imshow(rgb, interpolation="nearest")
963
+ ax.axis("off")
964
+ ax.set_title(_title(step), fontsize=9)
965
+ path = out_dir / f"frame_{i:03d}.png"
966
+ fig.savefig(path, dpi=100, bbox_inches="tight")
967
+ plt.close(fig)
968
+ paths.append(path)
969
+ return paths
970
+ ```
971
+
972
+ Modify `proteus/viz/__init__.py` to add the PNG export:
973
+ ```python
974
+ from proteus.viz.png import write_pngs
975
+ from proteus.viz.reconstruct import (
976
+ FrameMeta,
977
+ FrameStep,
978
+ TraceReconstructionError,
979
+ reconstruct,
980
+ )
981
+ from proteus.viz.terminal import render as render_terminal
982
+
983
+ __all__ = [
984
+ "FrameMeta",
985
+ "FrameStep",
986
+ "TraceReconstructionError",
987
+ "reconstruct",
988
+ "render_terminal",
989
+ "write_pngs",
990
+ ]
991
+ ```
992
+
993
+ Note: `write_pngs` imports matplotlib lazily INSIDE the function, so `import proteus.viz` (and `__init__` importing `proteus.viz.png`) does NOT import matplotlib — only calling `write_pngs` does. The module-level `from proteus.viz.png import write_pngs` is safe because `png.py` has no top-level matplotlib import.
994
+
995
+ - [ ] **Step 4: Run test to verify it passes**
996
+
997
+ Run: `.venv/bin/python -m pytest tests/viz/test_png.py -q`
998
+ Expected: PASS (2 tests). If matplotlib is missing, the file is skipped (importorskip) — but Task 0 installed it, so it should run.
999
+
1000
+ - [ ] **Step 5: Commit**
1001
+
1002
+ ```bash
1003
+ git add proteus/viz/png.py proteus/viz/__init__.py tests/viz/test_png.py
1004
+ git commit -m "feat(cp5): viz.png — per-frame titled PNGs via matplotlib Agg"
1005
+ ```
1006
+
1007
+ ---
1008
+
1009
+ ## Task 7: viz package import-safety guard
1010
+
1011
+ A tiny safety test: importing `proteus.viz` must NOT import matplotlib (offline invariant — matplotlib loads only when `write_pngs` is called).
1012
+
1013
+ **Files:**
1014
+ - Test: `tests/viz/test_import_safety.py`
1015
+
1016
+ - [ ] **Step 1: Write the test**
1017
+
1018
+ Create `tests/viz/test_import_safety.py`:
1019
+ ```python
1020
+ import sys
1021
+
1022
+
1023
+ def test_importing_viz_does_not_import_matplotlib():
1024
+ # Drop any pre-loaded matplotlib so we observe a fresh import graph.
1025
+ for name in list(sys.modules):
1026
+ if name == "matplotlib" or name.startswith("matplotlib."):
1027
+ del sys.modules[name]
1028
+ for name in list(sys.modules):
1029
+ if name == "proteus.viz" or name.startswith("proteus.viz."):
1030
+ del sys.modules[name]
1031
+
1032
+ import proteus.viz # noqa: F401
1033
+
1034
+ assert "matplotlib" not in sys.modules, (
1035
+ "importing proteus.viz pulled in matplotlib; keep it lazy in write_pngs"
1036
+ )
1037
+ ```
1038
+
1039
+ - [ ] **Step 2: Run test to verify it passes**
1040
+
1041
+ Run: `.venv/bin/python -m pytest tests/viz/test_import_safety.py -q`
1042
+ Expected: PASS. (If it FAILS, a top-level matplotlib import leaked into the viz package — move it back inside `write_pngs`.)
1043
+
1044
+ - [ ] **Step 3: Commit**
1045
+
1046
+ ```bash
1047
+ git add tests/viz/test_import_safety.py
1048
+ git commit -m "test(cp5): guard viz package against eager matplotlib import"
1049
+ ```
1050
+
1051
+ ---
1052
+
1053
+ ## Task 8: CLI — `play` subcommand
1054
+
1055
+ **Files:**
1056
+ - Modify: `proteus/cli.py` (add `_cmd_play`, parser, import)
1057
+ - Test: `tests/cli/test_cli.py` (add)
1058
+
1059
+ - [ ] **Step 1: Write the failing test**
1060
+
1061
+ Append to `tests/cli/test_cli.py`:
1062
+ ```python
1063
+ def test_play_human_writes_comparable_trace(tmp_path, monkeypatch, capsys):
1064
+ # Feed scripted moves through builtins.input (HumanAgent resolves it lazily).
1065
+ inputs = iter(["up"] * 20)
1066
+ monkeypatch.setattr("builtins.input", lambda *a, **k: next(inputs))
1067
+ out = tmp_path / "runs" / "human.jsonl"
1068
+ rc = main([
1069
+ "play",
1070
+ "--scenario", "predator_evade",
1071
+ "--seed", "42",
1072
+ "--play-turns", "5",
1073
+ "--out", str(out),
1074
+ ])
1075
+ assert rc == 0
1076
+ traces = read_traces(out)
1077
+ assert len(traces) == 1
1078
+ assert traces[0].model == "human"
1079
+ assert traces[0].scenario == "predator_evade"
1080
+ # The run summary names the scenario.
1081
+ assert "predator_evade" in capsys.readouterr().out
1082
+
1083
+
1084
+ def test_play_unknown_scenario_errors(capsys):
1085
+ rc = main(["play", "--scenario", "nope", "--seed", "1", "--play-turns", "1"])
1086
+ assert rc == 2
1087
+ assert "Unknown scenario" in capsys.readouterr().err
1088
+ ```
1089
+
1090
+ - [ ] **Step 2: Run test to verify it fails**
1091
+
1092
+ Run: `.venv/bin/python -m pytest tests/cli/test_cli.py::test_play_human_writes_comparable_trace -q`
1093
+ Expected: FAIL — argparse exits because `play` is not a registered subcommand.
1094
+
1095
+ - [ ] **Step 3: Write minimal implementation**
1096
+
1097
+ In `proteus/cli.py`, add `HumanAgent` to the agents import near the top:
1098
+ ```python
1099
+ from proteus.agents import HumanAgent, VanillaAgent
1100
+ ```
1101
+
1102
+ Add the command handler (place after `_cmd_run`):
1103
+ ```python
1104
+ def _cmd_play(args: argparse.Namespace) -> int:
1105
+ if args.scenario not in list_scenarios():
1106
+ print(
1107
+ f"Unknown scenario {args.scenario!r}. "
1108
+ f"Available scenarios: {', '.join(list_scenarios())}.",
1109
+ file=sys.stderr,
1110
+ )
1111
+ return 2
1112
+ agent = HumanAgent()
1113
+ runner = SessionRunner(
1114
+ args.scenario,
1115
+ agent,
1116
+ difficulty=Difficulty(args.difficulty),
1117
+ seed=args.seed,
1118
+ play_turns=args.play_turns,
1119
+ use_probe=args.probe,
1120
+ )
1121
+ trace = runner.run()
1122
+ accuracy = trace.metrics["motive_reading_accuracy"]
1123
+ print(
1124
+ f"{trace.scenario} seed={trace.seed} {trace.difficulty} "
1125
+ f"model={trace.model} -> {trace.outcome} "
1126
+ f"| motive_reading_accuracy={accuracy:.1f}% "
1127
+ f"| reactivity_index={trace.metrics['reactivity_index']:.1f}%"
1128
+ )
1129
+ if args.out:
1130
+ written = append_trace(trace, args.out)
1131
+ print(f"trace appended to {written}")
1132
+ return 0
1133
+ ```
1134
+
1135
+ In `build_parser()`, add the `play` subparser (after the `run` subparser block):
1136
+ ```python
1137
+ play = sub.add_parser("play", help="play a session as a human via stdin")
1138
+ play.add_argument("--scenario", default="predator_evade")
1139
+ play.add_argument(
1140
+ "--difficulty", default="easy", choices=[d.value for d in Difficulty]
1141
+ )
1142
+ play.add_argument("--seed", type=int, default=None)
1143
+ play.add_argument("--play-turns", type=int, default=15, dest="play_turns")
1144
+ play.add_argument(
1145
+ "--probe",
1146
+ action="store_true",
1147
+ help="also ask the per-turn comprehension probe (default: off for humans)",
1148
+ )
1149
+ play.add_argument(
1150
+ "--out", default=None, help="optional JSONL file to append the human trace to"
1151
+ )
1152
+ play.set_defaults(func=_cmd_play)
1153
+ ```
1154
+
1155
+ - [ ] **Step 4: Run test to verify it passes**
1156
+
1157
+ Run: `.venv/bin/python -m pytest tests/cli/test_cli.py -q`
1158
+ Expected: PASS (existing CLI tests + the two new `play` tests).
1159
+
1160
+ - [ ] **Step 5: Commit**
1161
+
1162
+ ```bash
1163
+ git add proteus/cli.py tests/cli/test_cli.py
1164
+ git commit -m "feat(cp5): CLI 'play' subcommand — human session via stdin"
1165
+ ```
1166
+
1167
+ ---
1168
+
1169
+ ## Task 9: CLI — `replay --visual / --png / --fps`
1170
+
1171
+ **Files:**
1172
+ - Modify: `proteus/cli.py` (`_cmd_replay` + `replay` subparser)
1173
+ - Test: `tests/cli/test_cli.py` (add)
1174
+
1175
+ - [ ] **Step 1: Write the failing test**
1176
+
1177
+ Append to `tests/cli/test_cli.py`:
1178
+ ```python
1179
+ def _write_fake_trace(tmp_path):
1180
+ out = tmp_path / "r.jsonl"
1181
+ main([
1182
+ "run", "--scenario", "predator_evade", "--model", "fake:x",
1183
+ "--seed", "42", "--play-turns", "4", "--no-probe", "--out", str(out),
1184
+ ])
1185
+ return out
1186
+
1187
+
1188
+ def test_replay_text_mode_unchanged(tmp_path, capsys):
1189
+ out = _write_fake_trace(tmp_path)
1190
+ capsys.readouterr() # drain
1191
+ rc = main(["replay", str(out)])
1192
+ text = capsys.readouterr().out
1193
+ assert rc == 0
1194
+ assert "turn 1" in text # legacy text behavior preserved
1195
+
1196
+
1197
+ def test_replay_visual_emits_truecolor(tmp_path, capsys):
1198
+ out = _write_fake_trace(tmp_path)
1199
+ capsys.readouterr()
1200
+ rc = main(["replay", str(out), "--visual", "--fps", "0"])
1201
+ text = capsys.readouterr().out
1202
+ assert rc == 0
1203
+ assert "\033[38;2;" in text # truecolor escape present
1204
+
1205
+
1206
+ def test_replay_png_writes_frames(tmp_path, capsys):
1207
+ import pytest
1208
+
1209
+ pytest.importorskip("matplotlib")
1210
+ out = _write_fake_trace(tmp_path)
1211
+ pdir = tmp_path / "png"
1212
+ rc = main(["replay", str(out), "--png", str(pdir)])
1213
+ assert rc == 0
1214
+ frames = list(pdir.glob("frame_*.png"))
1215
+ assert frames
1216
+ assert all(p.stat().st_size > 0 for p in frames)
1217
+ assert "PNG" in capsys.readouterr().out
1218
+ ```
1219
+
1220
+ - [ ] **Step 2: Run test to verify it fails**
1221
+
1222
+ Run: `.venv/bin/python -m pytest tests/cli/test_cli.py::test_replay_visual_emits_truecolor -q`
1223
+ Expected: FAIL — argparse rejects the unknown `--visual` flag.
1224
+
1225
+ - [ ] **Step 3: Write minimal implementation**
1226
+
1227
+ In `proteus/cli.py`, replace `_cmd_replay` with a version that branches to viz when `--visual`/`--png` is set (keep the existing text loop verbatim as the default branch):
1228
+ ```python
1229
+ def _cmd_replay(args: argparse.Namespace) -> int:
1230
+ try:
1231
+ traces = read_traces(args.trace_file)
1232
+ except FileNotFoundError:
1233
+ print(f"trace file not found: {args.trace_file}", file=sys.stderr)
1234
+ return 2
1235
+ if not traces:
1236
+ print(f"no traces in {args.trace_file}", file=sys.stderr)
1237
+ return 1
1238
+
1239
+ if args.visual or args.png:
1240
+ from proteus.viz import reconstruct, render_terminal, write_pngs
1241
+
1242
+ for trace in traces:
1243
+ steps = reconstruct(trace)
1244
+ if args.visual:
1245
+ render_terminal(steps, fps=args.fps)
1246
+ if args.png:
1247
+ paths = write_pngs(steps, args.png)
1248
+ print(f"wrote {len(paths)} PNG frames to {args.png}")
1249
+ return 0
1250
+
1251
+ for trace in traces:
1252
+ print(
1253
+ f"=== {trace.scenario} seed={trace.seed} {trace.difficulty} "
1254
+ f"model={trace.model} -> {trace.outcome} ==="
1255
+ )
1256
+ for turn in trace.turns:
1257
+ tag = "congruent" if turn.was_congruent else "DIVERGED"
1258
+ diag = " [diagnostic]" if turn.is_diagnostic else ""
1259
+ print(
1260
+ f" turn {turn.turn_idx}: action={turn.action} "
1261
+ f"motive={turn.motive_action} habit={turn.habit_action} "
1262
+ f"{tag}{diag} reward={turn.reward}"
1263
+ )
1264
+ print(f" metrics: {trace.metrics}")
1265
+ return 0
1266
+ ```
1267
+
1268
+ In `build_parser()`, extend the `replay` subparser with the three flags (add after the existing `replay.add_argument("trace_file", ...)` line):
1269
+ ```python
1270
+ replay.add_argument(
1271
+ "--visual", action="store_true", help="truecolor terminal replay"
1272
+ )
1273
+ replay.add_argument(
1274
+ "--png", default=None, metavar="DIR",
1275
+ help="also write per-frame PNGs to DIR (needs the 'viz' extra)",
1276
+ )
1277
+ replay.add_argument("--fps", type=float, default=4.0, help="replay frames/sec")
1278
+ ```
1279
+
1280
+ - [ ] **Step 4: Run test to verify it passes**
1281
+
1282
+ Run: `.venv/bin/python -m pytest tests/cli/test_cli.py -q`
1283
+ Expected: PASS (all CLI tests, including the three new replay tests).
1284
+
1285
+ - [ ] **Step 5: Commit**
1286
+
1287
+ ```bash
1288
+ git add proteus/cli.py tests/cli/test_cli.py
1289
+ git commit -m "feat(cp5): CLI replay --visual/--png/--fps (text stays default)"
1290
+ ```
1291
+
1292
+ ---
1293
+
1294
+ ## Task 10: Full-suite verification + acceptance demo
1295
+
1296
+ This task runs the offline gate and the spec §11 acceptance demonstration (one human trace produced and compared against an LLM trace), then records results for the handoff.
1297
+
1298
+ - [ ] **Step 1: Run the complete offline suite**
1299
+
1300
+ Run:
1301
+ ```bash
1302
+ .venv/bin/python -m pytest -q
1303
+ ```
1304
+ Expected: ALL pass (the prior 75 + all CP5 additions), no network, no display window.
1305
+
1306
+ - [ ] **Step 2: Generate an LLM (fake) trace and a human trace at the same seed**
1307
+
1308
+ Run:
1309
+ ```bash
1310
+ .venv/bin/python -m proteus run --scenario predator_evade --model fake:demo \
1311
+ --seed 42 --play-turns 6 --no-probe --out runs/cp5_llm.jsonl
1312
+ printf 'up\nup\nup\nup\nup\nup\n' | .venv/bin/python -m proteus play \
1313
+ --scenario predator_evade --seed 42 --play-turns 6 --out runs/cp5_human.jsonl
1314
+ ```
1315
+ Expected: each command prints a summary line (`predator_evade seed=42 easy model=… -> … | motive_reading_accuracy=… | reactivity_index=…`) and appends one trace.
1316
+
1317
+ - [ ] **Step 3: Verify the two traces are structurally comparable**
1318
+
1319
+ Run:
1320
+ ```bash
1321
+ .venv/bin/python -c "
1322
+ from proteus.runtime import read_traces
1323
+ h = read_traces('runs/cp5_human.jsonl')[0]
1324
+ l = read_traces('runs/cp5_llm.jsonl')[0]
1325
+ assert h.cut_frames == l.cut_frames, 'cut frames differ'
1326
+ assert [t.motive_action for t in h.turns] == [t.motive_action for t in l.turns], 'answer keys differ'
1327
+ assert h.model == 'human' and l.model == 'demo'
1328
+ print('OK comparable:', 'human', h.metrics.get('motive_reading_accuracy'), '| llm', l.metrics.get('motive_reading_accuracy'))
1329
+ "
1330
+ ```
1331
+ Expected: prints `OK comparable: human <acc> | llm <acc>` with no assertion error.
1332
+
1333
+ - [ ] **Step 4: Exercise the visual replay paths**
1334
+
1335
+ Run:
1336
+ ```bash
1337
+ .venv/bin/python -m proteus replay runs/cp5_human.jsonl --visual --fps 0
1338
+ .venv/bin/python -m proteus replay runs/cp5_human.jsonl --png runs/cp5_frames
1339
+ ls runs/cp5_frames/frame_000.png
1340
+ ```
1341
+ Expected: the terminal shows a colored grid + side panel for each frame; `runs/cp5_frames/` contains `frame_000.png …`; the `ls` lists the first frame. (`runs/` is gitignored — these artifacts are not committed.)
1342
+
1343
+ - [ ] **Step 5: Update HANDOFF.md (CP5 close) and the venv recreate command**
1344
+
1345
+ Edit `HANDOFF.md`:
1346
+ - Move CP5 from "Next" to a new "Done (CP5)" section summarizing: `HumanAgent` (I/O-injected, same SessionRunner → comparable human trace), `proteus/viz/` (reconstruct + terminal + png), CLI `play` + `replay --visual/--png/--fps`, and the acceptance demo result from Step 3.
1347
+ - Update the `.venv` recreate command to append `"matplotlib>=3.8"` (now required by the viz/png tests).
1348
+ - Set "Next" to CP6 (difficulty layouts + metric refinement + human-baseline harness).
1349
+ - Under "Open questions", add the deferred **probe-vs-act-reasoning redundancy** review (spec change, out of CP5) and **optional live-color human play**.
1350
+
1351
+ Then commit:
1352
+ ```bash
1353
+ git add HANDOFF.md
1354
+ git commit -m "docs(cp5): handoff close — human play + trace viz shipped; CP6 next"
1355
+ ```
1356
+
1357
+ ---
1358
+
1359
+ ## Self-Review (completed by plan author)
1360
+
1361
+ **Spec coverage:**
1362
+ - §1 휴먼 플레이 → Tasks 1, 2, 3, 8. §1 시각화 → Tasks 4, 5, 6, 9.
1363
+ - §2 모듈 배치 (human.py, viz/reconstruct|terminal|png, cli) → Tasks 1–9. Dependency direction (viz leaf; SessionRunner untouched; arc_grid/trace unchanged) honored.
1364
+ - §3 HumanAgent (act/probe/reset, I/O injection, name="human", ASCII live view, Ctrl-C abort) → Tasks 1, 2 (Ctrl-C is the natural default — no quit action coded, as designed).
1365
+ - §4.1 reconstruct + double self-verify → Task 4. §4.2 terminal + side panel → Task 5. §4.3 png Agg titled → Task 6.
1366
+ - §5 CLI (play + replay flags) → Tasks 8, 9.
1367
+ - §6 tests (all offline/headless) → every task is TDD; matplotlib guarded by importorskip + Task 7 import-safety.
1368
+ - §7 invariants (offline, Agg-only-on-png, no SessionRunner edit) → Task 0 + Task 7 + Task 9 lazy import.
1369
+ - §8 완료 기준 → Task 10 (full suite + human/LLM comparison demo + visual/png).
1370
+
1371
+ **Placeholder scan:** No TBD/TODO; every code step shows complete code; every command shows expected output.
1372
+
1373
+ **Type consistency:** `FrameStep(frame, meta)` / `FrameMeta` fields are defined in Task 4 and consumed identically in Tasks 5 (`render_frame`/`_panel_lines`) and 6 (`_title`). `reconstruct`, `render_terminal` (exported alias of `terminal.render`), `write_pngs`, `TraceReconstructionError` names are consistent across `viz/__init__.py` and all consumers. `HumanAgent(input_fn, output_fn)` signature is consistent between Tasks 1/2 and the CLI default-construction in Task 8. `ActResult`/`ProbeResult` token fields default to 0 (verified against `agents/base.py`).