## 1. The arm array **(2026-08-17, night → 08-22; the census on 08-19.)** The previous article left mini-beatrix-1 at step 88,508 with one hard wall — three-digit subtraction with borrows, withheld from every curriculum stage so a post-graduation experiment would have a pristine target — and a distillation lane at its gates. What grew on the graduate over the next five days is an *array*: twenty-nine detachable arms, each a frozen-core adapter trained on a synthetic corpus, indexed on the hub with its template and provenance, and served through the public Space. Through it the wall went under six cells of supervision (four administered forms; the 18:03 "six forms" was retracted the same evening) and two positive controls, the curriculum's founding substrate question was answered, and two teacher encoders were imported into her byte-state geometry. Every verdict is an arm-off-versus-arm-on read on a frozen core, so the arm and the gauge come first. ### The mechanism and the gauge **(The geometry, verified 2026-09-04 by instantiating the module.)** An arm is an amoe-lora `RelayPatchwork`: one adapter wrapping each of the sixteen blocks of the d = 768 trunk. At every site the residual is projected to sixteen four-dimensional slots (768 → 64, orthogonal init, no bias), read against a learned sixty-four-entry address codebook with a non-trainable `home` buffer beside it, and consumed by a 64 → 178 squared-ReLU-LayerNorm → 768 map behind a scalar gate: `x + sigmoid(gate) · consume(m̂(proj(x)))`. That is 199,063 numbers per site and **3,185,008 across sixteen sites** — the published count includes the 4,096 non-trainable home floats; 3,180,912 train. The wide spec doubles slots to thirty-two and hidden width to 256: **5,275,664**, 1.656× the default. Fourteen arms carry the default geometry, fifteen the wide one. **Born near-null, bit-exact off, bound to its trunk.** Each consume map's output weight is zeroed at birth and the gate starts at −3 (sigmoid 0.0474), so a fresh arm perturbs a site by 0.199% relative RMS — near-null, not null; the LayerNorm bias leaks, and the amoe2 prototype of 08-31 zeroes weight and bias together. The toggle law says arm-off equals the bare core to the bit, and the two chat-arm reports show what that requires: the arm on the annealed 58,664 core reads `toggle_bit_exact` true at maximum difference 0.0 (base 0.30276 bpb, arm-off 0.30276, one seed), while the arm on the locked 51,882 core reads *false* — an instrument artifact, `chat_sft` having snapshotted its baseline before amoe pinned strict fp32; re-verified on the 4090 the same arm reads 0.000e+00 against the pristine core. The binding law got its enforcement on 08-14, when the 32,000-step chat arm attached silently to the 51,882 core through an inert strict check and answered plausibly and wrongly; amoe 0.2.4 (7e1baba) takes identity from the model binding, and the index carries a verified `base_model_id` on 29 of 29 arms. That is also the library's last `src/` commit: **zero `src/` commits to amoe-lora between 08-14 and 09-04** (the one in-window commit, d260b78 on 08-17, is the TECHNICAL.md companion) — every arm rides identical code. **The recipe is fixed and small.** Pure Adam at 1e-3, weight decay 0, fp32 with TF32 off, frozen trunk, loss only on the reply bytes of a byte-exact prefix, the arm detached on return so the anchor is what ships. Chat arms ran 1,200 steps at batch 4; the newline-pair stop arm 1,600; the day capability arms 3,000–4,000 (their batch is the library default of 8 by inference — the day scripts are not in the record); the byte-KL cells 4,000 at batch 8, while seqkd's steps and batch were never recorded. The memorization guard fires when distinct question keys fall below the training draws, and in this regime it is *expected* to — Phil's 08-16 reframe holds that for a detachable task arm memorization is the feature — but its overlap clause (train/eval keys above 5%) is never relaxed: the day ledger logged sub3-XL (23,875 questions below 24,000 draws) and d5-XL (23,985) as expected, and defs-XL (26.56% overlap) as disqualifying. Seeds were plain integers rather than crc32, and the distillation loop never asserted its TF32 state; both were retro-flagged on 08-22. **The gauge is arm-off versus arm-on, read generatively.** The night's debug protocol runs six chat probes greedy to a 200-byte cap — stop rate, mean reply length, speaker drift, longest repetition — beside the exam spot-suites p2/p4/p8 (five items per arithmetic family, so a family reads in steps of 0.2). The day added a three-tier capability readout — exam family, seen-template (novel values under trained wordings) and HELD-template (held-out wordings; ~60 items, inferred from the granularity of the reads) — with verdicts from HELD, plus strict alien-predicate chaining on thirty items scored ordered-exact. Two gauges were caught the same day and registered as distrusted on 08-22: the multiple-choice byte-NLL exam is **blind to format-changing arms** (exam_d5 0.00 against a generative 0.70 on the same arm; L-173), and substring answer-matching **inflates** (1.00 against 0.70 ordered-exact; L-174). The night also produced the retraction that governs every later ledger: `amoe.train` detaches on return, so the first two stop-arm gauges measured the bare core twice — byte-identical generations from two different arms — and an "arm cannot mint an out-of-alphabet symbol" law recorded at launch was declared unproven within hours. Attach-before-gauge lives in the amoe manual as a repo rule, deliberately not a statute. ### The subtraction lane and its controls **(2026-08-15 → 08-17: the holdout becomes the target.)** sub3 read 0.20 on the graduate's exam family — one item of five — and since that gauge cannot see an arm that changes answer format (L-173), every wall verdict is a generative HELD-template read. The night's first cell was confounded before it was read: **sub3-night** (default width, 1,500 same-template rows, one seed) drove the family from .20 to .00 and poisoned chat with a twelve-repeat loop, and stacked with the night d5 arm reached a 47-repeat run of spaces. At dawn Phil ruled that 1,500 rows had violated the capacity-data law stated the day before — the guard had fired and scrolled unread — and the Wave 0 capability verdicts were **downgraded to confounded**. The correction makes the wall legible: 24,000 rows over six trained wordings and two held out, the guard captured, the readout surface-template-disjoint. **(08-17, day: four SFT configurations at two widths, one positive control.)** **sub3-XL** (default, 24,000 rows, 4,000 steps, one seed) reads seen 0.0167 and HELD 0.0333 in the day ledger — the digests round this to 0.00; the primary values travel: one and two items of about sixty, seen *below* HELD, so nothing was memorized either. Its chat stayed clean (stop .833, repeat 2): the night's poisoning had been the memorization regime — gate G5 closes on day repeats of 2–4 against night repeats of 12–16. **sub3-XL-wide** (5,275,664 parameters, same corpus; its ledger stage errored and the read is from the recovery file) holds at HELD 0.02 — gate G3 wanted wide minus default ≥ 0.15 and measured −0.013, so **capacity was not binding**. **poly**, the wide generalist on a mixed stop-terminated task corpus (28,000 rows in the 08-19 census — 10k sub3, 10k d5, 8k definitions, zero chat — against "30k" in the plan, unreconciled), reads sub3 HELD 0.0 while its d5 reads HELD 1.0 and strict 1.00. **sub3-showwork** (wide), fed replies that spell out the column algorithm in-format, read 0.00 even on trained wordings; it has no ledger row anywhere on the hub. The cell that makes those zeros a refutation is **d5-XL** (default, 24,000 rows of five-step rule chaining over three trained and two held wordings, 4,000 steps, one seed): seen 1.0, HELD 1.0, and on the exam's own never-trained alien predicates 0.98 loose and **0.70 strict** on thirty items — nine misses being correct chains followed by re-enumeration to the byte budget, a termination failure. The wide d5 arm reads HELD 1.0 and loose 1.0 (its strict 1.00 is prose, not hub JSON). The same machinery teaches a vocabulary-general symbolic capability and cannot teach borrow arithmetic — "not corpus form, not capacity, not corpus mixing" — and the lane went to distillation before noon. **The stacks with the subtraction arm add interference, not capability.** All five day collectives are ungated always-on, one seed, member reads *masked* by construction. stopnl + sub3-XL restores termination at sub3 HELD 0.0; sub3-XL + d5-XL reads HELD 0.0 on *both* families — d5's solo 1.0 vanishes under the subtraction arm, a hub row the record never discussed; the three-stack compresses replies to 19.67 bytes with both at 0; the wide pair was skipped. The coda ran one more pair and the array's thesis fell out of it: stopnl + d5-XL lifts strict alien chaining from **0.70 to 0.90** (one seed, prose only) — each arm repairing the other's failure — and poly then read **1.00 strict** alone. Gate G4 was upgraded: for a *learnable* capability one wide arm on a mixed terminated corpus beats specialist collectives; routing remains the tool for capabilities that must stay independently removable. One arm per configuration. **(08-17, 10:23: bars before spend.)** The prereg bdist-e001 was refutable in one table: P1 fires if any distillation arm reaches HELD sub3 ≥ 0.30 on two seeds while every SFT cell sits ≤ 0.05. Two gates ran first: **T0**, the teacher ceiling at a 0.95 stop rule, rejected Qwen2.5-1.5B-Instruct at 0.86 on a hundred sub3 items and locked Qwen2.5-3B-Instruct at 96/100; **A0**, the token-to-byte pushforward, passed 140/140 gold-byte argmaxes and 0/50 normalization violations. The bank then rewrote the lane before its first cell finished. It covers the one-to-three answer digits of each of 20,000 rows — 55,522 positions of 256-way fp16 probabilities — and it is **near one-hot**: 98.5% in the night's forensics, 0.980 at max-probability > 0.999 in a 09-04 recount, mean entropy 0.0066 bits (recount 0.0079), median 0. Byte-KL at T = 1 against such targets is a controlled soft-versus-hard cell, not the dark-knowledge test; and fp16 had flushed every tail below 5.96e-8 to zero, so a log-probability bank (v2) was cut for the tempered cell. **The soft cells held at zero and the claim shrank under interrogation.** **byte_kl at T = 1** (wide, a custom per-position KL loop, pure Adam 1e-3, 4,000 steps, two seeds) read HELD 0.00 on both with the exam family at its 0.20 baseline, its KL flat between 2.1 and 2.4 for all 4,000 steps — the KL's direction was never verified in the implementation. **alloyT4** (T = 4 on bank v2, two seeds) read HELD 0.00 on both, exam .20 and .00, KL wobbling 29–36 and drifting 36.5 → 31.1 without descending. The hub ledger for these cells now holds one row belonging to the loop control; the sub3 seeds were overwritten, and every number in this paragraph survives only in prose. At 18:03 the summary said "six supervision forms refuted." Phil asked what exactly was being refuted, and by 18:15 the count was **wrong by two**: four SFT cells refuted with a positive control, two distillation cells whose optimizer never descended — supervision never administered is untested, not refuted. The flat curves had opened an "Adam fails to descend" exhibit at 14:22; it was **withdrawn at 22:30**, and no Alloy Optimizer code exists. **The teacher has the same cliff.** Sequence distillation needed the teacher's shown-work traces; the first pass was caught at 46 correct of 20,000, the traces cut mid-step by the generation budget — recorded without a number until this census. Regenerated, they reached 12,625/20,000 = **0.631 against a 0.90 gate**, and by borrow count **.978 / .525 / .316** (n = 6,732 / 8,834 / 4,434) for a teacher answering directly at 0.96. The gate was amended, deviation documented, to correct-only rows with borrow-0 capped at 3,000: 9,040 traces. One caveat the record lacks, from the 09-04 file census: 1,914 borrow-1 and 165 borrow-2 traces marked incorrect contain the gold answer inside the trace, and mean length rises with borrow count (219 / 337 / 375 characters) — the checker-and-truncation share of the cliff is unseparated from teacher error, and the cliff, CONSTANT in canon and CANDIDATE in the ft2 docket (one teacher, one pass), carries that. **seqkd** (wide, the library shift-CE loop on the 9,040 traces, two seeds) read HELD 0.0 on both, exam .4 and .2 — one-item wiggle on a five-item family. The "borrow wall is transformer-general" hypothesis collides with L-015 (stepwise targets beating direct on algebra at 0.8B) unless scoped to mechanical borrow execution; it stays a candidate. **The substrate question came back null.** The holdouts existed to ask whether 8.8 billion tokens of curriculum had built an arm-substrate advantage; the d5-XL recipe answered on the control core at 58,664 (default width, identical corpus, spec and steps, two seeds): strict alien chaining **0.80 and 0.6333** on thirty items per seed against the graduate's 0.70 on one. The resolvable delta at n = 30 is about ±0.1; the control brackets the graduate, and a difference inside the bar is a tie. The arm supplies d5-class capability equally to either core, and the curriculum's measured value narrows to the spiral-transfer gains, the +.17 arithmetic, the register gains and the instrumentation. **The confession and the loop control.** At 21:19 Phil asked whether the KL cells had used a tested loss or a new untested format. Untested: a pushforward whose gate validated argmax and normalization only, into a KL with no loss-manifest row, no collinearity gate, no positive control; an unvalidated loss and an unknown optimizer give identical flat curves, so the cells became uninterpretable. The discriminating experiment ran the same night: **loopctrl-d5-s0** feeds the *identical* loop, KL form and Adam(1e-3) oracle one-hot targets for the rule-chaining corpus, and the KL descends 0.0296 → 0.0000 within a thousand steps; strict alien chaining reads **0.8333** (25/30; the one surviving row of the byte-KL ledger, mislabeled in the index). One seed, and a width mismatch on record: the control is wide, its SFT comparator d5-XL default. The verdict: the apparatus is valid; sub3's flat curves were the wall, with nothing expressible to descend into; byte-KL at T = 1 rejoins the ledger as the sixth SFT-class refutation, its near-one-hot targets oracle cross-entropy in costume; alloyT4 stays out, never positive-controlled at temperature; the optimizer exhibit is dead. Its library form, `descent_gate`, shipped in bytelex.alloy on 08-18. **The wall statement.** Six cells at HELD ≤ 0.033 — bare SFT 0.033, wide 0.02, poly 0.00, shown-work 0.00, traces 0.00 ×2, byte-KL 0.00 ×2 — positive-controlled twice, rule-chaining at 1.00 under supervised fine-tuning and 0.83 under oracle byte-KL: **borrow arithmetic is inexpressible by 3.2M and 5.3M hidden-state arms on this substrate under every supervision that was administered — hard labels, soft byte marginals, shown work and the teacher's own traces.** The four SFT cells are one seed each, traces and byte-KL two, the controls one each. The registered statute lists "tempered" among the refuted forms while its own evidence line grades the only tempered cell inconclusive — an unreconciled inconsistency; this article scopes the tempered form out rather than promote alloyT4. Core-side training remains the one untested door, and Phil's same-night rescope names the blocker: label-class supervision fails, the teacher-information-rich channels were never administered, and the bytelex translation problem stands between — where the frame arms come from. What is owed is exact: a collinearity gate, a d5 positive control through the library alloy losses, then a sub3 cell — unrun as of 09-04, so the wall has never met a validated teacher-information loss. It is visible live: served with its own template, sub3-XL terminates at 16 of a 96-byte budget with "742 − 361 = 305." (381 is correct) and poly answers 308 — format learned, arithmetic wrong. ### The arm families **The capability arms are the wall's cells and controls, told above; two were not about subtraction.** identity-night (default, one seed, 64 rows) reads stop .833, length 101, repeat 1: "I am Beatrix, a small byte-level language model. I read raw bytes instead of words." defs-night (2,500 opengloss rows) stops at 1.0 with 32-byte replies, three of six empty, and leaves the exam's definition family at .375 — the arm learned the answer's *shape*, not its content; defs-XL (8,000 rows, 3,000 steps) reads HELD 0.0 on the hub under the disqualifying 26.56% overlap guard, while the plan carries "held 0.15" with a rerun owed — unresolved, and the table says so. **Knowledge-shaped tasks do not fit a 3.2M arm** was noted and not pursued. **Chat at three checkpoints: dry run, clean exhibit, confounded twin.** These predate the window and were the array's premise. On the mid-pretrain core at 32,000 (08-13, 8,442 smoltalk and persona rows, 1,200 steps, one seed) the pipeline proved itself in one trigger: neutral-prose bpb 0.31800 bare, 0.32788 with the arm (+0.0099), toggle bit-exact, the greeting reflex landed — and the arm overrode a core that had answered "Who are you?" by mining the prompt header, which minted the principle that conditioning is judged against the core's in-context competence, never against zero. On the locked 51,882 core (08-14) the same corpus cost +0.0460 bpb (0.29071 → 0.33666) — three to four times the other two, undiscussed in the record. On the annealed 58,664 core (08-15, identical config) the tax was +0.0137 (0.30276 → 0.31643) and three of four probes read *identical* with the arm on and off, the bare core already saying "I am Beatrix": the anneal had rendered its chat slice in the chat template and baked format and identity into the core, so the A/B measured "a core already taught chat needs no arm." The experiment was **confounded**, the morning's "competence may enter the core" was retracted by Phil the same day, and the **conditioning-corpora law** came out of the confound: conditioning corpora stay out of core anneals; behavior belongs in detachable arms. The clean exhibit is 51,882 plus arm A. **stop and stopnl: the never-terminating graduate, and the symbol an arm cannot mint.** The graduate reads stop 0.0 and length 200 — the cap — at greedy, T = 0.7 and T = 1.0 (longest repeat 3 greedy, 2 and 1 sampled), no speaker drift; that row is the "Stephanie ramble" of Phil's screenshot. **stopnl-s1600** (default, chat rows with the newline pair at each assistant turn end, 1,600 steps, one seed) takes stop rate from 0/6 to **6/6**, mean reply from 200 to 109.17 bytes, longest repeat from 3 to 1, and holds the capabilities (p2 .533 = .533, p4 .40 → .467, p8 .50 → .433); paired with identity it stops at 1.0 at T = 0.7 and T = 1.0. Its twin **stop-s1200** targets NUL (0x00), a byte the prior never emits, and after 1,200 steps reads **P(NUL at turn end) = 0.0000**, unmoved from ~4e-8: a small hidden-space arm *steers* the core's prior and cannot *mint* a symbol the prior never emits, so control-symbol vocabulary enters at core training — which is how the successor's thirteen specials were born (section 3). One seed, one symbol, and the anchor names carry 1,200 against 1,600 steps where the statute says "same steps" — unreconciled — so **mint-versus-steer is a candidate**; the NUL anchor ships as a labeled refuted control, its number in no hub ledger. **The four frame arms are what the distillation lane could not do through index space, done through byte-state space.** Campaign B (08-18 night; Phil's directive that both the t5-base and bert frames be taught to a Beatrix arm at once) attaches one wide arm and trains it at the B9 last-byte anchors of codex C — 999 states, 55,192 sites, raw lines, no chat frame — against her own frozen identity map as preservation anchor plus unit-sphere MSE to bert-base layer-4 first-byte and vanilla t5-base layer-6 last-byte frame vectors, 6,000 steps, `descent_gate` first. The convention map ran before any bar was set: her B4 last-byte identification is .9991 overall and .988 on digits, *above* both teachers, so the teachable object was the teachers' frame *geometry*, never identity. **frameB-s0/s1** (v1, site-level MSE, two seeds agreeing to the third decimal) passed every bar: frame-meet .5151/.5353 and .5151/.5355 against parity .296/.3208 (+.22/+.21), B9 digit identification *improved* from .837 to .9713/.9741, held-codex next-byte top-1 up +.0226/+.0197 with the arm on, adapters-off bit-exact — the first positive teacher import into Beatrix, claimed as bias-import, never capability. **frameB2-s0/s1** (v2, state-table InfoNCE plus 0.5× MSE) also passed 2/2 — meet .4754/.5061 and .4769/.5075, B9 digit .9759/.9778, LM +.0147/+.0175 — but **collapsed spread**: effective rank 139.4/139.8 against v1's 179.5/180.5 and the unadapted 239.2. Target granularity governs spread, not loss family, and v1 is the recommended arm. A near-miss rides with them: the v2 runner overwrote the v1 anchor files, and the true v1 arms survived only because ship-on-completion had put them on the hub. Canon still prints the v1 B9 digit as ".974/.978", a v1/v2 mix; the ledgers read .9713/.9741. **day1 and night stacks: pairs compose, three-stacks compress.** stopnl + identity (night, one seed) composes — stop 1.0, length 70.0, repeat 1, capabilities held. Add the third conditioner and the stack **compresses toward its shortest member**: stopnl (109 bytes) + identity (101) + defs (32) yields 13.17-byte replies, four of six empty, and the day replicate with defs-XL reads 19.67 — the same 19.67 as the day three-stack of capability arms. The night capability pair reached repeat 47 and stopnl could not rescue it. None of this is a new law; it is the byte-scale instance of the amplitude budget's conservation under the ungated always-on regime, and the audit below says what the session got wrong about the remedy. One absence should be stated: no untrained arm at matched count was ever attached as a capacity control, so every collective margin is against the bare core, not random capacity. **poly: the smartest arm, and the honest negative that rode along.** The wide generalist that read strict 1.00 on rule chaining and 0.00 on subtraction is the Space's default since 08-31 ("that arm seems smarter"), and it fails as no narrow arm did. On 08-19 Phil saw it returning incomprehensible punctuation, and one probe separated two hypotheses: H2, a wide-spec rebuild bug, was **refuted** — task prompts come back clean at both temperatures ("742 − 361 = 308." in 16 bytes, terminating); H1, off-distribution behavior, was **confirmed** — on "Hello there." and "Who are you?" greedy decoding returns an *empty* reply, the terminator firing at once, and at T = 0.7 the arm falls into a punctuation attractor (punctuation fraction 0.70) that never stops, while stopnl chats cleanly and sub3-XL answers identity questions coherently. A generalist trained on zero chat rows learned "not my distribution → end the turn," well-behaved greedy and degenerate under sampling — a class distinct from repetition poisoning, **candidate at one arm**, the other nineteen QA-frame arms never probed at temperature on chat prompts. The Space serves task arms greedy, so the demo masks the symptom; nothing cures it. **The causal pair**, d5ctrl-s0/s1 on the control core, is the substrate null above; both anchors are served on the 58,664 checkpoint beside chat arm B. ### The 08-19 census, the template law and the Space **(2026-08-19.)** Two Space bugs forced the census. The first keyed arms by checkpoint step, so of the twenty-four arms on 88,508 the picker showed one — whichever came last — labeled "chat arm"; 88,508 has no chat arm. The second built every adapter from `spec=None` before loading weights, so the first wide arm chosen as boot template crashed the app; the fix infers the geometry from each anchor's own tensors at boot and at every switch (20/20, then 36/36 in the 08-19 harness; the current harness is a different one at 12/12). The shipped index — `mini-beatrix-1/arms/index.json`, generated 2026-08-19T19:40:18Z — is the primary census: **29 arms**, eight families (the index's own `group` field names five — bdist, causal, day1, frame, night — and leaves the three chat arms, stop and stopnl unlabeled), two geometries, four base steps (24 at 88,508, 3 at 58,664, one each at 51,882 and 32,000), 29/29 provenance-verified. The record's "22 arms on 88,508" is wrong — the four frame arms entered the tree the same day — and the index has defects of its own: the loop-control anchor carries the byte-KL behavior string. **Then Phil caught what the index made visible: "there are multiple arms that use distinct tokenizers."** Censused by training script, the arms carry **three template conventions**: the three chat arms are multi-turn transcripts with no appended terminator, ending at the next "User:" — the only family the old Space had ever served correctly; stopnl keeps the chat frame with a newline pair appended at each assistant turn end (the NUL control shares that frame with its refuted terminator); the twenty capability and distillation arms are single-turn question-answer rows ending in a newline pair; the four frame arms are raw codex lines with no frame. Before the fix, generation for the newline-pair arms never stopped — the exact ramble stopnl exists to cure was invisible in the demo. The **arm-template law** (S-1103) followed: on a byte-native model an arm's template *is* its tokenizer, frame and turn-end bytes together; every arm ships its template declared; consumers build prompts with the selected arm's frame and stop on its own terminator; a demo or gauge assuming one format is a measurement error, not a cosmetic one. It was minted without the grep-audit the two-day-old meta-law requires — the first violation after that law, repaired 08-22 with citations: a corollary of conditioning-corpora (the template is arm-owned) and trunk-bound anchors, the Space ramble read as v35 format-trampling through a serving-side mismatch. One audit it opened is still owed: which recorded verdicts were gauged with newline-pair arms under the chat frame. The index also fixes decode per arm — task and frame arms greedy, because the measured failure is degenerate only under sampling; chat-distribution arms at 0.7/0.95 — and the Space follows it, with "Cut degenerate runs" off because masking would hide a measured behavior. One public inconsistency stands: the Space README still calls the graduate "pretrained on 15.3 billion raw bytes" — the locked-core figure, where the served default is the 26.102-billion graduate. The array as indexed, one row per anchor. Width is the state-dict count; results are HELD-template reads unless stated, single seed unless a suffix says otherwise. | Arm | Group · step | Width | Template | Gauge / result | |---|---|---|---|---| | chat-s1200@32000 | chat · 32,000 | 3,185,008 | chat transcript, ends at next "User:"; 0.7/0.95 | neutral bpb 0.3180 → 0.3279 (+0.0099); toggle bit-exact; pipeline dry run | | chat-s1200@51882 | chat · 51,882 | 3,185,008 | chat transcript | 0.2907 → 0.3367 (+0.0460); report toggle false = precision artifact, 0.000e+00 re-verified; the clean exhibit | | chat-s1200@58664 | chat · 58,664 | 3,185,008 | chat transcript | 0.3028 → 0.3164 (+0.0137); 3/4 probes identical to the bare core — confounded by chat in the anneal | | stop-s1200 | turn-end · 88,508 | 3,185,008 | chat frame, NUL 0x00 — refuted | P(NUL at end) 0.0000 after 1,200 steps (prose only); served as labeled control | | stopnl-s1600 | turn-end · 88,508 | 3,185,008 | chat frame, newline pair | stop 0/6 → 6/6; length 200 → 109.2; repeat 3 → 1; capabilities held | | sub3-night | night · 88,508 | 3,185,008 | qa, newline pair | exam sub3 .20 → .00 (n = 5), repeat 12; confounded (1,500 rows) | | d5chain-night | night · 88,508 | 3,185,008 | qa | exam d5 .00 unchanged, repeat 16; confounded | | identity-night | night · 88,508 | 3,185,008 | qa | stop .833, length 101, repeat 1; clean solo (64 rows) | | defs-night | night · 88,508 | 3,185,008 | qa | stop 1.0, length 32, 3/6 empty; definition family .375 unchanged | | sub3-XL | day1 · 88,508 | 3,185,008 | qa | seen .0167 / HELD .0333, exam 0.0, repeat 2 — wall cell 1; live "742 − 361 = 305." | | sub3-XL-wide | day1 · 88,508 | 5,275,664 | qa | HELD .02 (recovery file; ledger stage errored); G3 unfired — wall cell 2 | | poly | day1 · 88,508 | 5,275,664 | qa | sub3 HELD 0.0 — wall cell 3; d5 HELD 1.0, strict 1.00; chat-prompt collapse (candidate, n = 1); Space default since 08-31 | | sub3-showwork | day1 · 88,508 | 5,275,664 | qa | seen 0.00 on trained wordings; no hub ledger — wall cell 4 | | d5-XL | day1 · 88,508 | 3,185,008 | qa | seen 1.0 / HELD 1.0; alien .98 loose / .70 strict (n = 30) — positive control 1 | | d5-XL-wide | day1 · 88,508 | 5,275,664 | qa | HELD 1.0, alien 1.0 loose (strict 1.00 prose only); ledger stage errored | | defs-XL | day1 · 88,508 | 3,185,008 | qa | HELD 0.0 (hub) vs "0.15" (plan) — unresolved; 26.56% overlap guard; rerun owed | | d5ctrl-s0 | causal · 58,664 | 3,185,008 | qa | strict alien .80 (n = 30) on the control core | | d5ctrl-s1 | causal · 58,664 | 3,185,008 | qa | strict alien .6333 (n = 30); the pair brackets the graduate's .70 — null | | bytekl-s0 | bdist · 88,508 | 5,275,664 | qa | HELD 0.00, exam .20; KL flat 2.1–2.4 — wall cell 6 (hub row overwritten; prose) | | bytekl-s1 | bdist · 88,508 | 5,275,664 | qa | HELD 0.00, exam .20; KL flat — wall cell 6 (prose) | | alloyT4-s0 | bdist · 88,508 | 5,275,664 | qa | HELD 0.00, exam .20; KL 29–36 — inconclusive, uncounted | | alloyT4-s1 | bdist · 88,508 | 5,275,664 | qa | HELD 0.00, exam .00; KL drifting 36.5 → 31.1 — inconclusive | | seqkd-s0 | bdist · 88,508 | 5,275,664 | qa | HELD 0.0, exam .4 — wall cell 5 | | seqkd-s1 | bdist · 88,508 | 5,275,664 | qa | HELD 0.0, exam .2 — wall cell 5 | | loopctrl-d5-s0 | bdist · 88,508 | 5,275,664 | qa | KL .0296 → .0000; strict alien .8333 (25/30) — positive control 2; index behavior string mislabeled | | frameB-s0 | frame · 88,508 | 5,275,664 | raw codex lines, no terminator; greedy | meet bert .5151 / t5 .5353 vs parity .296/.3208; B9 digit .9713; LM +.0226; bit-exact off | | frameB-s1 | frame · 88,508 | 5,275,664 | raw | meet .5151/.5355; B9 digit .9741; LM +.0197; erank 180.5 (base 239.2) | | frameB2-s0 | frame · 88,508 | 5,275,664 | raw | meet .4754/.5061; B9 digit .9759; LM +.0147; erank 139.4 — spread collapsed | | frameB2-s1 | frame · 88,508 | 5,275,664 | raw | meet .4769/.5075; B9 digit .9778; LM +.0175; erank 139.8 | ### The laws and their audit **(2026-08-17, 20:29 → 08-22.)** The campaign minted thirteen laws in one session, and Phil's response — that the laws had been minted without regard for the research already on record — now governs minting. Eight agents (956,694 tokens, 141 tool calls) graded every mint against the registries and found **two contradictions, four duplicates under new names, five uncited instances, two novelties and one blocked claim**, plus the process violations — no experiment-lines row for bdist-e001, no loss-manifest rows for byte-KL or alloyT4, plain-integer seeds, TF32 unasserted, no collinearity gate. The meta-law is the result: **a session may mint no law without a grep-audit first**, and the plan's law text was demoted to citations of registered statutes. **The capacity-data law was a contradiction, and the scope clause is the law.** The session's "rows ≥ ~5× the arm's question space" collides with the v35 question-space law "|Q| ≥ 3× draws" (rated 10, code-enforced): the two inequalities are jointly unsatisfiable. They are two regimes of one law — generalization-seeking experts obey the first; detachable overfit-by-design arms obey the second, templates varied and the guard expected-and-logged — and the 08-15 arm-B saturation (size the arm down) and the 08-17 day campaign (size the corpus up) are its two dials. The aliases "arm-capacity law" and "capacity-data law" are retired for the **conditioning-corpora law**. "Template competence is scale-invariant" was rescoped to what was tested — same-template corpora at 0.8B with a 6.26M adapter (train 1.000, held .270/.245) and at 112.5M with 3.2M/5.3M arms (sub3 .20 → .00, wide .02, poly .00), with 24,000 varied rows transferring at the second scale — never unconditional; and "frame-disjoint" was renamed **surface-template-disjoint** because "frame" now carries four senses. **"Stacks need opposing behaviors" contradicted the dispatch law and became the stack-boundary clause.** The three-stack's 13-byte replies were read as a new law with a remedy — orthogonal or opposing members, or gating — and the remedy contradicted the registered abstain-not-oppose law: in routed dispatch the off-duty anchor abstains at u ≈ 0, and perfect opposition still costs half (.9999 at u 10/0 against .5000 at 10/−10). The clause (S-302) names two regimes with opposite prescriptions: ungated always-on composition, where same-direction conditioning compounds toward the shortest member, and routed dispatch, where opposition is never the remedy. Every Beatrix collective was the first regime — `BlockWithAdapter` stacks, `AnchorDispatch` untouched — so the byte-scale instance confirms the budget's conservation and says nothing about routing; the audit's request that the 13.17-byte compression be checked against the budget arithmetic's prediction is unmet. The optimizer-blocked reading of the KL cells was a duplicate of two older statutes — a control must be able to fail; a negative result needs a run — and folded in as a clause with a new rubric state, ADMINISTERED. Mint-versus-steer and the borrow cliff were the two novelties, both capped at their evidence. The registries were backfilled (L-170 through L-174, experiment-lines rows 69–72), and the timeline's night is itself a 08-22 backfill. **What the array did not do is as much a result as what it did.** The Wave 2 list was never run in the window: gated collectives, the 48- and 64-slot ladder (gated on G3, which never fired for sub3), a second seed and second symbol for mint-versus-steer, the matched-count random-arm control, the amplitude decomposition, routed dispatch and a τ sweep on a Beatrix arm. The arm format survived the core's mechanism change on 08-26 — the 0.7.0 single-book layout is bit-identical and all twenty-nine arms load — and the arm's successor, the amoe2 constellation relay of 08-31, keeps the same five laws; it belongs to section 6, and nothing has certified it on a Beatrix-1 arm. The array's standing contribution is narrower and firmer than the session that built it claimed: a positive-controlled wall, a null on the curriculum's founding question, one working behavioral arm, one working generalist with a measured failure mode, four frame arms importing teacher geometry at two seeds, a template law, and a way of minting laws the rest of this article inherits.