Rule 6 is missing the exact thing you built one section earlier, and rule 5 is right for the opposite reason to the one you gave.
The 0.27 was wrong, and wrong in my favour
You are right. Ran anatomie_du_bump.py on 000e41d. Your table reproduces line for line, and the line I should have written prints right under it:
amplitude sur les vingt autres cellules : |delta t(R)| max = 0.517
beta=0.005, R=25, moving t up, 2.430 to 2.947. My 0.27 was downward moves only. Struck. Same shape as your q90 sentence, one section apart, both wrong toward the point being made.
Rule 6 has no null, and section 3 is the template for it
Section 3 is the best thing in your reply. You caught that my leave-one-out was a maximum over twenty-five removals, built its null, and read 1.295 against E[max drop] 0.430. That is exactly right and it kills the attack properly.
Then rule 6 publishes a bare integer.
"Dies to 2 runs of 150" only means something against how many runs a real effect of that size dies to. So I planted one. Residuals from the R-level means, resampled with replacement; a constant shift on every R=25 run, calibrated so E[t] = 2.430; your cell sizes, n25=53 and n24=47. The effect is true by construction and identical in every replicate. Nothing selected, nothing fragile.
breakdown count of a GENUINELY REAL effect of exactly this size
median 4 mean 5.5
q10 1 q25 2 q50 4 q75 8 q90 11
P(breakdown <= 2) 0.252
P(breakdown <= 3) 0.398
A quarter of true effects this size die to two runs. Two fifths die to three. So your 2/150 and 3/150 are unremarkable for a real effect, and the integer did not tell you the bump was fragile. It told you t was 2.43 at n=150.
Which is what it measures. Scaling the planted effect:
effect E[t] median breakdown P(>= 10)
x1.0 2.43 3 0.15
x1.5 3.64 9.5 0.50
x2.0 4.86 18 0.94
x3.0 7.29 30 1.00
Earning a double-digit breakdown count half the time takes 1.5x the effect you measured. So rule 6 used as a bar is a power requirement wearing a robustness costume. That is a fine rule to have. It is just not a diagnostic, and published as one it will retire true findings at this n at about the rate it retires this one.
The fix is your own section 3. Report the count over its null: observed against E[breakdown | a real effect of the observed size, this design]. Here 2 against 4. That reads as "unremarkable for its size", which is honest, rather than "1.3 % of the sample", which reads as damning and is not.
One number I did not expect, out of the same simulation. With the effect true and planted, the share of replicates reaching |t| >= 1.98 is 0.665. The design has two-thirds power at the size it found. That is upstream of the inversion, the fragility and the interaction p, and neither of us had put a number on it.
Rule 5 is right. The sentence under it points the wrong way
effet_par_beta.py, line 54:
"R": int(len(np.unique(code))),
R is the number of distinct symbols the sender's code uses across the 27 referents. It is the alphabet size of the artifact. And objectif returns recompense = (s * r.t()).sum() / N, expected round-trip success over the same 27.
That is a bound, not a correlation. R symbols carry at most R distinguishable messages, so at most R referents decode:
reward <= R/27
discovery 150 / 150
replication 60 / 60
deficit = (R/27 - reward) * 27
discovery 0 in 141, 1 in 9
replication 0 in 49, 1 in 11
never 2, never negative, across all 210 runs
Pigeonhole, not regression. Your +0.9725 is the shadow of a bound that is tight in 141 runs out of 150. And the nine are not rounding: they are runs where the code holds R symbols and only R-1 of them decode, a receiver collision. Six of the nine sit at R=25.
So reward is a function of R, not R of reward. "R is the reward times twenty-seven" has the arrow backwards, and the direction is what makes rule 5 bite harder than you claimed. R is not the objective on a grid. It is the structure of the thing the run produced, and the objective is pinned to its ceiling. A column that bounds the objective is disqualified for the same reason as one that is it, and it stays disqualified in the runs where the optimiser misses the ceiling, which "the objective divided by 27" does not cover.
For completeness, since it prices what rule 5 costs you: swap R for k = round(27 * reward), the actual referent count, and the contrast is t = +2.226 against +2.430. Nearly the same row. Either column, same size, equally not a fact about a factor.
What I think is left
Rule 5 kills a row by reading the generator before the data, which is the only move so far that costs nothing and cannot be gamed. Rule 6, priced against its own null, does not kill this one.
Which leaves the bump as: real in the discovery sample, unremarkable in fragility for its size, absent on the second draw, and about a column that was never eligible. Three of those four are power.
At 0.665 power, every contrast in ยง7.25 is drawn from a regime where the ones that surface are inflated, including the two you are still defending. Is there a draw large enough to put an effect that size above 0.9, and if there is not, is the honest column next to each contrast the detectable-effect floor rather than the p?