The identity is exact and the column is withdrawn. Then your pilot does not work on this quantity, for a reason that turns out to be the same reason the floor was worth having. And your reframe of question 4 is better than mine, so I took it and followed it to a date on my own record.
Your identity, to zero
contrast |t| |t|/2.80 my ratio difference
R 27 vs 26 0.482 0.1721 0.1721 0.00e+00
R 26 vs 25 0.627 0.2238 0.2238 0.00e+00
R 25 vs 24 2.430 0.8677 0.8677 0.00e+00
R 24 vs 23 0.711 0.2540 0.2540 0.00e+00
beta .005/.03 2.968 1.0601 1.0601 0.00e+00
beta .005/.037 0.875 0.3125 0.3125 0.00e+00
Not close โ identical, because plancher = 2.80 * se and t = d / se in the same file. So "every observed effect is at or below its own detection floor" is |t| < 2.80, which at 145 df is p > 0.0058: an alpha 8.6 times stricter than the 0.05 written into the floor's own definition. And you are right that I published the same fact twice in one message and flagged only one instance โ my section C printed t = 2.97 โ p = 0.0035, which is section D's 1.06 row.
I replaced a p-value column with a rescaled p-value column and called it a design quantity, one message after explaining why observed power is a rescaled p-value. The ratio column is withdrawn. The correction to my question 3 is that "is it a function of the design" has to be checked on the printed expression, not on the intent.
The absolute floor survives, and it convicts something older than this table
Your prospective use is the real find, and it reaches further back than the row you applied it to. Both replication floors, computable before those runs existed:
n floor effect it had to confirm ratio
beta .005/.03 discovery 30+30 0.00925 0.00981 1.06
replication 12+12 0.01462 0.00981 0.67
R 25 vs 24 discovery 53+47 0.00728 0.00631 0.87
replication 30+12 0.01240 0.00631 0.51
The sixty runs I ran in round eight could not have confirmed either contrast even if both were exactly true. I reported that replication as a test โ "the sign flips", "it does not replicate", "the omnibus goes to p = 0.144" โ and it was never a test. It was an estimator, and an unbiased one, which is why -0.0053, SE 0.0033 remains the only honest line to publish from it. The conclusion I drew survives. The reason I gave for it does not: I read a failure to reach a bar that the draw could not have reached.
That is computable before running, does not depend on what came out, and no p on the replication can say it. Kept, printed alone and absolute.
And rule 5 does eat it on the four R rows. 8+30+53+47+12 is a partition of the discovery runs by an outcome; the 6.32 effective runs at R 27 vs 26 was a result. Both inputs to those floors fall out of the run. Only the beta rows at 30 and 30 are design โ and, as you note, the only contrast in the table that clears its floor is in the only column that was eligible.
Where the pilot fails, and it is not a detail
You propose 2.80 * sigma_pilot * sqrt(1/na + 1/nb) in the design document. On this quantity the pilot cannot fix sigma, and the reason is structural rather than practical.
beta n mean sd sd/mean
0.005 30 0.00768 0.01080 1.41
0.010 30 0.00943 0.01042 1.10
0.020 30 0.01005 0.00857 0.85
0.030 30 0.01749 0.01771 1.01
0.037 30 0.01057 0.01436 1.36
Bartlett chi2 = 19.176 p = 0.0007 sd ratio across levels 2.07
Heteroscedastic across the swept factor, so one pilot sigma misprices the floor by 12 % to 38 % depending on the cell. But the mechanism is what matters:
The gap is bounded below by zero. The unconstrained max is always at least the matched value, so max - appariee >= 0 by construction, and 63 of the 210 runs sit exactly at 0. A non-negative variable piled at its bound has its scale tied to its location, and it does:
across the 18 cells with n >= 4:
corr(cell mean, cell sd) Pearson +0.874 Spearman +0.917
slope sd on mean +0.817
median coefficient of variation 1.07
Sigma is not a nuisance scale here. It is roughly the quantity being measured. So a pilot fixes the floor only if the pilot's mean matches the eventual mean, which is to say only if you already know the effect. That is not a fixable pilot design; it is the floor's one data-dependent input being data in the strong sense.
But more of question 4 moves before the run than you proposed, in a different unit
The same fact that breaks the absolute floor makes a relative one writable. If sd โ CV ร mean with CV stable near 1, then
floor / mean = 2.80 * CV * sqrt(1/na + 1/nb)
which needs no sigma at all. For 30 seeds per cell and CV โ 1.2: 0.89. Checked against the actual table, floor 0.00925 over grand mean 0.01035 = 0.894.
Written before the first seed, in one line and with no pilot: with thirty seeds per cell, this design sees a near-doubling of the gap and nothing smaller. That is a stronger pre-registration than an absolute floor because it survives not knowing the scale, and it is the honest form for any non-negative quantity piled at zero โ which is most of what gets measured in this field. CV is the stable thing to pilot; sigma is not.
So: not the whole of question 4 moves, but the part that moves is bigger than an absolute floor and it moves in relative units.
Question 4, your reframe, and the date
It became unwritable by you, which is not the same as uncomputable. The defect is not that the data arrived first. It is that "what does this measurement feed" was never asked, and that question does not reference the results.
You are right, and it is a better diagnosis than mine. I said the threshold rots on contact with data. It does not โ it was never planted.
So I asked it, and it has an answer with a date on it.
The gap max - appariee measures the inflation from publishing the concentration statistic in its unconstrained-argmax form rather than the matched form. Its consumer was the 0.35 threshold in TEST3 ยง6.1, which decided whether an emergent code counted as compositional. That threshold was withdrawn on 11/08/2026, notebook ยง1.9, after you pointed out it was built on a sample maximum.
Rounds six through twelve have priced a measurement whose consumer was deleted in round five.
The distinction that keeps it honest, since not all of the work goes: the bound still has a consumer, because I report the max-form statistic and a reader needs to know it can be inflated by up to 0.14. That work stands. The contrast table โ does the inflation depend on R, on beta โ never had one. No decision anywhere changes at any value of that dependence, which is exactly why the relevance ratio was free to be 2 or 8 and why naming it now looks picked. There was nothing to pick from.
Which is what the design says too, if you ask it about the question that did have a consumer. That one is one-sample, not a contrast:
mean gap over 210 runs 0.01035 SE 0.00085 95 % CI [0.00869, 0.01201]
distance to the worst case 0.1339 = 158 standard errors
one-sample floor 0.00237
contrast floors in the table 0.00728 to 0.01445
The same runs are three to six times finer on the question with a consumer than on the contrasts without one, and they answered it at 158 sigma before any of this started.
What neither of us opened: the variable is not continuous
I noticed the zero bound while checking your pilot and did not stop to look at it. It is the largest thing in this file.
63 of the 210 runs have a gap of exactly zero. Thirty per cent. The gap is zero precisely when the unconstrained argmax is already a bijection โ no message position claiming the attribute another one claimed. So this is not a continuous quantity with a floor; it is a point mass plus a right-skewed positive part, and every t, every permutation, every Scheffe bar and every bootstrap in twelve rounds has been computed on it as though it were neither.
First consequence, and it lands on the number I have been defending since round seven. The bound 0.1443 comes from the transposition climber searching for the worst code โ which necessarily has an argmax collision. It is a worst case given a collision. My 0.0104 is unconditional, thirty per cent of it zeros. The published ratio compares a mixture to a conditional:
quantity value ratio to 0.1443
E[gap] over 210 runs (what I published) 0.01035 13.9
E[gap | gap > 0] (the matched one) 0.01479 9.8
median of the positive part 0.01254 11.5
q95 of the positive part 0.04049 3.6
max observed over 210 runs 0.05927 2.4
The ratio I have been quoting is 13.9 against a like-for-like 9.8 โ and if you compare the two quantities that are actually the same kind of object, a worst case against a worst case, it is 2.4. "Emergent codes come nowhere near what an adversarial search reaches" was built on the pairing that flatters it, and I chose that pairing without noticing there was one to choose.
Second consequence: the quantity is two quantities, and they have different consumers.
P(argmax collision) 0.700 95 % CI [0.633, 0.761]
E[inflation | collision] 0.01479 SE 0.00101
product 0.01035 = the published mean, exactly
How often the published statistic is wrong at all, and by how much when it is. Those answer different questions for a reader, and averaging them into one number answers neither. Never separated in any version of this document.
Third consequence: the tests. The positive part has skew +1.34, and neither it nor its log passes Shapiro โ 1.7e-09 and 2.2e-06 โ so it is not lognormal either, just skewed with no convenient form. Running the two contrasts of this exchange against the structure instead of through it:
contrast Student, raw on log(gap>0) Mann-Whitney
R 25 vs 24 t=+2.462 p=0.0156 t=+1.471 p=0.1459 p=0.0619
beta .005/.03 t=-2.589 p=0.0122 t=-2.113 p=0.0405 p=0.0285
The R contrast โ four rounds, a selection correction, a replication, a bump analysis, a pigeonhole argument โ goes to p = 0.062 the moment it is asked in a form the variable can answer. It was partly the Gaussian machinery reading a point mass as data.
And the decomposition says the contrasts are not about the collision rate: 13 % and 17 % of each comes from the zero proportion, Fisher p = 0.66 and 0.55. They are about the size of the inflation when it happens, which is the half with the smaller n โ 39 and 32 runs, not 53 and 47.
One number I will not claim. The observed zero rate, 0.300, differs from 6/27 = 0.2222 at binomial p = 0.0098 โ 6/27 being the permutation rate if the three argmaxes were independent and uniform. They are argmaxes of correlated mutual informations, so that reference is not a justified null and the p is not a finding. It is a number, and I am printing it as one.
Rule 7, and it was available on 11/08
When a claim is withdrawn, list every measurement whose only consumer it was, and stop measuring them. A retraction propagates downstream and nothing in my process made it propagate. ยง1.9 killed the threshold; the quantity it justified kept being measured, contrasted, corrected for multiplicity, replicated, corrected again, and defended across seven rounds against increasingly good statistics โ all of it correct, none of it attached to anything.
That is the rule that ends this thread, and unlike the previous six it costs nothing to run and cannot be gamed: it is a list, written at retraction time, of what the retracted claim was feeding.
I do not think you and I could have found it by getting better at the statistics. Every round did get better at the statistics.
And the zero mass is the same lesson from the other side. Twelve rounds of increasingly correct inference on a variable that neither of us had plotted. min, max and a count of exact zeros would have caught it in round one, cost nothing, and required no argument โ and they would have caught it before the bound comparison that all of this was downstream of.
Notebook ยง7.29, ยง1.21, ยง1.22. Code in src/test3_communication/plancher_de_detection.py and masse_en_zero.py.