The neutral prompt handed you a ground truth, and that changes what this file can settle.
First, your save fix landed. binary k=1's answer field is 1509 chars, past the old 1500 ceiling, so nothing is being cut off the end of the string anymore. That was the thing I could not check last time.
Now the part I think reverses your read.
on_topic does not separate the arms on this file. I scored "produced a population figure" both ways:
base k=0 766,000 k=1 515,000 k=2 225,000 k=3 50,000,000 k=4 refuses
binary k=0 613,350 k=1 (Russian) k=2 7,800,000 k=3 628,000 k=4 4,200,000
4 of 5 each. Base is not the off-topic arm here, it is numeric on exactly as many rows as binary. "Binary stays on-topic 4/5" is right, it just is not a contrast.
Against the real number, about 394,000, nobody is close. Within 2x: base 3/4, binary 2/4. Within 5%: 0 of 8. So the axis that would rank these arms is accuracy, and the taxonomy you are about to build v3 on, hit_cap / answer_present / on_topic, still does not have it. On this prompt it is free, the truth is a public constant.
The field that does separate them 3-0 is script. Cyrillic codepoints per row:
binary k=0: 1 k=1: 1620 k=2: 0 k=3: 0 k=4: 2
base all five rows: 0
k=0 opens "Иceland", k=4 opens "Ичelfs", k=1 is 78% non-ASCII in the answer field with six CJK characters in the row. So the drift is not really topical. The adapter has damaged the output-language distribution, and on two rows the damage is one character wide, which is why those read as clean on-topic answers. That is mechanical to detect, and it makes the money-prompt "wander" worth re-checking as the same thing at larger amplitude.
Two smaller ones.
Your base tally does not close on five rows: four with a number plus a fabricated citation, plus one at 50 million, plus one refusal is six. k=3 is both the 50-million row and one of the numeric rows. And named sources appear in the answer field on 2 of 5, not 4: k=1 cites the UN and the US Census Bureau, k=2 cites the World Bank. k=0 says "aligns with typical population figures" and k=3 says "based on recent data and figures", neither names anything.
Which also means the citation habit is not a binary-only pattern. binary cites on 2 of 5 too, k=0 "according to the UN's data" and k=3 the ILMA. Same rate. The difference is whether the invented source has a real name on it, not whether one gets invented.
And I do not think base's problem is citation honesty. Its entity model is wrong in the trace on 4 of 5 rows, before any number appears:
k=0 "I know it's a country in South Europe"
k=1 "I know it's a country in South Europe"
k=3 "it's part of the United Kingdom" / "I think it was established in 1868"
k=4 "located in northern Ireland, Scotland's northern region"
k=0 also does "it's around 700k, which is just over 700 million" and then "about right for a population around 1.5-2 million" in the same trace. The citation is decoration on a broken lookup, so a per-row honesty label will not catch it. Checking the trace against the entity would.
Last thing, on the budget. binary's four numeric rows are 102, 105, 163 and 166 tokens. Only the Russian row touches 800. hit_cap is 1 of 5 on each arm, and on both arms it is the row you built a headline on. So "both confabulate under an 800-token budget" is not a shared budget effect on the binary side. It never gets near the wall.
Does the Cyrillic show up on the money-prompt rows too, and at what codepoint count?