Publish the model card as README so the limitations render on the repo page
Browse files
README.md
ADDED
|
@@ -0,0 +1,617 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- taiko
|
| 5 |
+
- rhythm-game
|
| 6 |
+
- chart-generation
|
| 7 |
+
- music
|
| 8 |
+
- audio-to-symbolic
|
| 9 |
+
---
|
| 10 |
+
|
| 11 |
+
# BarScript (experimental preview) β C1.3, denominator augmentation, seed 1234
|
| 12 |
+
|
| 13 |
+
Audio log-mel + a bar grid β a Taiko no Tatsujin chart, decoded under a
|
| 14 |
+
finite-state grammar over bar-scoped plan, skeleton and realization stages.
|
| 15 |
+
|
| 16 |
+
**BarScript** is the release name of this charting model; it is also the name of
|
| 17 |
+
the bar-scoped multi-stage token encoding it decodes into. It is a **preview
|
| 18 |
+
under evaluation**, published so it can be tried, not as a finished model: the
|
| 19 |
+
champion comparison below is `inconclusive`, and Β§2 lists measured defects β one
|
| 20 |
+
of which fires a pre-registered kill condition. The seven published SoftChart
|
| 21 |
+
1.x models are a separate, unaffected line.
|
| 22 |
+
|
| 23 |
+
| | |
|
| 24 |
+
|---|---|
|
| 25 |
+
| source checkpoint | `runs/sc2_c1_v13_s1234/best.pt` |
|
| 26 |
+
| sha256 | `44980bf64b1611ea73c1433c41adceff7596cfeece30dfaad5d228eae4a14e6e` |
|
| 27 |
+
| seed / step | 1234 / 96 000 of 100 000 |
|
| 28 |
+
| best validation CE | 0.5769365892960475 |
|
| 29 |
+
| parameters | 9 107 291 |
|
| 30 |
+
| shipped weights | bfloat16 safetensors, 18.2 MB (fp32 checkpoint preserved in the run dir) |
|
| 31 |
+
| training data | `JacobLinCool/taiko-1000-parsed-clean`, revision `b72da4616d643018e81f372cea06ce51349285e0` |
|
| 32 |
+
| label spec | `barscript_labels_v1`, frozen, shipped as `spec.json` (sha1 `cd876768β¦`) |
|
| 33 |
+
| training flags vs C1 | `--cal-axes big span density` Β· `--balance-exclude-big` Β· `--denom-augment 0.5 --denom-augment-mode lattice` |
|
| 34 |
+
| external pretraining | none |
|
| 35 |
+
|
| 36 |
+
**This card is written for THIS checkpoint.** Every number below was measured on
|
| 37 |
+
C1.3's own generations under the serving code this package ships
|
| 38 |
+
(`generate.py` md5 `73b6e49bβ¦`, `min_onset_gap_sec = 0.020`). Nothing is
|
| 39 |
+
inherited from the C1 card. Where an item on the C1 card has no C1.3
|
| 40 |
+
measurement, it is marked **not measured on this arm** rather than carried over.
|
| 41 |
+
|
| 42 |
+
**Read the KNOWN LIMITATIONS section before using any number above it.** This
|
| 43 |
+
model has real, measured defects, one of which fires a pre-registered kill
|
| 44 |
+
condition, and its difficulty ladder is roughly 25 coverage points less nested
|
| 45 |
+
than an authored one.
|
| 46 |
+
|
| 47 |
+
---
|
| 48 |
+
|
| 49 |
+
## 0. Status: the champion comparison is `inconclusive`
|
| 50 |
+
|
| 51 |
+
`scripts/sc2_gates.py champion`, three pairings on one serving code base
|
| 52 |
+
(`experiments/sc2_eval/FINAL_CHAMPION.md` Β§1):
|
| 53 |
+
|
| 54 |
+
| pairing | `champion_verdict` | reason code |
|
| 55 |
+
|---|---|---|
|
| 56 |
+
| C1.2 vs C1 | `inconclusive` | `no_seed_replicate` |
|
| 57 |
+
| **C1.3 vs C1** | **`inconclusive`** | `no_seed_replicate` |
|
| 58 |
+
| C1.3 vs C1.2 | `inconclusive` | `no_seed_replicate` |
|
| 59 |
+
|
| 60 |
+
`MIN_SEEDS_FOR_VERDICT` is 2 in **both** arms and C1.3 has one seed. The rule
|
| 61 |
+
fires before any metric is read. **No arm has been declared champion.**
|
| 62 |
+
|
| 63 |
+
What the clause ledger says, which is one-directional: C1.3 **passes** the
|
| 64 |
+
required Β§4 clause `plan_token_calibration`, which **C1 fails on both its
|
| 65 |
+
seeds**. No clause is passed by C1 and failed by C1.3. The two `a2` span clauses
|
| 66 |
+
fail on **all four arms** and belong to the campaign, not to this checkpoint.
|
| 67 |
+
|
| 68 |
+
The shipping rationale, the two named compromises and the decision rule for
|
| 69 |
+
choosing a different checkpoint are in
|
| 70 |
+
`experiments/sc2_eval/SHIP_DECISION.md`.
|
| 71 |
+
|
| 72 |
+
Gate items on the 71-chart hard+oni population, this arm against the C1 seed-1234
|
| 73 |
+
arm under identical code:
|
| 74 |
+
|
| 75 |
+
| | C1 | **C1.3** |
|
| 76 |
+
|---|---:|---:|
|
| 77 |
+
| fail / pass / na | 9 / 10 / 6 | **8 / 11 / 6** |
|
| 78 |
+
| newly passing | β | `control_own_axis_density`, `control_named_regression_density_big_count` |
|
| 79 |
+
| newly failing | β | `a2_span_placement_fit` (knife-edge, Β§2.6) |
|
| 80 |
+
|
| 81 |
+
---
|
| 82 |
+
|
| 83 |
+
## 1. What it does
|
| 84 |
+
|
| 85 |
+
* **Input** β 128-bin log-mel at 22 050 Hz (`preprocessor_config.json` freezes
|
| 86 |
+
the exact decode/STFT/mel contract), plus a bar grid with explicit measure
|
| 87 |
+
edges, exact rational meters and per-bar lattice denominators.
|
| 88 |
+
* **Conditions** β course, authored level, density bucket, and the split-axis
|
| 89 |
+
knobs `big_rate` / `span_rate` / `stream` / `sync`.
|
| 90 |
+
* **Output** β a BarScript token stream decoded under `BarscriptFSM`, exported
|
| 91 |
+
as TJA on the 96-slot lattice.
|
| 92 |
+
* **Family decoding** β `generate_chart_family` decodes a song's courses as one
|
| 93 |
+
nested ladder, hardest first, each easier course biased toward its harder
|
| 94 |
+
sibling's onsets. **Off by default**; the campaign's headline population was
|
| 95 |
+
decoded independently (Β§2.1).
|
| 96 |
+
|
| 97 |
+
Capabilities this checkpoint carries (`config.json:capabilities`): `aux`,
|
| 98 |
+
`beat_head` (hi-res), `hierarchical_ctx`, `axis_knobs`, `sync_token`,
|
| 99 |
+
`stage_emb`, `span_duration_head`, `tempo_head`, `rich_section_stats`. It
|
| 100 |
+
carries **no** `style`, `sibling`, `ctx`, `plan`-prefix, `slot`, `dual`,
|
| 101 |
+
`align`, `mask_infill`, `func_time`, `global_ctx` or `complexity` conditioning;
|
| 102 |
+
requests on those axes reach nothing.
|
| 103 |
+
|
| 104 |
+
### What works
|
| 105 |
+
|
| 106 |
+
Evidence marks: **β** n β₯ 71 charts with a measured C1 seed band; **β** n = 19β71
|
| 107 |
+
charts or 8β12 sweep cells, one seed; **β** n < 19 or a CI covering zero.
|
| 108 |
+
|
| 109 |
+
| axis | this checkpoint | reference | mark |
|
| 110 |
+
|---|---|---|---|
|
| 111 |
+
| **Density knob** β median within-song Ο **0.850**, sign p **0.0078**, 8/8 cells positive, **+1.413 nps** firstβlast, 12 % monotone adjacency | C1 moves the chart +0.024 / +0.151 nps over the same range and **fails** the gate; C1 seed-to-seed ΞΟ 0.003 (p 1.000) | 8 songs Γ 5 buckets, oni | β |
|
| 112 |
+
| **`big_rate` knob** β Ο **1.000**, p 0.00049, 91.7 % monotone, and **`calibrated: true`** β mean absolute bucket error **0.533**, within Β±1 bucket **93.3 %** | first axis in the whole campaign to clear `calibrated`; C1 0.533 vs 1.200 / 0.967 | 12 songs Γ 5 buckets, oni | β |
|
| 113 |
+
| **`span_rate` knob** β Ο 0.810, p 0.0117, `calibrated: true`, mean abs err 1.483; and unlike C1 it **no longer drags hit density with it** | C1's span knob flags `onset_nps` / `hit_nps` / `n_onsets` as interference; C1.3 flags only the definitional `span_count` | 12 songs Γ 5 buckets, oni | β |
|
| 114 |
+
| **Span budget** β 739 spans against the authored 762 (**0.970Γ**); `span_rate_gen_per_min` gap to authored **0.049** against a C1 seed band of 0.198 | C1 1.160β1.244Γ, C1.2 1.259Γ | 97 charts | β |
|
| 115 |
+
| **FSM guarantees** β **0 unclosed, 0 orphan ends, 0 swallowed hits** over 739 generated spans | holds on all four arms | 97 charts | β |
|
| 116 |
+
| **Placement does not degrade** β precision 0.6477 (C1 band 0.6480β0.6502), median \|offset\| **0 ms**, `exact_slot_lift_over_null` **1.1755** (above the C1 band), long-song drift tests significant **2** vs C1's 6 / 5 | see Β§2.3 for why apparent recall drops | 71 charts | β |
|
| 117 |
+
| **Accent over-emission reduced furthest** β pooled hard+oni big share **0.0873** against C1's 0.1082 / 0.1027 (3.3 seed bands) and an authored 0.0589 | still outside the pre-registered stop window 0.045β0.075 | 71 charts | β |
|
| 118 |
+
| **Cheaper** β 1 618.6 tokens/chart, 19.22 tokens/bar, `max_window_tokens_p99` **449** | C1 1 731.2 / 20.51 / 515 | 71 charts | β |
|
| 119 |
+
| **Deployment grid repaired** β on a forced `/16` BPM grid, notes Γ· authored **1.130 / 1.112 / 1.067 / 1.027** and onset F1 **0.333 / 0.374 / 0.466 / 0.580** | C1 1.688 / 1.461 / 1.297 / 1.210 and F1 0.318 / 0.376 / 0.477 / 0.599; C1.3 beats C1.2 on **12 of 12** deployable cells | 13 songs Γ 4 courses, one seed | β |
|
| 120 |
+
| **Empty-bar declaration restored** β `/16` empty bars 12.82 / 8.60 / 4.55 / 3.14 % against C1.2's 5.96 / 3.23 / 1.08 / 0.33 %, sign-significant on 3 of 4 courses | authored 15.69 / 12.30 / 9.08 / 6.27 % | 13 songs | β |
|
| 121 |
+
|
| 122 |
+
---
|
| 123 |
+
|
| 124 |
+
## 2. KNOWN LIMITATIONS
|
| 125 |
+
|
| 126 |
+
Nothing in this section is softened. Where a limitation is invisible to the gate
|
| 127 |
+
suite, that is said.
|
| 128 |
+
|
| 129 |
+
### 2.1 The difficulty ladder is 25β32 coverage points less nested than authored
|
| 130 |
+
|
| 131 |
+
Coverage = the fraction of the easier chart's notes that have a note in the
|
| 132 |
+
harder chart within tolerance. **Two decode modes, two different numbers, and
|
| 133 |
+
both belong on the record.**
|
| 134 |
+
|
| 135 |
+
| adjacent pair | authored charts | **independent decode** (package default) | **family decode Ξ² = 2** (authored grid) | **family decode Ξ² = 2** (`/16` deploy grid) |
|
| 136 |
+
|---|---:|---:|---:|---:|
|
| 137 |
+
| easy β normal | 0.970 | **0.658** | 0.730 | 0.726 |
|
| 138 |
+
| normal β hard | 0.979 | **0.667** | 0.797 | 0.811 |
|
| 139 |
+
| hard β oni | 0.985 | **0.660** | 0.822 | 0.834 |
|
| 140 |
+
|
| 141 |
+
Independent-decode figures: `FINAL_CHAMPION.md` Β§8.1, n = 13 / 13 / 35 songs, one
|
| 142 |
+
seed, same serving code as this package (β/β). Family-decode figures:
|
| 143 |
+
`C13_PREDICTION.md` Β§3.3, n = 13 songs, one seed (β).
|
| 144 |
+
|
| 145 |
+
Five things make this worse than the headline:
|
| 146 |
+
|
| 147 |
+
* **`nesting_coverage` and `hand_agreement` FAIL the family gate on every arm
|
| 148 |
+
measured**, C1 included. `difficulty_monotonicity` passes on all of them.
|
| 149 |
+
* **C1.3 is slightly worse than C1 here.** C1's independent-decode coverage on
|
| 150 |
+
the same population is 0.690 / 0.704 / 0.737; C1.3 is β0.032 / β0.037 / β0.077
|
| 151 |
+
against C1 seed bands of 0.010 / 0.021 / 0.025. The move is small against the
|
| 152 |
+
~0.30 gap to authored that every arm shares, but it is in the wrong direction.
|
| 153 |
+
* **Hand agreement at coinciding hits is near chance on hardβoni**: C1.3 0.539
|
| 154 |
+
against a marginal-chance null of 0.500 and an authored 0.753. When two
|
| 155 |
+
generated courses agree that a note belongs somewhere, which drum they pick is
|
| 156 |
+
near-independent across courses.
|
| 157 |
+
* **The `/16` figure overstates the model's nesting.** The same checkpoint on a
|
| 158 |
+
`/96` grid drops to 0.649 / 0.730 / 0.752, worse on 11β12 of 13 songs
|
| 159 |
+
(p = .003 / .022 / .022). A coarse lattice manufactures agreement by leaving
|
| 160 |
+
few places to disagree (`C13_PREDICTION.md` Β§5.2).
|
| 161 |
+
* **`family_bias = 2.0` / `family_hand_bias = 1.5` are PROVISIONAL** β described
|
| 162 |
+
in `generate.py` as logit offsets on a decoder never trained at that setting,
|
| 163 |
+
and never calibrated against the authored target.
|
| 164 |
+
|
| 165 |
+
### 2.2 The note-type channel carries almost no information
|
| 166 |
+
|
| 167 |
+
71 charts, **22 260 matched slots** (`FINAL_CHAMPION.md` Β§5.2, β).
|
| 168 |
+
|
| 169 |
+
| quantity | this model | its own floor | corpus Β§2.1 floor |
|
| 170 |
+
|---|---:|---:|---:|
|
| 171 |
+
| 4-way accuracy | **0.4776** | majority **0.5559** | majority 4-way **0.563** |
|
| 172 |
+
| β margin over majority | **β0.0783**, 95 % CI [β0.1022, β0.0530] | | |
|
| 173 |
+
| hand accuracy | **0.5493** | always-don **0.5975** | always-don **0.620** |
|
| 174 |
+
| β margin | **β0.0482**, CI [β0.0721, β0.0229] | | |
|
| 175 |
+
| **MI(generated; authored), 4-way** | **0.01028 bits** | β | **0.79 % of H(authored) = 1.300 bits** |
|
| 176 |
+
| MI, hand | 0.00534 bits | β | 0.55 % of 0.972 bits |
|
| 177 |
+
|
| 178 |
+
* **The model sits below its own majority floor on both endpoints**, and below
|
| 179 |
+
the campaign's corpus baselines (56.3 % / 62.0 %), which are a fixed reference
|
| 180 |
+
line from a separate 59 961-hit census, not this population.
|
| 181 |
+
* **The honest caveat, which cuts the other way.** The published floors are
|
| 182 |
+
*argmax* predictors scored against a temperature-1.0 top-p-0.95 *sample*.
|
| 183 |
+
Against the sampler-appropriate i.i.d. floor (Ξ£ pΒ² = 0.4582 4-way, 0.5190
|
| 184 |
+
hand) this model is **above** baseline by **+0.0194** and **+0.0303** β the
|
| 185 |
+
largest 4-way excess of any arm in the campaign. Both readings belong on the
|
| 186 |
+
record.
|
| 187 |
+
* **MI is flat at ~0.010β0.012 bits on all four arms** (seed band 0.0008). The
|
| 188 |
+
accuracy differences between arms are marginal-matching, not information.
|
| 189 |
+
* Accuracies are conditional on coverage **0.6442** (precision side) /
|
| 190 |
+
**0.5842** (recall side); more than a third of generated hits have no authored
|
| 191 |
+
partner, and coverage is 9 % lower than C1's, so this population is smaller
|
| 192 |
+
and differently selected than C1's.
|
| 193 |
+
* **No inter-charter agreement ceiling exists** for this split β no
|
| 194 |
+
(song, course) carries two independent authored charts β so 100 % is not a
|
| 195 |
+
legitimate target and is not used as one.
|
| 196 |
+
* **Greedy re-decoding was not repeated on this arm.** The C1-era finding that
|
| 197 |
+
greedy closes part of the gap at an unacceptable `motif_reuse` cost is **not
|
| 198 |
+
measured on this arm** and must not be quoted for it.
|
| 199 |
+
|
| 200 |
+
### 2.3 The note budget is 9 % short, and it is the whole of the apparent timing regression
|
| 201 |
+
|
| 202 |
+
| | C1 (s1234 / s4321) | **C1.3** |
|
| 203 |
+
|---|---:|---:|
|
| 204 |
+
| `onset_precision` | 0.6480 / 0.6502 | **0.6477** |
|
| 205 |
+
| note budget (gen Γ· authored notes) | 1.0046 / 0.9912 | **0.9069** |
|
| 206 |
+
| `onset_recall` | 0.6510 / 0.6445 | **0.5874** |
|
| 207 |
+
| `frac_of_authored_within_jnd` | 0.6463 / 0.6397 | 0.5840 |
|
| 208 |
+
| `type_coverage_recall` | 0.6465 / 0.6401 | 0.5842 |
|
| 209 |
+
|
| 210 |
+
`precision Γ budget` equals `onset_recall` to four decimals on **every** arm, by
|
| 211 |
+
construction. Those three "timing" rows are one quantity, and it is the note
|
| 212 |
+
budget, not placement: precision is inside the C1 seed band, median offset is
|
| 213 |
+
0 ms, `exact_slot_lift_over_null` is *above* the band, and significant drift
|
| 214 |
+
tests fall from 6 / 5 to 2.
|
| 215 |
+
|
| 216 |
+
**The shortfall itself is a real regression and its cause is unknown.** It
|
| 217 |
+
appears on both C1.2 and C1.3, so it belongs to the `--cal-axes` /
|
| 218 |
+
`--balance-exclude-big` pair rather than to the denominator augmentation. It is
|
| 219 |
+
the largest practical regression in this release.
|
| 220 |
+
|
| 221 |
+
### 2.4 Raising density squeezes drumrolls out β this is the compromise a user can hit
|
| 222 |
+
|
| 223 |
+
Median spans/min over eight oni songs, per requested density bucket
|
| 224 |
+
(authored **3.885**; `FINAL_CHAMPION.md` Β§4.3, 8 songs Γ 5 buckets, β):
|
| 225 |
+
|
| 226 |
+
| requested density | 0 | 3 | 7 | 11 | **15** |
|
| 227 |
+
|---|---:|---:|---:|---:|---:|
|
| 228 |
+
| C1 (s1234) | 3.629 | 3.055 | 4.153 | 2.818 | **2.971** |
|
| 229 |
+
| **C1.3** | 3.550 | 4.413 | 2.172 | 2.727 | **1.279** |
|
| 230 |
+
|
| 231 |
+
Ο(density β `span_per_min`) = **β0.759**, sign p 0.0078, firstβlast β2.443 /min.
|
| 232 |
+
At the top of the density request C1.3 delivers **33 %** of the authored span
|
| 233 |
+
rate; C1 delivers 76β79 %, C1.2 44 %.
|
| 234 |
+
|
| 235 |
+
* **Do not advertise the density knob at its top setting** until this is fixed.
|
| 236 |
+
* There is a named suspect: `cal_span` never left **0.1667** on either
|
| 237 |
+
challenger, so this is plausibly a calibration term that never trained rather
|
| 238 |
+
than an intrinsic cost of the knob.
|
| 239 |
+
* The mirror image is a *gain*: C1.3's `span_rate` knob no longer drags hit
|
| 240 |
+
density, which C1's does.
|
| 241 |
+
|
| 242 |
+
### 2.5 Absolute bucket calibration is still wrong on the density axis
|
| 243 |
+
|
| 244 |
+
| axis | exact bucket | within Β±1 | mean abs error | mean signed | `calibrated` |
|
| 245 |
+
|---|---:|---:|---:|---:|:--:|
|
| 246 |
+
| density | 5.0 % | 25.0 % | 4.125 | β1.075 | **false** |
|
| 247 |
+
| `big_rate` | **53.3 %** | **93.3 %** | **0.533** | +0.267 | **true** |
|
| 248 |
+
| `span_rate` | β | β | 1.483 | β | true |
|
| 249 |
+
|
| 250 |
+
The density knob orders the request correctly and lands in the wrong bucket,
|
| 251 |
+
under-shooting by ~1 bucket on average β more than C1.2 does. Rank control is
|
| 252 |
+
real; absolute density targeting is not delivered.
|
| 253 |
+
|
| 254 |
+
Note the cost that comes with the `big_rate` calibration win: C1.3 has the
|
| 255 |
+
**shortest** `big_rate` ladder of any arm (firstβlast +0.159 against C1's
|
| 256 |
+
0.183 / 0.224), so the top of that request now delivers less accent than C1's
|
| 257 |
+
did.
|
| 258 |
+
|
| 259 |
+
### 2.6 It fails its own span-quality gate and triggers a pre-registered kill condition
|
| 260 |
+
|
| 261 |
+
71 charts hard+oni (467 closed spans) and 26 charts easy+normal (272):
|
| 262 |
+
|
| 263 |
+
| axis | hard+oni | easy+normal |
|
| 264 |
+
|---|---|---|
|
| 265 |
+
| `duration_low_tail` | **fail** (rate verdict passes; balloon 0.250 beats = 0.75Γ the shortest authored balloon, roll 0.250 = 0.60Γ) | **fail** (balloon 0.250 beats = 0.20Γ the shortest authored balloon; 45/159 balloons below authored support) |
|
| 266 |
+
| `forced_close` | **fail** β **7 / 467 = 1.499 %** (0 unclosed, 7 clamp-forced) | **fail** β **7 / 272 = 2.574 %** |
|
| 267 |
+
| `over_span_max` | **fail** β one balloon 7.72 s = 1.19Γ the 6.5 s ceiling | **fail** β one balloon 7.58 s = 1.17Γ |
|
| 268 |
+
| `placement_fit` | **fail** β AUC **0.560** (CI [0.529, 0.593]) against a 0.56 floor; authored reference 0.668 | pass β AUC 0.604 |
|
| 269 |
+
| `orphans` / `swallowed` / `rate` / `duration` | pass | pass |
|
| 270 |
+
| **`a2_kill_condition_not_triggered`** | **fail β the kill condition FIRED** | fail |
|
| 271 |
+
|
| 272 |
+
The pre-registered kill condition (`SOFTCHART2_DESIGN.md` Β§4-A2) reads: *rate is
|
| 273 |
+
calibrated but placement/length quality is bad β concede that the bar-scoped
|
| 274 |
+
commitment mechanism is insufficient and escalate to an independent span-state
|
| 275 |
+
objective.* It fires on **all four arms**, including C1. Shipping does not
|
| 276 |
+
discharge it.
|
| 277 |
+
|
| 278 |
+
**`placement_fit` is a knife-edge, not a finding.** C1.2 passes at 0.561 and
|
| 279 |
+
C1.3 fails at 0.560 with CIs overlapping every other arm across their whole
|
| 280 |
+
width. One thousandth of AUC decides the verdict. It is reported, not used.
|
| 281 |
+
|
| 282 |
+
### 2.7 One span in seven is a micro-span
|
| 283 |
+
|
| 284 |
+
All generated spans, 97 charts, `FINAL_CHAMPION.md` Β§2 (β):
|
| 285 |
+
|
| 286 |
+
| | n spans | p1 | p5 | p50 | p90 | min | **share < 0.25 s** |
|
| 287 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 288 |
+
| authored | 762 | 0.177 | 0.273 | 0.963 | 2.419 | 0.083 | **3.8 %** |
|
| 289 |
+
| **this model** | 739 | 0.089 | 0.150 | 0.818 | 2.258 | **0.072** | **14.7 %** |
|
| 290 |
+
|
| 291 |
+
Per course: easy 6.6 %, normal 8.1 %, hard 13.7 %, oni **25.2 %** β authored
|
| 292 |
+
3.0 / 3.3 / 2.1 / 7.1 %.
|
| 293 |
+
|
| 294 |
+
* C1.3 does **not** inherit C1.2's worsening (16.4 %); 14.7 % is exactly at the
|
| 295 |
+
top of the C1 seed band (13.6β14.7 %). It is not an improvement either.
|
| 296 |
+
* **Mechanism is known and unfixed on every arm.** Duration bucket 0 is open at
|
| 297 |
+
the bottom; 123 of C1.3's 739 spans sit in it, at a median 0.5 beats against
|
| 298 |
+
the authored 0.917.
|
| 299 |
+
* **The shipped floor does not fix it.** `span_min_sec` / `span_min_beats` fired
|
| 300 |
+
263 times (`span_close_floor`) plus 33 carries across 97 charts and the
|
| 301 |
+
< 0.25 s share is still 14.7 %.
|
| 302 |
+
* **The mitigation is unbuilt.** `span_bucket0_profile` is `null` in this
|
| 303 |
+
package because `scripts/build_span_bucket0_profile.py` has never been run on
|
| 304 |
+
the train split. This is the shortest available fix and it is not in the box.
|
| 305 |
+
* **The gate cannot see it directly.** `a2_span_duration` compares p50/p90/p99
|
| 306 |
+
only and passes; `duration_low_tail` catches the support violation but not the
|
| 307 |
+
rate (its rate verdict passes on this arm).
|
| 308 |
+
* On a `/16` deployment grid the same quantity is 14.7 / 12.4 / **30.6** / 21.8 %
|
| 309 |
+
β hard is untouched by the augmentation and is the worst cell in the release.
|
| 310 |
+
|
| 311 |
+
### 2.8 The deployment grid has a two-role conflict that no checkpoint dissolves
|
| 312 |
+
|
| 313 |
+
`BAR_DENOM` is simultaneously an input the grid supplies and a **trained output
|
| 314 |
+
token** the model reads as a density announcement. `bar_denoms="supplied"`, the
|
| 315 |
+
shipped default, teacher-forces it.
|
| 316 |
+
|
| 317 |
+
13 songs Γ 4 courses, one seed, `--cond authored`, `C13_PREDICTION.md`:
|
| 318 |
+
|
| 319 |
+
| easy / normal / hard / oni | authored grid | **forced `/16`** | `/96` | `/24` |
|
| 320 |
+
|---|---|---|---|---|
|
| 321 |
+
| notes Γ· authored | 0.917 / 0.919 / 0.875 / 0.899 | **1.130 / 1.112 / 1.067 / 1.027** | 0.944 / 0.950 / 0.943 / 1.027 | 1.102 / 1.070 / 0.950 / 0.921 |
|
| 322 |
+
| onset F1 | 0.613 / 0.565 / 0.595 / 0.647 | **0.333 / 0.374 / 0.466 / 0.580** | 0.303 / 0.325 / 0.393 / 0.535 | 0.342 / 0.370 / 0.442 / 0.542 |
|
| 323 |
+
| genuine triplets (den 3/6/12/24) % | 1.13 / 2.48 / 6.34 / 9.43 | **0.00** (arithmetically impossible) | 3.45 / 5.78 / 10.57 / 11.46 | 14.13 / 20.31 / 28.57 / 38.70 |
|
| 324 |
+
| ultra-fine (den 48/96) % | ~0 | 0.00 | **11.53 / 13.22 / 15.62 / 6.60** | 0.00 |
|
| 325 |
+
| nesting (family Ξ² = 2) | 0.730 / 0.797 / 0.822 | 0.726 / 0.811 / 0.834 | 0.649 / 0.730 / 0.752 | 0.714 / 0.769 / 0.779 |
|
| 326 |
+
|
| 327 |
+
Authored triplet rates for reference: 4.20 / 5.72 / 7.43 / 9.95 %; ultra-fine
|
| 328 |
+
~0.00 %.
|
| 329 |
+
|
| 330 |
+
**The recommended deployment configuration is `--grid bpm --grid-bpm-denom 16`
|
| 331 |
+
with the package default `bar_denoms="supplied"`.** Its two costs, stated
|
| 332 |
+
plainly: **triplets are exactly 0 and cannot be otherwise**, and hard micro-spans
|
| 333 |
+
stay at 30.6 % against an authored 2.1 %.
|
| 334 |
+
|
| 335 |
+
Four consequences a deployer must carry:
|
| 336 |
+
|
| 337 |
+
* **Deployment costs 10β46 % of onset F1** relative to the authored grid:
|
| 338 |
+
easy β46 %, normal β34 %, hard β22 %, oni β10 %. Every `--grid authored`
|
| 339 |
+
number in this card is therefore optimistic for a user who supplies only BPM.
|
| 340 |
+
* **`p90 |offset|` is 0 ms on the authored grid and 14.6β36.7 ms on the bpm
|
| 341 |
+
grids.** That residual belongs to the bar edges; no denominator choice touches
|
| 342 |
+
it.
|
| 343 |
+
* **The model's off-dyadic rate is a property of the model, not the lattice.**
|
| 344 |
+
Given any 3-divisible grid it puts 2β3Γ the authored share off the dyadic
|
| 345 |
+
grid; the grid only decides whether that mass gets called "triplets" (`/24`)
|
| 346 |
+
or "ultra-fine jitter" (`/96`). This is a training-side defect and no
|
| 347 |
+
deployment grid fixes it.
|
| 348 |
+
* **The gate suite cannot score the deployment path.** `robustness.py:629-642`
|
| 349 |
+
raises `ContractError` on uniform-BPM arms, so every gate timing metric is
|
| 350 |
+
`null` for the path the model ships on; the `/16` figures above come from
|
| 351 |
+
`timing.py`, a second implementation that reproduces the gate on the authored
|
| 352 |
+
grid.
|
| 353 |
+
|
| 354 |
+
**The `deploy_bpm_grid` profile in this package is NOT the recommendation.** It
|
| 355 |
+
sets `bar_denoms="model"`, `bar_denom_mask="div96"`, per `DEPLOY_GRID_FIX.md`
|
| 356 |
+
Β§7 β a recommendation made explicitly conditional on C1.3 failing, which it did
|
| 357 |
+
not. That arm was measured on **C1 only**: density 1.06 / 1.08 / 0.97 / 1.01 and
|
| 358 |
+
nesting 0.635 / 0.706 / 0.778. **It has never been run on this checkpoint.** The
|
| 359 |
+
profile is kept because it is implemented and because the arm is the obvious
|
| 360 |
+
next experiment, not because it is advised here.
|
| 361 |
+
|
| 362 |
+
**One contradiction inside this package, stated so nobody has to find it.**
|
| 363 |
+
`config.json:serving_profiles.deploy_bpm_grid.basis` still quotes
|
| 364 |
+
`DEPLOY_GRID_FIX.md` Β§7 verbatim, including the words *"Recommended now"*, and
|
| 365 |
+
still carries C1's nesting figures. That string is generated by
|
| 366 |
+
`scripts/sc2_package.py`, which is read-only for this release. **This section
|
| 367 |
+
supersedes it.** The generated string is left intact rather than silently
|
| 368 |
+
diverging from the shipped code.
|
| 369 |
+
|
| 370 |
+
### 2.9 Playability floor: gaps no human hand can play
|
| 371 |
+
|
| 372 |
+
Recomputed on this arm's own events, 71 charts hard+oni, generated vs authored
|
| 373 |
+
on the same songs (producer `experiments/sc2_eval/ship_c13/playability.py`,
|
| 374 |
+
records `.../records/playability_all_arms.json`; β):
|
| 375 |
+
|
| 376 |
+
| | this model | authored |
|
| 377 |
+
|---|---:|---:|
|
| 378 |
+
| inter-onset gaps < 40 ms | **102** across **29** charts | 2 across 1 chart |
|
| 379 |
+
| β as a share of adjacent pairs | **0.296 %** (102 / 34 484) | **0.0053 %** (2 / 38 032) |
|
| 380 |
+
| gaps below the series record 29.4118 ms | **19** | 0 |
|
| 381 |
+
| shortest gap | **20.8 ms** (the serving floor) | 39.7 ms |
|
| 382 |
+
| peak burst in any 1 s window | **17 notes** | 15 notes |
|
| 383 |
+
|
| 384 |
+
C1 on the same population: 151 gaps (0.395 %), shortest 16.7 ms, peak 22
|
| 385 |
+
notes/s. C1.3 is better on every row and still **56Γ the authored rate**.
|
| 386 |
+
|
| 387 |
+
* `min_onset_gap_sec = 0.020` is a **degeneracy guard, not a fix**: it is set
|
| 388 |
+
below the fastest thing the series has shipped (29.4118 ms,
|
| 389 |
+
TAIKO-TONGUE-TWISTER oni, BPM 170, 48th notes) precisely so it cannot refuse a
|
| 390 |
+
chart the domain writes. It fired **125 times** across 97 charts with **0**
|
| 391 |
+
fail-opens. *A mask cannot fix a distribution.*
|
| 392 |
+
* The `generate.py` constant comment quotes the generated sub-40 ms rate as
|
| 393 |
+
~0.35 % against ~0.045 % authored. **That authored figure does not reconcile
|
| 394 |
+
with the 0.0053 % measured here**, and the two have never been put on the same
|
| 395 |
+
denominator. The excess is 8Γ on the comment's accounting and 56Γ on this one.
|
| 396 |
+
* **Good news that replaces a C1-card claim.** The C1 card reported the model
|
| 397 |
+
placing hits in 19.6 % of sub-0.25 s scaffolding bars. Under today's serving
|
| 398 |
+
floors this arm places **0 hits in all 112 such bars** across 97 charts, exactly
|
| 399 |
+
as the authored side does; `degenerate_bar_hits_blocked` fired 12 times with 0
|
| 400 |
+
fail-opens.
|
| 401 |
+
|
| 402 |
+
### 2.10 Pattern proxies regress against C1 on hard+oni
|
| 403 |
+
|
| 404 |
+
71 charts, one code version, C1 band = |s1234 β s4321| (`FINAL_CHAMPION.md` Β§3):
|
| 405 |
+
|
| 406 |
+
| metric (lower is better) | C1 mean | C1 band | **C1.3** | Γband | relative |
|
| 407 |
+
|---|---:|---:|---:|---:|---:|
|
| 408 |
+
| **`motif_reuse_gap`** | 0.3630 | 0.0042 | **0.4222** | +14.1 | **+16.3 %** |
|
| 409 |
+
| `compression_gap` | 0.0976 | 0.0006 | 0.1116 | +23.3 | +14.3 % |
|
| 410 |
+
| `motif_ref_marginal_js` | 0.0331 | 0.0046 | 0.0452 | +2.6 | +36.6 % |
|
| 411 |
+
| `ioi_js_per_chart` | 0.0295 | 0.0006 | **0.0352** | **+9.7** | +19.3 % |
|
| 412 |
+
| `ul_4gram_js_per_chart` | 0.2402 | 0.0055 | 0.2530 | +2.3 | +5.3 % |
|
| 413 |
+
| `class_4gram_js_per_chart` | 0.1337 | 0.0080 | 0.1357 | +0.2 | +1.5 % |
|
| 414 |
+
|
| 415 |
+
The `Γband` column overstates the case β several bands are under 1 % of their own
|
| 416 |
+
level β so the relative column is the honest one. On `motif_reuse` the bootstrap
|
| 417 |
+
CIs are nonetheless **disjoint**: C1 [β0.392, β0.334] / [β0.388, β0.334] against
|
| 418 |
+
C1.3 [β0.447, β0.396]. **The regression is real and not seed noise.**
|
| 419 |
+
|
| 420 |
+
Four bounds, all measured, none of them a dismissal:
|
| 421 |
+
|
| 422 |
+
* **It does not happen on easy+normal.** Every pattern proxy there is inside the
|
| 423 |
+
C1 seed band, `motif_reuse_gap` is marginally *better* than C1's mean, and
|
| 424 |
+
`class_4gram_js` is better on both challengers.
|
| 425 |
+
* **The decomposition puts the loss on rhythm and hand, while the accent layer
|
| 426 |
+
improves**: rhythm 0.1251 β 0.1640, hand 0.1774 β 0.2102, accent 0.0606 β
|
| 427 |
+
**0.0480**.
|
| 428 |
+
* **`ioi_js_per_chart` is where the augmentation shows up** β it reads the exact
|
| 429 |
+
rational IOI lattice, and it is C1.3's worst pattern metric relative to C1.2
|
| 430 |
+
(+9.7 band against +1.5).
|
| 431 |
+
* **The proxy reads mostly a channel carrying ~0.010 bits** about the authored
|
| 432 |
+
type (Β§2.2), and it has **never been validated against a listener**. It is
|
| 433 |
+
quoted because Β§4 names it, not because it is strong.
|
| 434 |
+
|
| 435 |
+
Realized `motif_reuse` is **0.3165** against an authored **0.7387** β under half
|
| 436 |
+
the authored repetition.
|
| 437 |
+
|
| 438 |
+
### 2.11 The note-type marginal, and one C1 finding that does NOT reproduce
|
| 439 |
+
|
| 440 |
+
Big notes are **1.48Γ the authored rate** pooled hard+oni (0.0873 vs 0.0589),
|
| 441 |
+
worst on the easiest course:
|
| 442 |
+
|
| 443 |
+
| course | authored | **C1.3** | C1 (s1234 / s4321) |
|
| 444 |
+
|---|---:|---:|---:|
|
| 445 |
+
| easy | 18.80 % | **26.22 % (1.40Γ)** | 32.68 % / 29.65 % |
|
| 446 |
+
| normal | 11.75 % | **17.41 % (1.48Γ)** | 19.21 % / 20.77 % |
|
| 447 |
+
| hard | 7.38 % | **10.61 % (1.44Γ)** | 13.11 % / 12.73 % |
|
| 448 |
+
| oni | 4.88 % | **7.40 % (1.52Γ)** | 9.29 % / 8.58 % |
|
| 449 |
+
|
| 450 |
+
The pre-registered stop window (0.045β0.075 pooled hard+oni) is **not met**.
|
| 451 |
+
|
| 452 |
+
**The C1 card's Β§2.8 claim does not reproduce here and is corrected.** On C1 the
|
| 453 |
+
realized class marginal was 13Γ closer to the training loss weights than to the
|
| 454 |
+
corpus. On C1.3 it is the other way round: JS(gen β authored) = **0.00588**
|
| 455 |
+
against JS(gen β class-weight prediction) = **0.00746**
|
| 456 |
+
(JS(authored β weights) = 0.01452). `--balance-exclude-big` moved the marginal
|
| 457 |
+
off the loss weights and toward the corpus. The log-log fit of realized share
|
| 458 |
+
against class weight collapses from slope 0.378 (Ο 0.436) on C1 to slope 0.078
|
| 459 |
+
(Ο 0.156) on C1.3.
|
| 460 |
+
|
| 461 |
+
### 2.12 Fine-lattice, tuplet and syncopation accuracy are much worse than average
|
| 462 |
+
|
| 463 |
+
`robustness.py` timing strata, 71 charts, exact-slot rate (overall **0.5840**):
|
| 464 |
+
|
| 465 |
+
| stratum | n authored | exact-slot | C1 (s1234) |
|
| 466 |
+
|---|---:|---:|---:|
|
| 467 |
+
| lattice step 11.6β23.2 ms | 114 | **0.342** | 0.456 |
|
| 468 |
+
| lattice step 23.2β46.4 ms | 790 | **0.443** | 0.465 |
|
| 469 |
+
| lattice step β₯ 92.9 ms | 24 228 | 0.594 | 0.651 |
|
| 470 |
+
| positions with a denominator divisible by 3 | 3 480 | **0.345** | 0.496 |
|
| 471 |
+
| syncopation band 0 β band 5 | 20 846 β 1 669 | **0.615 β 0.434** | 0.672 β 0.514 |
|
| 472 |
+
| local IOI 25β50 ms | 189 | **0.402** | β |
|
| 473 |
+
| BPM > 250 | 2 570 | **0.499** | β |
|
| 474 |
+
|
| 475 |
+
Every stratum is lower than C1's, which is the note-budget effect of Β§2.3 acting
|
| 476 |
+
on a per-stratum recall, but the *shape* is the finding: tuplet positions and
|
| 477 |
+
fast lattice steps are 0.24 below the overall rate, and accuracy decays
|
| 478 |
+
monotonically with syncopation.
|
| 479 |
+
|
| 480 |
+
### 2.13 Smaller, but on the record
|
| 481 |
+
|
| 482 |
+
| finding | number | n | source |
|
| 483 |
+
|---|---|---|---|
|
| 484 |
+
| `long_song_no_drift` fails | **2** of 21 metrics significant at BH q β€ 0.05 (`dens_signed_err`, `plan_dens_bucket_tvd`) β C1 fails 6 / 5 | 71 charts | `robustness.json` β |
|
| 485 |
+
| Harness-level generation failures | **6** of 71 song-courses refused (`#BRANCHSTART` and non-representable meters); **1** chart exported `gen_tja = null` β identical to C1 on the same population | 71 charts | `index.json` β |
|
| 486 |
+
| Slot-export off-lattice events | 8 across 97 charts (C1: 5) β the documented contract path, not a crash | 97 charts | decoder counters β |
|
| 487 |
+
| `escape_hatch` fires | 91 times across 97 charts (C1: 84) | 97 charts | decoder counters β |
|
| 488 |
+
| Generation wall time | 9.29 s/chart, 4.33 s per audio-minute β **contended**, the GPU was shared, quoted as a declaration only | 71 charts | β |
|
| 489 |
+
| Complete-bar rest placement | **not measured on this arm** (C1: Jaccard 0.474, recall 0.623) | β | β |
|
| 490 |
+
| Greedy-decode contrast | **not measured on this arm** | β | β |
|
| 491 |
+
| Micro-span seed sensitivity | **not measured on this arm** (C1: 24.0 % vs 32.4 % across two sampling seeds on 19 oni charts) | β | β |
|
| 492 |
+
|
| 493 |
+
### 2.14 The evidence base is thinner than it looks
|
| 494 |
+
|
| 495 |
+
* **One seed.** Every number in this card is n = 1 in the training seed. The C1
|
| 496 |
+
seed band quoted throughout is C1's, used as a reproducibility scale; it is a
|
| 497 |
+
range over n = 2 and carries no confidence statement.
|
| 498 |
+
* **easy and normal are evaluated on 13 songs.** The 71-chart population is
|
| 499 |
+
`hard` + `oni` only; the easy/normal population is 26 charts from 13 songs, and
|
| 500 |
+
its seed band is 4β5Γ wider than hard+oni's.
|
| 501 |
+
* Every difficulty-family and deployment-grid number is **13 songs, one seed**
|
| 502 |
+
(β). Every knob sweep is **8β12 songs Γ 5 buckets** (β).
|
| 503 |
+
* At these sizes a gate `pass` carries little information. Worked example from
|
| 504 |
+
this campaign: the `big_share_of_hits` gate **flips between two C1 training
|
| 505 |
+
seeds** while the failing seed's point estimate is *smaller*. At n = 36 that
|
| 506 |
+
gate cannot rank checkpoints, and no pass/fail on it should be quoted as
|
| 507 |
+
evidence for any arm.
|
| 508 |
+
* **Nothing here is audio-referenced** except `a2_span_placement_fit`. Every
|
| 509 |
+
other quantity is chart-vs-chart on the exact rational lattice or on hit order.
|
| 510 |
+
No timing window and no game judgement parameter appears anywhere in this path.
|
| 511 |
+
* **`best_val` is comparable across C1 / C1.2 / C1.3** (one training cache) and
|
| 512 |
+
**not** comparable to C1.1 or to any arm on a different cache. C1.3's 0.5769 is
|
| 513 |
+
**2.9 % worse** than C1.2's 0.5605; that was the registered trade and the
|
| 514 |
+
deployment side won it.
|
| 515 |
+
* **No plan-neutral arm was generated**, so the `plan_neutral_fallback` Β§4 clause
|
| 516 |
+
is `na` and the ledger is incomplete by one required clause.
|
| 517 |
+
|
| 518 |
+
### 2.15 Serving regime: what `motif_constraint: auto` resolves to, and why it matters less here
|
| 519 |
+
|
| 520 |
+
`motif_constraint: "auto"` resolves to **OFF** for this checkpoint: the training
|
| 521 |
+
cache's `barscript_md5` (`634e3dc3β¦`) differs from the serving `barscript.py`
|
| 522 |
+
(`9afc715aβ¦`), so the gold sequences it trained on never satisfied the MOTIF hard
|
| 523 |
+
constraint. `train_serve_matched = true` β OFF is the train/serve-matched choice
|
| 524 |
+
and it is the right default.
|
| 525 |
+
|
| 526 |
+
**Unlike the C1 package, this card needs no correction for it.** Every number in
|
| 527 |
+
this card was measured with the constraint **OFF**, i.e. under exactly this
|
| 528 |
+
package's resolved default; the 20 evaluation runs all record
|
| 529 |
+
`serving_fsm.motif_constraint: off, verified: true`. In particular
|
| 530 |
+
`motif_ref_marginal_js` here is **0.0452** (hard+oni) and **0.0114**
|
| 531 |
+
(easy+normal), both measured OFF.
|
| 532 |
+
|
| 533 |
+
Two things that still belong on the record:
|
| 534 |
+
|
| 535 |
+
* **The ON/OFF blast radius was measured on C1 only**, where switching the
|
| 536 |
+
constraint on moved `motif_ref_marginal_js` 0.0311 β 0.0189 (β39 %). On the C1
|
| 537 |
+
package the shipped default therefore makes the correct value **0.0311, not the
|
| 538 |
+
0.0189 in the older tables**. **That contrast has never been measured on
|
| 539 |
+
C1.3**, so no ON-constraint number should be quoted for this checkpoint.
|
| 540 |
+
* Serving under the constraint ON would be train/serve **mismatched** for this
|
| 541 |
+
checkpoint and is not a supported configuration.
|
| 542 |
+
|
| 543 |
+
---
|
| 544 |
+
|
| 545 |
+
## 3. Serving contract
|
| 546 |
+
|
| 547 |
+
Every value below is in `config.json:serving`, with its justification in
|
| 548 |
+
`config.json:serving_basis`. Pass them explicitly β the package's
|
| 549 |
+
`serving_kwargs()` does β so the recorded contract is the one that reaches the
|
| 550 |
+
decoder.
|
| 551 |
+
|
| 552 |
+
| parameter | value | one-line basis |
|
| 553 |
+
|---|---|---|
|
| 554 |
+
| `min_onset_gap_sec` | **0.020** | Degeneracy guard, **not** a corpus percentile. Must stay strictly below the series record of 29.4118 ms (BPM 170, 48ths). Two earlier corpus-derived values (0.0395, 0.0300) were both wrong. See Β§2.9. |
|
| 555 |
+
| `degenerate_bar_sec` | **0.25** | Authored charts place zero hits in any bar under 0.375 s (3 010 bars). Largest round threshold with zero counterexamples and 1.5Γ margin; identical to the gate's threshold. Blocks hits only, never span geometry. |
|
| 556 |
+
| `span_min_sec` / `span_min_beats` | 0.0833 / 1β3 | Authored population minima. Fail-open, counted. Does not fix Β§2.7. |
|
| 557 |
+
| `span_bucket0_profile` | **null** | Not built β see Β§2.7. |
|
| 558 |
+
| `family_bias` / `family_hand_bias` / `family_mode` | 2.0 / 1.5 / `bias` | **Provisional**, never calibrated β see Β§2.1. Family decode is opt-in. |
|
| 559 |
+
| `motif_constraint` | `auto` β resolves **off** here | Train/serve matching by `barscript.py` md5 β see Β§2.15. |
|
| 560 |
+
| `bar_denoms` / `bar_denom_mask` | `supplied` / `lattice` | The evaluated regime, and the recommended deployment regime on a uniform `/16` grid. The `deploy_bpm_grid` profile exists but is **not** recommended for this checkpoint β see Β§2.8. |
|
| 561 |
+
| `greedy` / `temperature` / `top_p` | false / 1.0 / 0.95 | As evaluated. |
|
| 562 |
+
| `plan_temperature` | `null` (follows `temperature`) | Never exercised in evaluation; shipped unset rather than tuned. |
|
| 563 |
+
|
| 564 |
+
`_has_sync` **must** be forced on at load. `train.py` records
|
| 565 |
+
`sync_token=False` while the training prefix carries the SYNC slot, so a loader
|
| 566 |
+
that trusts the recorded flag drops the slot and every sync bucket reaches the
|
| 567 |
+
decoder as an identical prefix. The package loader does this and records why in
|
| 568 |
+
`config.json:serving_prefix_fix`. This is a workaround; the fix belongs in
|
| 569 |
+
`train.py`.
|
| 570 |
+
|
| 571 |
+
### Recommended deployment configuration
|
| 572 |
+
|
| 573 |
+
```
|
| 574 |
+
--grid bpm --grid-bpm-denom 16 # uniform /16 grid
|
| 575 |
+
bar_denoms = "supplied" # the package default
|
| 576 |
+
motif_constraint = auto -> off # resolved at package time
|
| 577 |
+
family decode: optional; Ξ² = 2.0 / 1.5 is provisional
|
| 578 |
+
```
|
| 579 |
+
|
| 580 |
+
Expected behaviour under it, 13 songs Γ 4 courses, one seed: notes within
|
| 581 |
+
13 / 11 / 7 / 3 % of authored; onset F1 0.333 / 0.374 / 0.466 / 0.580;
|
| 582 |
+
**zero triplets**; hard micro-spans ~30 %.
|
| 583 |
+
|
| 584 |
+
---
|
| 585 |
+
|
| 586 |
+
## 4. Weights are bfloat16
|
| 587 |
+
|
| 588 |
+
Training ran bf16 autocast and CUDA inference runs bf16 autocast, so fp32
|
| 589 |
+
storage carried no information the forward pass could use β the argument the
|
| 590 |
+
v1.5 release made and verified. Measured cast cost: max absolute delta
|
| 591 |
+
**0.00711** on `frontend.2.weight` (0.349 % of that tensor's max) over 9 107 291
|
| 592 |
+
float elements.
|
| 593 |
+
|
| 594 |
+
This is an argument about the **compute** dtype, not a proof that the two
|
| 595 |
+
checkpoints decode identically. Sampling is chaotic in the logits, so individual
|
| 596 |
+
charts can differ. The paired fp32-vs-bf16 evaluation has **not** been repeated
|
| 597 |
+
for BarScript. The fp32 checkpoint is preserved in the source run directory and
|
| 598 |
+
is the reference for any bit-level comparison.
|
| 599 |
+
|
| 600 |
+
---
|
| 601 |
+
|
| 602 |
+
## 5. Provenance
|
| 603 |
+
|
| 604 |
+
| | |
|
| 605 |
+
|---|---|
|
| 606 |
+
| serving code | `generate.py` md5 `73b6e49b00fb60f3229c8539bc7a1029`; all 18 `src/softchart/*.py` md5s in `config.json:code.module_md5` |
|
| 607 |
+
| evaluation harness | `sc2_generate_eval_v2_2_serving_floors`, contract `sc2_eval_v2` |
|
| 608 |
+
| evaluation runs | 20 run indices, 1 028 charts, one `serving_floors` signature, `comparable: true`, 6/6 floor counters present |
|
| 609 |
+
| headline population | 71 charts hard+oni (36 songs) + 26 charts easy+normal (13 songs), `--cond authored`, `--grid authored`, seed 1 |
|
| 610 |
+
| deployment population | 13 songs Γ 4 courses Γ 10 arms, `--cond authored`, one serving tree, one signature |
|
| 611 |
+
| label spec | `spec.json`, field-for-field equal to the cache manifest's copy (`spec_parameters_match_cache_manifest: true`) |
|
| 612 |
+
| uploaded | **no** (`release_manifest.json: "uploaded": false`) |
|
| 613 |
+
|
| 614 |
+
Reports: `experiments/sc2_eval/SHIP_DECISION.md`,
|
| 615 |
+
`experiments/sc2_eval/FINAL_CHAMPION.md`,
|
| 616 |
+
`experiments/sc2_eval/C13_PREDICTION.md`,
|
| 617 |
+
`experiments/sc2_eval/C12_EVAL.md`, `experiments/sc2_eval/DEPLOY_GRID_FIX.md`.
|