# Execution notes The final measurements were made on 2026-09-16. No model training or weight saving was performed. GPU measurements used transient L40S and H100 benchmark jobs. Initial diagnostic measurements were retained under `results/exploratory/`. The final engine adds JSON Schema meta-validation at request time so malformed enum definitions are rejected before inference. Final benchmark runs include that overhead equally in both methods. The final harness additionally records full dependency inventories and includes the separate scaling suite. The initial and final task labels, model revision, precision, prompts and candidate-scoring method are unchanged. During validation, a test expecting ValueError for an unsupported `allOf: []` schema instead received jsonschema.SchemaError because the schema itself was invalid. This was a test-fixture issue, not a model/cache failure. The corrected test uses a valid but unsupported `allOf: [{}]` schema and separately tests rejection of a malformed enum schema. The failed GPU jobs aborted before benchmarking. All final device checks pass. The larger Mac probes show sizeable per-repeat variation. We retain every measurement and report means without discarding outliers or selecting faster runs. The final tables use exactly the final run per device/suite, not the faster of the exploratory and final measurements. Inspect raw latency rows before interpreting the aggregated ratios. Final GPU source hashes match the packaged inference and benchmark code. Numerical comparison tolerances account for FP16 differences from batch and kernel ordering. No optimized causal-conv1d kernel was installed; both paths use the reference implementation. This can materially affect absolute latency and relative speedups versus a fully optimized serving stack. The 28-field constrained decision differs by one field on L40S compared with MPS/H100 (64.3% versus 60.7% field accuracy). FP16 decision margins can be sensitive to backend numerical differences. Each device was internally consistent across its three repeats; bitwise cross-device equivalence is not claimed. After the benchmarks, the original model weights, configuration and tokenizer were bundled byte-for-byte from the same pinned revision. This packaging update does not change the measured inference code, model parameters, or results. SHA-256 checksums are recorded in `BASE_MODEL_MANIFEST.json`.