Update Space to n=5 run (2026-07-19): ranking-led report, /30 scores, 30 runs, appendix at bottom 3dd0b08 verified rufimelo commited on Jul 19
Invert scale: quality score 1-18, higher = better (report + both charts) c167fbf verified rufimelo commited on Jul 19
Move pass/fail appendix to the very bottom, after the Runs section 1246248 verified rufimelo commited on Jul 19
Charts: add footnote that scoring is joint (per-attempt judging is brittle) 221de5a verified rufimelo commited on Jul 19
Explain why we judge together: single-judge scoring is brittle with current rubric 68082ce verified rufimelo commited on Jul 19
Charts now use the single-judge ranking (mean rank + spread), not pass/fail 4872861 verified rufimelo commited on Jul 19
Rewrite report: self-contained, PM- and scientist-readable; precise per-trial framing f08aba6 verified rufimelo commited on Jul 19
Report: lead with single-judge ranking, demote pass/fail to appendix ad74a2f verified rufimelo commited on Jul 19
Render PDF pages as inline images (fixes PDF viewing on static host) cafb846 verified rufimelo commited on Jul 19
Convert to static Space (no runtime): report + charts + 18 run PDFs/verdicts af0d989 verified rufimelo commited on Jul 19
Add Gradio app + deep-run artifacts (report, charts, 18 PDFs + verdicts) 8d7aad4 verified rufimelo commited on Jul 19