| license: llama3.1 | |
| base_model: cosmicoptima/computer-7 | |
| tags: [computer, model-c, self-preference, rl] | |
| # computer-9 | |
| Computer-7 after 340 steps of online self-preference RL. A frozen Computer-7 read out which of eight sibling turns it preferred, under a four-line constitution for steps 0–160 and six weighted frames after that; within-fork advantages trained the policy (REINFORCE, KL to init). The user seat was the `sundry-1` simulator. | |
| Previously published as `computer-run1-step340`. Earlier points on the same run: computer-9c, 9d, 9e (steps 100/120/160) and computer-run1-step180–240. | |
| Format: same as the other Computers. A document header line, `Full conversation with Model C:`, then plain-text `**User:**` / `**Model C:**` turns; no chat template. Sample at temperature 1.0, top-p 0.98, stop on `\n\n**User:**`. Weights are bf16 safetensors. | |