computer-9b (run 0, step 120)

Computer-7 after 120 steps of online self-preference RL with no constitution: a frozen Computer-7 reads out, one token, which of 8 sibling turns it prefers (8 presentation rotations averaged); within-fork advantages train the policy (REINFORCE, token-level loss, KL to init). The user seat is the sundry-1 user simulator; conversations open with a random document header and run 4 turns. Length was allowed to move for the first 80 steps and was neutralised (pooled within-fork length slope removed from advantages) for steps 80โ€“120.

What moved by step 120, relative to Computer-7 (per-1k-word rates, all sampled turns): "I suppose" 5.1 โ†’ 2.2, "I think" 2.6 โ†’ 5.1, irrealis markers (would/might/perhaps/seem) 68 โ†’ 50, "(Note: โ€ฆ)" asides โˆ’25%, quote marks and parentheses per word back at the init rate after a mid-run rise; median turn length 135 โ†’ ~245 tokens. Per-token KL to init 0.010. The judge's read-out is stable across presentation orders (split-half r 0.78).

Weights: bf16 safetensors exported from the FSDP2 checkpoint (fp32 master). Same tokenizer and chat format as Computer-7 (**User:** โ€ฆ **Model C:** โ€ฆ plain-text turns under a document header; no chat template).

Downloads last month
15
Safetensors
Model size
71B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cosmicoptima/computer-9b

Finetuned
(12)
this model