HuggingEnvs/watercolour-grpo-hps-only
Reinforcement Learning • Updated • 16
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
Note GRPO adapter for Qwen3.5-35B-A3B. Reward is 0.90 HPSv3 with the pairwise judge off.
Note The 178 paintings that define the reward, with the sketch that produced each one.
Note Every rollout of the HPS-only run: 470 paintings, sketches and rewards, by step.