andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 1, 2023 • 35.3k • 224 • 10
sabaridsnfuji/repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation Viewer • Updated Jul 25 • 1 • 62 • 3
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 11 days ago • 162
leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset Viewer • Updated Aug 12, 2025 • 14k • 121 • 10
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 19 days ago • 159
jaehyeokdoo2/openpi-droid-pnpcarrot-singetask-qflow-offlinerl-criticwarmup2000-alpha100-bs8-test Updated Mar 4 • 3
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 16 days ago • 41