andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 1, 2023 • 35.3k • 185 • 10
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding Paper • 2609.30670 • Published 7 days ago • 11
leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset Viewer • Updated Aug 12, 2025 • 14k • 105 • 9
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis Paper • 2609.29444 • Published 8 days ago • 20
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 9 days ago • 18
StudentSim: Training LLM-based Student Simulators Paper • 2609.01591 • Published about 1 month ago • 494
imflash217/proximal_policy_optimization_huggy_unity Reinforcement Learning • Updated Jan 14, 2023 • 87 • 3
MRNH/proximal-policy-optimization-LunarLander-v2 Reinforcement Learning • Updated Aug 10, 2023 • 16 • 1
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 16 days ago • 82
ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models Paper • 2609.13231 • Published 30 days ago • 18
RULER: Instance-aware Rubric Rewards for SVG Generation Paper • 2609.25270 • Published 11 days ago • 101