JonusNattapong/Reinforcement-Learning-for-Gold-Trading-Model Reinforcement Learning • Updated Dec 23, 2025 • 95 • 20
VisionHOPE: Visual Backbones as Self-Modifying Learning Systems Paper • 2609.33325 • Published 5 days ago • 307
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 8 days ago • 13
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding Paper • 2609.30670 • Published 7 days ago • 11
saurabh5/saurabh5-rlvr_acecoder_filtered-offline-results-full-chunk-0 Viewer • Updated Jul 5, 2025 • 10k • 109 • 3
Anandbheesetti/Lunar_Lander_By_using_reinforcement_learning Reinforcement Learning • Updated Nov 4, 2023 • 11 • 3
jarguello76/reinforcement_learning_lunar_landing Reinforcement Learning • Updated Aug 17, 2025 • 4 • 8
Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents Paper • 2609.29892 • Published 8 days ago • 32
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds Paper • 2609.30199 • Published 8 days ago • 28
crumb/bloom-560m-RLHF-SD2-prompter-aesthetic Text Generation • 0.6B • Updated Mar 19, 2023 • 301 • 26