-
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
Paper • 2508.08221 • Published • 50 -
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Paper • 2504.20571 • Published • 99 -
RLPR: Extrapolating RLVR to General Domains without Verifiers
Paper • 2506.18254 • Published • 35
Igor Kilbas
kaleinaNyan
AI & ML interests
Computer Vision, NLP
Recent Activity
liked a model about 2 months ago
prism-ml/Ternary-Bonsai-27B-gguf liked a dataset 9 months ago
Skywork/Skywork-OR1-RL-Data liked a dataset 9 months ago
t-tech/ruAIME-2025Organizations
None yet