-
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
Paper • 2510.01132 • Published • 6 -
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Paper • 2604.13602 • Published • 32 -
GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
Paper • 2410.05254 • Published • 85
Yashaswi Sharma
yashu2000
AI & ML interests
Large Language Models, Vision Models, Applied Models, MLOPS, Differentiable Economics, Differentiable Physics
Recent Activity
updated a collection 1 day ago
Thesis Papers updated a collection 1 day ago
Thesis Papers liked a dataset 28 days ago
lambda/hermes-agent-reasoning-tracesOrganizations
None yet