-
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
Paper β’ 2510.01132 β’ Published β’ 6 -
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Paper β’ 2604.13602 β’ Published β’ 32 -
GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
Paper β’ 2410.05254 β’ Published β’ 85
Yashaswi Sharma
yashu2000
AI & ML interests
Large Language Models, Vision Models, Applied Models, MLOPS, Differentiable Economics, Differentiable Physics
Recent Activity
updated a collection about 10 hours ago
Thesis Papers updated a collection about 10 hours ago
Thesis Papers liked a dataset 27 days ago
lambda/hermes-agent-reasoning-tracesOrganizations
None yet