hanoz bhathena
bh9052
AI & ML interests
None yet
Recent Activity
updated a collection 4 days ago
Continual learning updated a collection 4 days ago
Continual learning updated a collection 4 days ago
Continual learning Organizations
None yet
Agent harness
Continual learning
-
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Paper • 2605.20025 • Published • 192 -
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 89 -
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 47 -
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
Paper • 2605.13941 • Published • 25
Evaluation
CUA
-
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 89 -
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
Paper • 2605.18652 • Published • 8 -
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
Paper • 2605.12481 • Published • 28 -
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
Paper • 2605.25624 • Published • 36
Post training
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 7 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 60 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 52
Foundation models
Evaluation
Agent harness
CUA
-
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 89 -
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
Paper • 2605.18652 • Published • 8 -
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
Paper • 2605.12481 • Published • 28 -
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
Paper • 2605.25624 • Published • 36
Continual learning
-
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Paper • 2605.20025 • Published • 192 -
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 89 -
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 47 -
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
Paper • 2605.13941 • Published • 25
Post training
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 7 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 60 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 52