Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Paper • 2608.12781 • Published • 35
None defined yet.
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward