The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 12 days ago • 12
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 6 days ago • 50
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 30 days ago • 106
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 30 days ago • 125
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Paper • 2604.14142 • Published Apr 15 • 30
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Paper • 2603.24440 • Published Mar 25 • 99
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation Paper • 2603.18886 • Published Mar 19 • 6
Grounding and Enhancing Informativeness and Utility in Dataset Distillation Paper • 2601.21296 • Published Jan 29 • 21
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence Paper • 2511.18538 • Published Nov 23, 2025 • 307
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique Paper • 2511.09067 • Published Nov 12, 2025 • 2
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models Paper • 2510.27196 • Published Oct 31, 2025
Grounding Computer Use Agents on Human Demonstrations Paper • 2511.07332 • Published Nov 10, 2025 • 107
SPICE: Self-Play In Corpus Environments Improves Reasoning Paper • 2510.24684 • Published Oct 28, 2025 • 18
InstructCoder: Empowering Language Models for Code Editing Paper • 2310.20329 • Published Oct 31, 2023 • 2
MMCode: Evaluating Multi-Modal Code Large Language Models with Visually Rich Programming Problems Paper • 2404.09486 • Published Apr 15, 2024 • 2