HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 3 days ago • 116
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 10 days ago • 335
Codifying the Judge: Scalable Evaluation via Program Distillation Paper • 2607.22561 • Published May 29 • 8
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published Jul 18 • 143