CAFE: Self-Improving Search Agents Need Co-Evolving Feedback Paper • 2608.24794 • Published 22 days ago • 6
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening Paper • 2605.19597 • Published May 19 • 22
DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training Paper • 2602.05890 • Published Feb 5 • 1
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation Paper • 2506.04078 • Published Jun 4, 2025 • 1
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models Paper • 2508.05452 • Published Aug 7, 2025
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment Paper • 2601.01576 • Published Jan 4 • 19
PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts Paper • 2503.06706 • Published Mar 9, 2025