Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published 5 days ago • 26
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus Paper • 2606.15345 • Published Jun 13 • 16
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale Paper • 2409.15637 • Published Sep 24, 2024
Implicit Personalization in Language Models: A Systematic Study Paper • 2405.14808 • Published May 23, 2024
Automatic Generation of Model and Data Cards: A Step Towards Responsible AI Paper • 2405.06258 • Published May 10, 2024
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data Paper • 2410.16491 • Published Oct 21, 2024 • 2
Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba Paper • 2406.12754 • Published Jun 18, 2024
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding Paper • 2412.17729 • Published Dec 23, 2024