Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 5 days ago • 59
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 7 days ago • 44
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning Paper • 2609.35505 • Published 6 days ago • 24
RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation Paper • 2609.18703 • Published 18 days ago • 54
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings Paper • 2609.25165 • Published 13 days ago • 77
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 18 days ago • 30
Emergent Collusion in Long-Horizon LLM Agent Interaction Paper • 2609.24967 • Published 13 days ago • 19
xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning Updated Dec 9, 2025 • 240 • 11
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 13 days ago • 131
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 16 days ago • 35
wnkh/vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-RLHF-v6 Viewer • Updated Jul 10, 2025 • 100 • 56 • 4
wnkh/vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-original-only Viewer • Updated Oct 18, 2025 • 244 • 50 • 4
LangAGI-Lab/magpie-reasoning-v1-10k-step-by-step-rationale Viewer • Updated Jan 31, 2025 • 10k • 108 • 4
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 19 days ago • 815
CreitinGameplays/magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtoken-mistral Viewer • Updated Feb 9, 2025 • 10k • 38 • 2