agi-noobs/chess-sft-20k-llm-reasoning-enriched-dpo-hard-negatives-v1 Viewer • Updated Dec 30, 2025 • 1.57k • 48 • 4
ticoAg/llm-complex-reasoning-train-qwen2-72b-instruct-correct Viewer • Updated Aug 8, 2024 • 7.11k • 62 • 5
Yuhan123/vicuna-13b-self_consistency_random_var_3 Text Generation • 13B • Updated Mar 14, 2025 • 11 • 5
s-emanuilov/LLMBG-Llama-3.1-8B-BG-Reasoning-v0.1 Text Generation • 8B • Updated Feb 9, 2025 • 172 • 14
Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure Paper • 2609.23551 • Published 12 days ago • 7
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 11 days ago • 156
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 11 days ago • 130
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation Paper • 2609.20758 • Published 15 days ago • 4
StudentSim: Training LLM-based Student Simulators Paper • 2609.01591 • Published about 1 month ago • 494
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 14 days ago • 138
xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning Updated Dec 9, 2025 • 148 • 11
Yuhan123/vicuna-13b-self_consistency_neg_exp_var_5 Text Generation • 13B • Updated Mar 14, 2025 • 107 • 8
Yuhan123/vicuna-13b-self_consistency_neg_exp_var_4 Text Generation • 13B • Updated Mar 14, 2025 • 107 • 7
Learning Foresight without Explicit Trajectories for 3D Diffusion Policies Paper • 2609.20669 • Published 15 days ago • 9