Papers
arxiv:2609.08183

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

Published on Sep 8
· Submitted by
JarvisPei
on Sep 9
#1 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

NeoHorse-1 uses agentic post-training with intelligent routing, structured feedback loops, and curriculum-based distillation to improve model capabilities across agent benchmarks.

Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.

Community

NeoHorse-1 offers an interesting look at how agentic systems can improve through recursive learning and better routing. In gaming, 3 Patti Land also focuses on creating a smooth and engaging experience for card-game players. Those interested in 3 Patti Land can visit https://3patti-land.pk/ to explore the game and its features.

NeoHorse-1: Towards Recursive Self-Improvement highlights an interesting approach to improving AI agents through agentic post-training and adaptive routing. In a different digital space, 3 Patti Vegas is a gaming-related website focused on providing an engaging online card-game experience for players. Those interested in 3 Patti Vegas can learn more at https://3pattivegass.pk/.

·

Thanks for your comment. The parallel is actually an interesting one — card games like 3 Patti and agent training share a common loop: read the table (gather evidence), adjust your strategy (update constraints), act (execute), then use the outcome as feedback. That observe–judge–act–verify cycle is exactly what Agentic Post-Training in NeoHorse-1 tries to strengthen in the model.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.08183
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 10

Browse 10 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.08183 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 1