StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
Abstract
StarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.
We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration. StarHarness constructs a compact evolution pool by stratifying tasks according to baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution improves full-benchmark performance by 20-35 percentage points over the default harness after 4-12 accepted changes per environment. These gains persist on tasks excluded from evolution and transfer without re-evolution across GPT and Qwen model families. Trace analysis links the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, with fewer false-positive diagnoses and shorter trajectories in several settings. StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.
Community
On ITBench, Qwen3.5-27B with an evolved harness beats GPT-5.5 on the baseline harness by 19.2 points.
More capable models do not always make better enterprise agents.
With the right harness, a smaller open-weight model can outperform a larger frontier model running the default scaffold.
In our new preprint, we present StarHarness, a method for directly searching for an environment-optimal harness around a fixed model. It evolves prompts, tools, schemas, execution logic, state handling, skills, MCPs, subagents, context management, and more, while keeping model weights fixed.
Across three enterprise benchmarks, the evolved harness improved performance by 20–35 percentage points and reduced inference cost by 17–53%.
Our results suggest that some of the capability gap we attribute to models may actually come from how well the surrounding harness fits the environment.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MemoHarness: Agent Harnesses That Learn from Experience (2026)
- DarwinX: Evolving Agent Harnesses Through Natural Selection (2026)
- Rethinking the Evaluation of Harness Evolution for Agents (2026)
- Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories (2026)
- Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents (2026)
- HarnessOpt-Bench: Evaluating LLMs at Harness Optimization (2026)
- TTHE: Test-Time Harness Evolution (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.24804 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper