Papers
arxiv:2608.24804

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

Published on Aug 25
· Submitted by
Esakkivel Esakkiraja
on Aug 31
Authors:
,
,
,
,
,

Abstract

StarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.

We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration. StarHarness constructs a compact evolution pool by stratifying tasks according to baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution improves full-benchmark performance by 20-35 percentage points over the default harness after 4-12 accepted changes per environment. These gains persist on tasks excluded from evolution and transfer without re-evolution across GPT and Qwen model families. Trace analysis links the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, with fewer false-positive diagnoses and shorter trajectories in several settings. StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.

Community

Paper author Paper submitter
edited about 21 hours ago

On ITBench, Qwen3.5-27B with an evolved harness beats GPT-5.5 on the baseline harness by 19.2 points.

More capable models do not always make better enterprise agents.

With the right harness, a smaller open-weight model can outperform a larger frontier model running the default scaffold.

In our new preprint, we present StarHarness, a method for directly searching for an environment-optimal harness around a fixed model. It evolves prompts, tools, schemas, execution logic, state handling, skills, MCPs, subagents, context management, and more, while keeping model weights fixed.

Across three enterprise benchmarks, the evolved harness improved performance by 20–35 percentage points and reduced inference cost by 17–53%.

Our results suggest that some of the capability gap we attribute to models may actually come from how well the surrounding harness fits the environment.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.24804
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.24804 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.24804 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.24804 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.