X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
Abstract
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less fully than its content allows. Recent agents do use that structure, but only as LLM-written skills in context, never in the weights, so their gains do not generalize beyond retrieval. We instead recover this hierarchy from the data itself and train on it, with no LLM calls. Following text tokenizers, which build a vocabulary by counting alone, we score action spans by reusability and merge canonicalized actions into a reusable eXperience tree (X-Tree). Each X-Tree node captures how a frequent and success-bearing skill is composed from sub-skills, guiding efficient generalization. We integrate X-Tree into three training settings: offline RL, with each node as a training instance; online RLVR, with an adaptive skill bonus; and on-policy self-distillation, with X-Tree as the self-teacher's privileged context. Across WebArena, ScienceWorld, and WebShop at three model scales, X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop. Matched analyses attribute the gains to the X-Tree structure and the three integrations.
Community
Instead of training agents on flat action sequences, we mine the reusable hierarchy latent in existing trajectories, with zero LLM calls, and use it to guide training as data, as reward and as context.
Takeaways
- Agent trajectories hide a reusable hierarchy. The same routines, such as filling two fields and submitting, or walking to the kitchen to fetch a thermometer, recur across tasks and sites, yet flat training relearns them each time.
- The hierarchy can be mined by counting, with zero LLM calls. X-Tree merges canonical actions by how reusable the result is, the way a tokenizer builds a vocabulary, so the tree is deterministic, auditable and reproducible.
- Training on the tree beats training on flat sequences, at the same data and budget. It works as data, as reward and as context: +4.5 SR on WebArena, +4.9 SR on ScienceWorld, and as good as an LLM-written skill bank.
The same routines, again and again
Flat training relearns every routine from scratch. SFT and RL weight every action the same, so a model learns “fill a date, fill a date, click apply” anew each time it sees it, on every site where it appears. Experience is expensive: every trajectory needs an environment, a task and a verifier that can be trusted, and none of that scales the way web text does.
The structure is already in the data. Below is a real WebArena trajectory with the tree mined from 7,974 Go-Browse trajectories. Its last three steps form the node S3, fill two fields and submit, which occurs 411 times in the corpus: as a date filter on the admin panel and as a route query on the map. With the navigation node before it, it forms S127, which filters a table by a date range. The explorer below opens on this trajectory with S127 selected: click fill_two_and_submit in the tree to highlight its three steps, or switch to ScienceWorld to see one node cover the whole trip to fetch the thermometer.

Counting, not prompting
- A tokenizer for actions. Language models already solve a version of this problem before training starts: a tokenizer builds a vocabulary by repeatedly merging the pair that co-occurs most often. We apply the same idea to actions, with two changes.
First, actions are made comparable. Each raw action becomes a typed token such as type⟨date⟩ or click⟨button⟩, and the element ids and values move into slots, so two date filters on different pages read as the same symbol. Second, merges are chosen by reusability, not frequency alone. We score an adjacent pair by how often it recurs, how long the merged span is, and how often it appears in successful episodes. We call this score X-Score. A merge must also shorten the corpus by more than it costs. The merges stack into a tree: a reusable experience tree, or X-Tree.
- Other agent systems ask an LLM to write skills instead. Those skills stay in the prompt, never enter the weights, and cannot be rebuilt from the data. X-Tree is deterministic, auditable, and trained into the model.
Why the tree helps training
As data, a node is a dense, graded target. With only trajectories and no environment, each node becomes an RL instance. The policy continues from the gold steps before the node and earns credit for every matched step, plus a completion bonus that grows with depth, instead of a single reward at the end of the episode.
As reward, the tree gives a signal when the verifier is silent. Early in RL, a group of rollouts often has no success, so outcome rewards are all equal and GRPO learns nothing. Every mined skill a rollout executes earns a bonus, which separates the rollouts. The bonus fades as successes appear.
As context, the tree tells the teacher which routine to follow. In on-policy self-distillation the model is its own teacher. Given a rendered X-Tree card, the teacher raises the tokens of the next routine step, and the student is pulled toward it on exactly those tokens. The card replaces the LLM-written skill bank such methods normally use.
The demo below walks through each integration on real tasks. The panel on the right is the hint: the part of the X-Tree the rollouts are matched against, in the same colors.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents (2026)
- EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents (2026)
- HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory (2026)
- Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents (2026)
- AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning (2026)
- LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks (2026)
- MemoryWalker: Stop Training Agents on Contexts They Never Saw (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.32993 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper

