Online Learning with LLM Experts from Limited Feedback
Abstract
Adaptive routing of prompts to LLM experts is formulated as a contextual bandit problem with limited feedback, yielding algorithms with sublinear regret bounds and effective routing strategies.
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with K actions that represent experts and d features that encode prompts, over a horizon of T rounds. We propose algorithms that strategically select and observe rewards to minimize regret. In the full-information setting, we achieve a regret of O(d T / m), while in the bandit setting we achieve O(d T K / m), where m ll T is a budget on feedback. Our experiments show that we efficiently learn high-quality routing strategies across diverse LLMs from limited feedback.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- WISERouter: LLM Routing with Workload Budget Constraint (2026)
- Coverage-Maximizing Multinomial Subset Routing under Operational Constraints (2026)
- Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation (2026)
- Generalized Linear Bandits with Memory (2026)
- Lipschitz Bandits with Arbitrary Feedback Delays (2026)
- Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits (2026)
- Top-k Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.05820 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper