Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention
Abstract
RO-PnR learns when to ask clarifying questions versus deliver corrections by weighing interaction costs against expected gains, improving cost-adjusted outcomes in health misinformation dialogues.
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.
Community
Should an AI ask a clarifying question before correcting health misinformation or respond immediately?
We introduce Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth the interaction cost. By accounting for differences in users’ health literacy and belief commitment, RO-PnR delivers more adaptive interventions while using approximately 30% fewer turns than always-probe approaches.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs'Clarification Policies (2026)
- Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation (2026)
- Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations (2026)
- ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors (2026)
- Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising (2026)
- Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives (2026)
- Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.21721 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper