Papers
arxiv:2606.05976

The Self-Correction Illusion: LLMs Correct Others but Not Themselves

Published on Jun 4
Authors:
,
,

Abstract

Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a capability deficit or a role-label artifact: does an agent's willingness to correct a wrong claim depend causally on the chat-template role that carries it, rather than on the claim's content? Our setup keeps the erroneous claim byte-identical across all conditions (SHA-256 verified) and varies only its wrapping role: the agent's own <thought>, a user message, a tool response, or a system <memory> block. Across 13 model-domain cells covering seven model families and three domains (n{=}30 paired tasks per cell), relabeling the claim from <thought> to an external role lifts the explicit-correction rate by 23 to 93 percentage points, with 10 of 13 cells reaching p{<}0.001. Further experiments confirm that the effect is asymmetric, mechanistically decomposable, and robust across domains. The failure to self-correct is not a cognitive deficit; it is a chat-template artifact. We exploit this artifact by designing a prompt-structure-only intervention that requires no training and no model modification, with its strongest role label being domain-dependent: <memory> dominates on math, while a plain user message dominates on logical deduction.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.05976
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.05976 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.05976 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.