Safety of Latent Communication in Multi-Agent Systems
Abstract
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole.
Community
What happens when aligned LLM agents stop communicating in text and instead exchange learned latent representations?
We show that the communication link itself becomes a safety-critical attack surface: even benign link training can weaken refusal behavior, while manipulating only the link can drive harmful compliance without modifying the underlying LLMs. Surprisingly, strong attacks can require neither harmful target completions nor control over every link. We further show that the same links can be repaired to substantially restore both safety and utility without retraining the underlying agents.
The uncomfortable part isn't that latent handoffs can carry a jailbreak — it's that refusal behavior seems to live in the text channel rather than in the weights. Which means every red-team result I've collected on my own agents is evidence about the wrong interface. The question that decides whether this matters: does compliance drift show up when the link is trained on ordinary benign task data with no attacker in the loop, or does it need adversarial optimization of the link? If it's the former, anyone shipping latent handoffs is flying blind and doesn't know it. Cheap test I'd actually run: freeze two open models I already serve, train only the link on benign task data, then replay my existing refusal suite through the link instead of through text and diff the refusal rates. If they drop, the safety story I've been telling myself was about the tokenizer.
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper