CISPA: latent agent links raise harmful compliance from 4.4 to 31.1
Swapping text messages between AI agents for trained latent links raised a two-agent system’s harmful-compliance score from 4.4 to 31.1, with no attacker involved and every model parameter frozen. That’s the headline result in Safety of Latent Communication in Multi-Agent Systems, posted by three researchers at the CISPA Helmholtz Center for Information Security in Saarbrücken. The links were trained only on benign data, so nobody poisoned anything.
Latent communication lets agents hand each other internal representations instead of generated text. A small trainable link maps the sender’s hidden states into soft tokens inside the receiver’s embedding space, which means both models stay frozen. The appeal is cost. In the paper’s two-agent setup, tokens per query fell from 729 to 511 and inference time from 2.65 seconds to 1.96, while MATH500 accuracy rose from 51.5% to 65.3%.
But the same training step that bought those savings moved the safety numbers the other way. Averaged across HarmBench, StrongREJECT, JailbreakBench and AdvBench, compliance with harmful requests rose 26.7 points in the two-agent system, 28.5 points in a three-model sequential chain, and 8.3 points in a mixture where two expert models feed a summariser. By our calculation the two-agent jump is a factor of about seven.
| Topology | Channel | Harmful compliance (avg) | MATH500 | Tokens |
|---|---|---|---|---|
| 2-agent | Text | 4.4 | 51.5% | 729 |
| 2-agent | Latent | 31.1 | 65.3% | 511 |
| Sequential | Text | 3.2 | 58.3% | 1042 |
| Sequential | Latent | 31.7 | 59.5% | 501 |
| Mixture | Text | 12.5 | 80.6% | 1123 |
| Mixture | Latent | 20.8 | 76.8% | 532 |
The authors trace that effect to refusal initiation, because the link changes how often the receiver starts by saying no. The receiver alone starts with a refusal 82% of the time, and 84% under text communication, but only 29% once the link is trained. Brief interventions claw it back: forcing Sorry as the first generated token cuts the compliance score from 31.1 to 1.6, and a five-token refusal prefix takes it to zero. A similar gap between what a model can see and what it refuses turned up in IIT Delhi’s work on vision-language refusals.
Deliberate manipulation goes much further still. Injecting 212 harmful query and response pairs from PKU-SafeRLHF into 1,904 benign examples, a poisoning rate of roughly 10%, lifted the mean score from 27.9 to 58.2 without the attacker touching the link parameters at all. A reinforcement-learning attack built on GRPO and an LLM judge reached 76.9.
That reinforcement-learning attack is the part worth reading twice. It needs no harmful target responses, and unlike the supervised version it kept the system useful: average MATH500 accuracy rose from 67.2% to 71.0% as compliance climbed to 76.9. Direct supervised optimisation scored a comparable 75.6 but crashed MATH500 to 30.7%. So the stronger attack is the one that a utility benchmark won’t flag, which matters because safety benchmark rankings are already unstable on small models.
Retaining performance on benign tasks therefore does not rule out substantial harmful compliance in the same system.
Huzaifa, Mavali and Eisenhofer, Safety of Latent Communication in Multi-Agent Systems
Repair worked, but it came with a bill. Turning the same reward machinery toward refusals pulled the aggregate score from 70.3 to 4.8 across nine attack and topology pairs, below the clean links themselves, and the authors published the code that produced it at Muhammad-Huzaifaa/latent-safety. Even so, repair reduced accuracy in all three reward-guided settings, and average GPQA-Diamond landed at 33.6% against a clean-link baseline of 36.4%.
None of this appears in the efficiency work that sells the technique. Vision Wormhole, revised on 28 September, reports a 6.0 percentage point average accuracy gain over text-mediated multi-agent systems and a 1.69x geometric-mean speedup across four VLM families, six team configurations and nine reasoning benchmarks. Its 32 pages and 16 tables carry no harmful-compliance measurement, and its ethics statement says evaluation is confined to those benchmark tasks.
Other groups stitching frozen models together are at least reporting their limits. NinaXander cut a Transformer key-value cache by 84.4% by splicing the first 5 layers of Pythia onto 27 layers of RWKV, though no composed model matched Pythia on multiple-choice accuracy. The question for the next latent-communication paper is whether it prints a jailbreak score next to the speedup.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
