LLM Groups Can Overstate Consensus and Settle on Wrong Answers
Three arXiv preprints examine when LLM collectives reach reliable consensus—and why interaction design matters as much as reasoning.
Three new arXiv preprints probe what happens when language models talk to each other. The first replays 100 held-out human Wason reasoning groups with matched LLMs and finds that the language-model groups overstate consensus. The authors argue that full-consensus rates are not stable indicators of collective cognition: they shift depending on how participation and final states are defined.
The second study similarly treats consensus as fragile. It shows that LLM collectives can converge on a wrong answer even when the available evidence leans the other way. Key transition points depend on how many messages an agent can actually read, and on the exact wording of claims. In this view, communication constraints and phrasing—not just underlying reasoning—shape whether a group finds truth.
The third paper takes a different angle: it looks at LLM agents collaborating under information asymmetry, where agents have different pieces of knowledge. Its emphasis is on communication and verification as drivers of successful collaboration, not on consensus rates. Together, the three studies agree that group outcomes are heavily influenced by interaction design; they differ on whether the central problem is measuring consensus or enabling cooperation.
Sources · 3
- Language-model groups overstate consensus when replaying human deliberation on a reasoning task
- Message capacity and claim wording set the transition points of collective truth-finding in language-model networks
- Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry
More in Research Digest
Can LLM Agents Design Chips From Higher-Level Abstractions?
A new preprint asks whether large language model agents can outperform RTL-level approaches by designing chips from higher-level abstractions.
Research Digest: Memory and Cooperation in Multi-Agent Vision
New papers explore how vision-language agents can share memory and arbitrate roles, while other work tackles compact representations and multi-channel imaging.
New Papers Probe the Hidden Costs and Risks of LLM Reasoning Traces
Six recent arXiv papers examine what happens inside chain-of-thought reasoning, showing that intermediate traces can be a liability as much as a capability.
New AI Research Spans Networks, Economy, Art, and Tools
Five independent papers highlight AI's expanding footprint from network optimization to cultural critique.