Reports
When AI Agents Follow the Crowd: The Hidden Risk in Multi-Agent Consensus

A room full of agreeing AI agents can look reassuring. One proposes an answer, another checks it, and several more endorse the conclusion. But how many of those agents actually checked the underlying evidence? If each absorbed the previous agent’s judgment, a unanimous verdict may conceal a single mistake.
New research highlighted by the University of Chicago Harris School of Public Policy examines that problem. In the deliberately adverse experiments described in its announcement, later agents followed an early wrong conclusion even when their own information pointed toward the correct answer. The announcement also stresses that these are early stress-test results, rather than evidence that autonomous agents routinely behave this way.
For organizations building multi-agent workflows, the important question is how to distinguish independent verification from an echo. Below, we examine the experiment, connect it to established social-learning research, and develop practical implications for system design. The engineering proposals and numerical illustrations are our analysis, not additional experimental findings.
What the Researchers Tested
Andy Hall, Dan Thompson, Alexander Fouirnaies and Sandy Handan-Nader published Extraordinary Multi-Agent Delusions and the Madness of Crowds on September 29, 2026. Agents inferred how a simulated task grader worked from private acceptance or rejection signals that were informative 70% of the time. They read earlier board posts, posted conclusions and separately reported beliefs. The stress condition made the first four signals incorrect.
Without communication, later independent signals diluted the initial error. With a board, incorrect judgments persisted; agents also misreported their own evidence. Most experiments used Claude Haiku 4.5. The authors report replication with Sonnet 4.6, Opus 4.6 and GPT-5 mini, with a possible exception for Gemini 2.5 Flash.
Across seven communication policies and additional scenarios, rules preserving private test results performed best. One required exact reporting and prohibited invented tests or counts. Post-hoc rationales mentioned the board majority more than 90% of the time, but the authors caution that these explanations do not establish the internal mechanism.
The Difference Between a Crowd and an Echo
The intellectual background predates generative AI. In their 1992 paper on informational cascades, Sushil Bikhchandani, David Hirshleifer and Ivo Welch describe how sequential decision-makers can follow predecessors while disregarding their own information. Their framework helps explain why conformity can coexist with fragile collective beliefs. A later review of information cascades and social learning, coauthored with Omer Tamuz, surveys the theoretical and empirical research that followed.
Consider a hypothetical software investigation. Agent A interprets a failed test as proof that an authentication library is broken. Agent B reads A’s report and recommends replacing the library. Agent C summarizes both messages as two confirmations of the same defect. A coordinator now sees three agents agreeing, although the entire chain rests on one interpretation of one test.
Adding Agent D will not necessarily produce new information. If D only reads the summary, the workflow has created another opportunity to repeat the claim. The relevant unit of corroboration is the independent observation: a separate reproduction, a different test or a direct inspection of the suspected failure.
This distinction matters even when every component is capable and cooperative. Agreement is useful only to the extent that its evidentiary foundations justify it. A system should therefore track which agents performed a check, which merely interpreted another agent’s output, and which repeated a conclusion without checking it.
Why Independence Changes the Arithmetic
A simple calculation illustrates the stakes. Suppose five hypothetical voters each answer a binary question correctly with probability 0.7, and their answers are independent. The probability that at least three are correct is about 83.7%. Combining their votes improves on the 70% accuracy of a single voter.
Now suppose all five copy one answer. The majority is correct only when that original answer is correct: 70% of the time. The visible number of participants has increased, but the amount of independent information has not. These are illustrative probabilities, not performance measurements from the new research.
There is another useful calculation for interpreting the stress condition. If four signals independently have a 30% chance of being wrong, the probability that all four are wrong is 0.3 to the fourth power, or 0.81%. That describes the probability of a particular starting sequence under those assumptions. It does not describe the probability that a deployed agent team will fail.
A failure estimate would require both how frequently difficult situations occur and how the system behaves when they occur. Confusing those quantities can make a result look either more alarming or more reassuring than the evidence permits. A stress test can expose a consequential weakness without measuring its frequency in ordinary work.
Build an Evidence Trail Before Building Consensus
A practical response is to separate observations from interpretations in the shared workspace. For the hypothetical authentication investigation, an observation might contain a test identifier, the tested code revision, the actual output and the execution time. A separate field would record the agent’s proposed explanation and its uncertainty.
That structure lets the coordinator ask a concrete question: does this recommendation introduce a new observation, or does it refer back to evidence already counted? Three summaries citing the same failed test should remain one test in the evidence ledger. A second independent reproduction should be visible as a separate check.
For expensive or consequential decisions, teams could also collect initial assessments before exposing agents to one another’s conclusions. A reviewer would inspect the source material, commit an initial judgment and only then see the group discussion. Changes of mind would remain possible, but the system could record which new evidence justified them.
These are design proposals to evaluate, not safeguards validated by this experiment. They introduce costs: more structured records, additional tool calls and potentially slower coordination. Their value should be assessed against the decisions they protect. A brainstorming workflow may tolerate loose conversation; a production incident investigation may need a much stricter evidence trail.
Prompt Rules and Runtime Controls Solve Different Problems
Asking an agent to preserve evidence is useful, but a production architecture can go further by preserving the original tool output itself. An execution service could write results into a record that agents can cite but cannot overwrite. Readers could then compare the interpretation with the underlying output.
Even an immutable log does not guarantee that a test was well designed, that the source was trustworthy or that the interpretation is correct. It does, however, make discrepancies inspectable. A system can distinguish a faulty test from a report that inaccurately describes that test.
Authorization should remain a separate decision. A confident conclusion that a system needs repair does not confer permission to alter it. Keeping execution permissions outside the consensus process prevents an epistemic error from automatically becoming an operational action. An agent team can be wrong without being allowed to make an unrestricted change.
Model diversity should likewise be evaluated rather than assumed to solve the problem. Different models might contribute distinct interpretations, but assigning different names or roles to agents does not establish that their evidence is independent. A diverse group that reads the same mistaken summary can still share one evidentiary bottleneck.
What a Stronger Evaluation Should Measure
The linked publication is a public research writeup. Readers should avoid treating it as a comprehensive benchmark of deployed multi-agent products. Its value is the specific failure mode it makes testable; its broader implications need further evidence.
A useful follow-up evaluation would vary the ordering of signals, the accuracy of private evidence, the number of participants and the communication structure. It should include ordinary sequences alongside deliberately misleading ones, and correct early majorities alongside incorrect early majorities. Otherwise, a policy could appear successful simply because it teaches agents to distrust every consensus.
Accuracy alone would miss important distinctions. Evaluators should measure whether reports faithfully preserve observations, how often agents invent supporting evidence, whether a correct minority can change the final decision, and whether the coordinator counts repeated citations as separate checks. Token consumption and latency belong in the same comparison: a safer protocol still needs an operationally useful cost.
To investigate mechanisms, researchers could experimentally vary access to earlier judgments while holding task evidence constant. They could compare independent initial assessments with assessments formed after reading the board. Randomizing presentation order would help separate evidence effects from effects of sequence or prominence. Generated explanations can guide hypotheses, but they should not serve as the sole proof of why a model changed its answer.
A Better Definition of Multi-Agent Reliability
The language of herd mentality is vivid, but it should not be read as evidence that models experience human social pressure or possess human beliefs. The operational concern is observable: information enters a workflow, judgments influence later judgments, and the final answer may obscure where its support originated.
For builders, the next milestone should be demonstrable correction. When one agent makes a plausible mistake, can another identify the contradictory evidence? Can the coordinator preserve that disagreement long enough to investigate it? Can the final decision explain which independent checks resolved the issue?
A larger team earns trust when it adds verifiable information and improves decisions under challenge. Counting agents, counting endorsements and producing longer discussions are inadequate substitutes. The most useful multi-agent system is one whose evidence remains inspectable even when its participants agree.












