Published October 1, 2026 | Version 1.0

Perspectives Without Independence: Multi-Agent and Multi-Persona Reasoning Under Compute-Normalised Comparison, Why the Gains That Survive Are Not the Ones Diversity Predicts, and the Controls That Would Tell Them Apart

Authors/Creators

  • 1. SONYTECH

Description

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.

Multi-agent debate, multi-persona prompting and related schemes are usually justified by diversity of viewpoint: several perspectives err differently, so their combination is more reliable than any one of them. That justification is a claim about independent errors, and a scheme that samples all of its perspectives from one model has to earn the independence it needs. This paper contributes an audit of the published results that hold inference compute constant, an analysis of the six cost units those results use, and a procedure for attributing a matched-cost gain to one mechanism. Where debate or persona schemes are compared with self-consistency or majority voting at a matched number of calls, tokens or generations, most of the reported gain disappears, and voting over the agents' independent first answers accounts for most of what remains; in one replication, self-consistency with nine samples scored 88.2 percent on GSM8K against 83.0 for debate at the same count. Against that, a 2026 study reports mixture-of-agents and debate ahead of self-consistency at equal compute, and a controlled agentic study reports gains of up to 80.8 percent on decomposable tasks. The paper argues that these positive results do not rest on diversity of viewpoint: the strongest one uses a single model in every role, and the agentic gains track whether the task decomposes. It separates four mechanisms that "multi-agent" names (voting, synthesis by an aggregator, decomposition, and viewpoint diversity) and shows that the compute-normalised comparisons with self-consistency isolate the fourth almost nowhere for single-model personas, where the few matched tests give small and mixed results. Model heterogeneity is different: at equal numbers of calls, mixed-model debate beats single-model debate, and a clone-controlled study on estimation and forecasting tasks found that deliberation among different models improved on their own pooled answers where deliberation among copies of one model did not. Two controls are still missing: a heterogeneous vote at matched cost on a reasoning benchmark, and a measurement of conditional error correlation among persona agents, including how often they agree on the same wrong answer.

Files

perspectives-without-independence.pdf

Files (550.4 kB)

Name Size Download all
md5:0d9fe67ae6f84c53350fd4536698763a
550.4 kB Preview Download