Collective Decision Making
Status: draft.
Opening Question
At a 1906 country fair, statistician Francis Galton recorded nearly 800 guesses of the weight of an ox, expecting to demonstrate the foolishness of the crowd. The median guess was 1,207 pounds; the ox weighed 1,198 — closer than nearly any individual guess, including the experienced farmers and butchers in attendance. A century later, groups deliberating together on politically charged topics routinely produce worse judgments than their individual members would have reached alone. Both are real, replicated findings about collective judgment. What separates the crowd that outperforms its members from the committee that underperforms them?
Historical Perspective
James Surowiecki's The Wisdom of Crowds (2004) revisited Galton's result and a wide range of similar cases to identify the conditions under which aggregated independent judgments outperform individual experts: diversity of opinion, independence (individual judgments not influenced by others' answers), decentralization (people can draw on local and specialized knowledge), and a working aggregation mechanism to combine individual judgments into one collective decision. Notice that three of these four conditions concern how information is gathered before aggregation, not how skilled any individual participant is — which is why average, non-expert crowds can outperform credentialed individual experts under the right structural conditions, as in the ox-weighing case.
Formal social choice theory adds a sobering constraint: Kenneth Arrow's Impossibility Theorem (1951) proved that no voting system aggregating three or more ranked options can simultaneously satisfy a small set of seemingly reasonable fairness conditions (no dictator, respecting unanimous preference, independence from irrelevant alternatives, among others) — meaning every real-world voting system embeds some tradeoff or compromise, not a solvable engineering defect waiting for a clever fix. This does not make collective decision-making hopeless, but it does mean no voting mechanism can be treated as a neutral, assumption-free way of discovering "what the group really wants."
What Modern Evidence Suggests
- Groupthink (Irving Janis's 1972 study of U.S. foreign policy
fiascoes, most notably the Bay of Pigs invasion) identified the specific conditions under which face-to-face deliberation degrades judgment rather than improving it: high group cohesion, insulation from outside expert opinion, directive leadership that signals a preferred outcome, and high stress with a perceived lack of viable alternatives. Under these conditions, groups suppress dissent, especially self-censorship, and converge on a decision faster and with more confidence than the underlying evidence justifies.
- Deliberative polling and citizens' assemblies (pioneered
academically by James Fishkin, applied at national scale in Ireland's Citizens' Assembly on abortion, 2016–2018) show that structured deliberation — a randomly selected, demographically representative sample, given balanced expert briefings and facilitated small-group discussion under transparent process — can produce recommendations that shift participants' views substantially and are broadly perceived as legitimate across the political spectrum, in ways ordinary legislative debate on the same topic did not achieve.
- Prediction markets (studied extensively in behavioral economics and
used operationally by some organizations for internal forecasting) aggregate dispersed private information through price signals rather than deliberation or voting, and have shown competitive or superior accuracy to expert panels on a range of well-defined, verifiable questions — though they perform poorly on questions with thin trading volume or unclear, hard-to-verify resolution criteria.
Where the Principle Fails
The conditions Surowiecki identifies for crowd wisdom — independence and diversity above all — are precisely what social media and modern political polarization erode: when everyone's judgment is influenced by the same viral content or partisan signal before they answer, the "crowd" degrades into a single correlated opinion wearing many faces, losing the statistical benefit of aggregating genuinely independent estimates. Wise- crowd effects and groupthink are not opposite phenomena occurring in different populations; they are two outcomes of the same aggregation process depending entirely on whether independence is preserved.
Civilization Design Principle
> Structure collective decisions to preserve independent judgment before > aggregation, and match the decision mechanism (voting, deliberation, > markets) to the kind of question being asked rather than defaulting to > one method for everything.
Different collective-decision tools are suited to different problems: markets and prediction aggregation work well for well-defined, verifiable questions with liquid participation; structured deliberation with balanced briefings works well for value-laden questions requiring legitimacy and buy-in; simple independent-judgment aggregation (Galton's ox) works well for quantitative estimation problems. Treating any one of these as a universal solution — pure direct democracy, pure technocratic expert panels, pure market mechanisms — reproduces a specific, well-documented failure mode rather than avoiding it.
Institutional Translation
- Citizens' assemblies and sortition-based bodies for high-stakes,
value-laden questions requiring broad legitimacy, modeled on Ireland's process: random representative selection, balanced expert briefing, facilitated deliberation, transparent process (see 05_power_and_governance/07_local_governance.md).
- Structured elicitation protocols (e.g., the Delphi method, or
simply requiring written independent estimates before group discussion begins) as standard practice for expert panels, specifically to preserve the independence condition before groupthink dynamics can take hold.
- Internal prediction markets or forecasting tournaments for
well-defined organizational and policy questions, modeled on Philip Tetlock's Good Judgment Project (referenced in 01_defining_civilizational_intelligence.md).
Metrics
- Divergence between individual pre-deliberation judgments and post-
deliberation group consensus, as a diagnostic: healthy deliberation should show belief updating toward evidence, not uniform convergence toward whatever the highest-status participant initially believed.
- Calibration (Brier score) of aggregated forecasts against actual
outcomes, by aggregation method, allowing direct comparison of markets, panels, and crowds on shared questions.
- Perceived legitimacy of a decision's process, surveyed separately from
agreement with its outcome — a proxy for whether deliberative processes are achieving their distinctive advantage (buy-in across disagreement) rather than merely re-deriving the median position.
Questions Still Unresolved
Sortition-based and deliberative bodies address groupthink and polarization well but have no obvious mechanism for the kind of continuous accountability that elections provide — a citizens' assembly cannot be "voted out" if its recommendation turns out badly. How much authority these bodies should hold relative to elected representatives, and how their legitimacy interacts with democratic accountability, is a genuinely unresolved institutional design question, developed further in 05_power_and_governance/03_democracy_and_its_failure_modes.md.