CIVILIZATION CODEX
Home / Part IV — Civilizational Intelligence

Learning Governments

Status: draft.

Opening Question

A city tries a new policing strategy, a new school curriculum, or a new housing subsidy. Five years later, almost no one — not the officials who proposed it, not the public who lived under it — can say with confidence whether it worked. This is the normal condition of most government policy, not the exception. What would it take for policy-making to routinely produce an actual answer to "did this work," rather than an assumption, an anecdote, or a purely political verdict?

Historical Perspective

Governments have periodically rediscovered the idea that policy should be tested rather than merely legislated. Public health is the clearest long-running example: the controlled comparison — try an intervention in one place, compare it to a similar place without it — has roots at least as far back as James Lind's 1747 scurvy trial aboard HMS Salisbury, widely regarded as one of the first controlled clinical trials, which found that citrus fruit cured scurvy while several other contemporary remedies did not. It took the British Navy several more decades to actually adopt citrus rations fleet-wide despite the trial's clear result — an early, concrete illustration that generating good evidence and institutionalizing it into policy are two different achievements, and the gap between them can be measured in decades even when the evidence is unambiguous.

What Modern Evidence Suggests

particularly with Abhijit Banerjee, Esther Duflo, and Michael Kremer, jointly awarded the 2019 Nobel Memorial Prize in Economics) applied randomized controlled trials systematically to development policy questions — deworming programs, microfinance, teacher incentives — and found, repeatedly, that intuitive, well-funded interventions often underperformed or failed outright when actually tested, while some cheap, unglamorous interventions (deworming, in particular) substantially outperformed expectations. The broader significance is methodological: policy intuition, even from well-informed experts, is frequently wrong in ways only a real test reveals.

onward across policy domains (education, crime reduction, early intervention, local economic growth, among others), institutionalized evidence review and synthesis as a standing government function rather than a one-off study — an attempt to build the "publish and critique" stages of this book's policy cycle directly into ongoing government operation rather than treating evaluation as exceptional.

administration research: organizations run pilot after pilot, generate genuinely useful evidence, and then never scale the interventions that worked — often because the institutional incentives reward launching visible new initiatives far more than the less visible work of scaling and maintaining a proven one. A pilot that never scales, however well-evaluated, has not completed this book's policy cycle; it has stalled partway through it.

Where the Principle Fails

Evidence-based policy language is not immune to capture. "Evidence-based" can become a rhetorical shield for a predetermined conclusion — a sponsor commissions studies until one produces the desired result, then labels only that one "the evidence," a pattern sometimes called policy- based evidence-making. Randomized trials themselves, while a strong tool for isolating a specific intervention's effect, can also miss important context-dependent or systemic effects that only show up at a scale or timeframe no pilot can capture — a small, well-run pilot program can succeed for reasons (extra attention and resources, motivated staff) that quietly disappear the moment it is scaled to normal operating conditions, a phenomenon documented in implementation science as "voltage drop."

Civilization Design Principle

> Treat policies as hypotheses, run through a complete and > institutionalized cycle — propose, pilot, measure, publish, critique, > correct, scale — with an explicit organizational owner for each stage, > especially the scaling stage most often skipped.

PROPOSE → PILOT → MEASURE → PUBLISH → CRITIQUE → CORRECT → SCALE

Each arrow in this cycle is a place institutions typically underinvest: measurement without independent publication invites the policy-based- evidence-making failure above; publication without structured critique wastes the "organized skepticism" this book's science chapter (03_science_and_error_correction.md) identifies as essential; and critique without a genuine scaling pathway produces exactly the pilot-itis failure mode. A learning government is one that has built organizational ownership and funding for the entire cycle, not just the visible, popular first step.

Institutional Translation

default, with funding for the evaluation itself guaranteed at the time of passage rather than left to future discretion.

Centres, with a mandate covering the full cycle rather than one-off reports.

launches: what result would trigger scale-up, what result would trigger discontinuation, and who is accountable for acting on either outcome.

should be publishable and citable without political cost to the officials who piloted them in good faith — mirroring the scientific reforms discussed in 03_science_and_error_correction.md, since punishing honest negative results guarantees that future pilots will be designed, consciously or not, to avoid producing them.

Metrics

and dedicated evaluation budget at time of launch.

policy, and from failed pilot to formal discontinuation — both transitions, not only the successful one, since an indefinitely extended pilot is itself a scaling failure.

effect size measured after scale-up), tracked as a standard part of scaling evaluation rather than assumed away.

Questions Still Unresolved

Not every policy question is well suited to randomized testing — some interventions are too large, too slow, too entangled with other simultaneous changes, or too ethically fraught to pilot in a controlled way. What the right evidentiary standard is for exactly these higher-stakes, harder-to-test cases (constitutional change, large infrastructure commitments, irreversible ecological decisions) is not answered by the randomista toolkit, and remains a genuinely contested question in public administration, picked up again in 13_institutional_design/02_government.md.