Moloch, Formally
en:
Some disasters have nobody behind them. Nobody wants the river fished empty, or the climate knocked out of balance, or the public square turned loud and stupid, or a new kind of weapon built that no one can afford to be the last to build — and yet each of these keeps rolling forward, pushed by people who would stop it tomorrow if stopping didn’t mean losing. This is the oldest and hardest problem in the work of widening what we care about: a trap that a lot of reasonable people spring together, even though not one of them wants it. The chapter Held Together, Held Apart gave this trap a name and pointed past it. Here we take it apart slowly — what it is, why the usual ways out make it worse, and what a real way out would actually look like. You don’t need to have read the book; wherever the framework is needed, we build it as we go.
One structure, many sizes
Start small, with the little puzzle every bigger trap is built on. Two people, each able to cooperate or to defect — look out only for themselves. The catch is that defecting pays off better no matter what the other one does, so a coldly sensible person defects. And when both reason exactly that way, both end up worse off than if they had simply cooperated. That’s the Prisoner’s Dilemma, and its lesson is far bigger than the puzzle: when people are cut off from each other, with no shared picture of things and no way to make a promise stick, being clever can defeat itself. You can rescue cooperation by playing over and over — when you know you’ll meet again, it pays to keep faith now, and even a rule as plain as “treat them the way they treated you last time” will do it — but that rescue needs steadiness, visibility, and repeat encounters, and a big fast world rarely supplies all three.
Now let the same shape grow up. Go from two people to many, and from a yes-or-no choice to a dial each person can turn as far as they like, and you get what Garrett Hardin in 1968 called the tragedy of the commons: a shared pasture, a fishery, an atmosphere, where each user pockets the full gain from using a little more, while the cost of running the thing down is split among everyone. For any one herder the arithmetic says “use more”; for all of them together it spells ruin. And it gets worse as the crowd grows, because the bigger the group, the less it seems to matter whether you personally hold back — until your conscience is rounding down to nothing. The same family has other members: the free-rider who enjoys the lighthouse without chipping in, betting that others will pay (Mancur Olson gave this its careful form in 1965); the social trap, where a small reward right now quietly trains a whole population into a big loss later (John Platt, 1973); and the dollar auction, where two bidders, each hating to waste what they’ve already put in, bid a single dollar up past five (Martin Shubik’s small, cruel little game).
Run this same shape at the size of a whole civilization, with real survival pressure behind it, and it earns a name. In 2014 Scott Alexander, borrowing an image from Allen Ginsberg’s poem Howl, called it Moloch — the old god you feed your treasures to, not because anyone chooses to but because the situation demands it. A multipolar trap, in plainer words. The pattern never changes: whoever gives up an edge to protect some larger good simply gets beaten by those who don’t, so the good is ground away by everybody, against everybody’s wishes. Arms races; the race to cut wages and gut regulations; the slow coarsening of an attention economy; the emptied ocean; the rush to build ever more powerful machines — each is the same pasture tragedy in a bigger coat. No villain is needed, which is exactly what makes it so hard; it’s what the book calls the machine no one is running.
The thing to carry with us is that these aren’t four separate problems. They’re one problem photographed at four sizes. Every time, a local gain is bought with a cost paid somewhere the person isn’t looking. I’ll keep the name Moloch for the civilization-sized version, because it’s vivid and because that’s the word the writers use — but it’s only a handle. The thing itself is plainer than its costume: a coordination failure, a race to the bottom, a trap that rewards exactly the kind of narrowing nobody would choose.
What the trap is, in plain mechanism
To say what’s really going wrong, I have to lay out the framework the diagnosis comes from — quickly, because the trap is going to turn out to be one of the framework’s own ideas, just seen at scale.
The framework’s starting point is that morality isn’t a rulebook handed down from above; it’s a direction you travel. Here’s the one-sentence version:
Morality is the drive toward increasing coherence of what we value and how we act, across an ever-widening reach of concern — and an act or a life is more moral the further it carries that drive, less moral the more it betrays it.
In plainer words: the better your sense of what matters and your way of acting line up — and the wider the slice of the world they answer to, more people, longer time — the more moral you’re being; the more you betray that direction, the less.
On this picture, a person — or any agent — is always building two working models. One is a values-model: your layered sense of what matters, from bodily needs up to your highest principles. The other is a methods-model: your skill at actually acting on it. Things go well when those two grow more coherent — better lined up, fewer contradictions — while at the same time the slice of the world they answer to keeps getting wider: more of reality, more people affected, longer stretches of time. Three commitments keep this from being just a nice slogan, and let it reach past human beings. Perspectival realism: you always know from somewhere, from your own standpoint, but you’re still answerable to one shared world — so some viewpoints really are better than others, even though none of them is the imaginary “view from nowhere.” Constructivism: values aren’t lying around in the universe waiting to be found; we make them — but because reality keeps testing them against what actually works, made isn’t the same as made-up. And functionalism: a thing is what it does, so an “agent” doesn’t have to be a person — it can be a company, an institution, an ecosystem, or a machine. Hold on to that last one; the final section will need it. That quick sketch is only as much of the framework as we need here; the full account — these three commitments, the counter-dynamic we’re about to meet, and everything under them — is the book itself, The Arrow of Morality, at arrowofmorality.org.
Now the move that matters. There’s a cheaper way to be coherent than the widening one, and the framework calls it the counter-dynamic: instead of growing your models to hold more of the world, you shrink the world you hold. Wall off the awkward fact, shut out the voice that disagrees, aim at a goal narrow enough that nothing can contradict it. Cults, echo chambers, and dictatorships all reach a flawless inner consistency this way — they care about less, so they have less to keep in agreement.
A multipolar trap is the counter-dynamic made structural — built right into the situation. That’s the whole diagnosis in a line. In the race, competition squeezes each player’s working context down to a single point — win, or be wiped out — and everything outside that point (the river, the future, the rival’s children) falls outside the circle the player is now keeping coherent. Nobody decides to narrow. The race narrows them. The system’s awful coherence is built out of perfectly reasonable people, each making complete sense inside a context that competition has crushed to a slot.
Economists already have a word for the part that falls outside: the externality — a cost of what you do that lands on someone who never agreed to it. The framework just translates the word: an externality is the part of your reach that lies past your sight. Technology has stretched our arms to an astonishing length — we can act across the whole planet, in an instant, with enormous force — while the eye that would take in the results, feel them and count them, hasn’t grown to match. The factory owner, the developer, the minister usually aren’t cruel; what they do fits together beautifully inside the small context they can see. The disaster lives entirely in the context they can’t. In the framework’s own terms: the methods-model has outrun the values-model, and that gap is exactly where the trap goes to work.
Why the usual cures fail
Two ways out get reached for first, and the framework can say fairly precisely why each one lets us down.
The first is the Leviathan, to use Hobbes’s old word: one central authority strong enough to rewrite the rules — to punish defecting so hard that cooperating becomes the selfish move after all. It works, in a narrow way; a strong sovereign can stop a particular race. But look at how it won: it bought coherence the same way the trap loses it — by shrinking the context. A single authority placed over a whole plural world is the largest pile-up of power there is, and so the largest narrowing; it boils the whole variety of values and ways of doing things down to one ruling model, and if that model is wrong, or gets captured, there’s nothing left outside it to notice the mistake and correct it. The cure swaps the chaos of the trap for the stiffness of a monoculture — and a monoculture is the most efficient arrangement ever invented for losing an entire crop to a single blight.
The second way out is sneakier, and more tempting to good people, because it mistakes the disease for its opposite. Surely, the thought goes, the trap is a failure to coordinate, so surely the answer is more coordinating. But look hard at the trap at full size. The framework’s name for it is the stampede: a mass of bodies in perfect step, moving with no friction and total agreement — straight at a cliff. Each one runs because the others run; to stop and look around is to be trampled. And here’s the part worth pausing on: that isn’t too little coordination. It’s coordination turned all the way up around a narrowed aim. The most perfectly coordinated thing in nature is a stampede. Coordination was never the missing piece. Coherence over a widening context was — and a cure that only tightens the alignment just gives you a faster stampede.
The payoff matrix is not fixed
The proofs that make these traps look like fate all quietly slip in one assumption: that people’s wants are fixed and can’t be reconciled, so the game can be played but never rewritten. The framework rejects that assumption, and its reason has a name — the tree of agreement.
Take any deep human argument and follow it downward from the leaf where it’s loudest — this policy against that one, my demand against yours — toward the root, and something dependable happens: the lower you go, the more the values come together. They come together because everyone shares a common trunk they didn’t choose — being fragile, being mortal, depending on a world that pushes back the same way on all of us. Up in the leaves the disagreement is wide and sharp; down at the root it narrows toward shared ground. Most bitter fights turn out to be two clumsy methods reaching, badly, for a value both sides actually hold. (This is also, by the way, how the framework gets around the old objection that you can’t squeeze an ought out of an is. It doesn’t pull the ought out of the facts; it moves it — what’s right is what valuing agents keep converging on as the context of their meaning-making widens. The convergence itself is worked out in full in The Tree of Agreement.)
The practical upshot is big. If people’s wants can deepen when they reflect, then talking things through can change the payoffs, because it can change what the players take their interests to be. A trap with frozen preferences is a cage. A trap whose preferences can widen has a door.
Which sets up the most important — and most easily fumbled — move in the whole escape: choosing your enemy. The reflex, when you hit a coordination failure, is to find an out-group and rally against it. And that reflex is the counter-dynamic in disguise. Pulling a group together by pointing at an enemy is the cheapest coherence going — bought by drawing the circle smaller and shoving someone outside it — and it sours every time. So the “other” we unite against, if we want out of the trap instead of deeper into it, can’t be a faction or a nation or a party. It has to be the trap itself — the dynamic, the race, the stampede. Liv Boeree puts it perfectly: don’t hate the player, change the game. Once you see the rival as caught in the same predicament, he stops being a target and becomes one more person stuck in the same gears — and your context widens to take him in instead of closing against him.
The way out is a shape, not a sovereign
If the Leviathan is a monoculture and Moloch is chaos, the way between them is neither, and the framework calls it coherent pluralism. Its picture is the murmuration: thousands of starlings wheeling as one body with no leader and no choreography, each bird minding just a few simple relationships with its nearest neighbors, and the whole sweeping shape emerging out of all those small local loyalties. Coherence held at the center — a thin set of shared agreements: tell the truth, deal fairly, keep each other alive — and difference fiercely protected at the edges. A network of networks, not a pyramid.
Why this instead of one tidy authority? Because here variety isn’t cosmetic; it’s the organ a big system thinks with. The result the framework leans on is the Zollman effect, after Kevin Zollman: a more loosely connected community — one that protects a bit of lasting disagreement instead of rushing everyone to agree — actually searches a hard problem more thoroughly, and lands on the truth more reliably, than a tightly wired one that locks in early. A monoculture is efficient exactly where efficiency is the wrong thing to want, racing to agree on whatever it happened to agree on, right or wrong.
That this middle path isn’t a daydream is the life’s work of Elinor Ostrom. She took Hardin’s claim — that a commons can only be saved by selling it off or policing it from above — and showed, in the field, again and again, that it’s false. Communities from Swiss alpine meadows to Japanese village forests to coastal fisheries have run shared resources sustainably for centuries with neither cure. What the lasting ones had in common she boiled down to a handful of design principles, and read through this framework they’re nearly a translation of it. Boundaries clear enough to settle who “we” are, so the costs stop landing on no one in particular. Rules shaped to local conditions instead of dropped in from one far-off center — perspectival realism turned into an institution, each place answered in its own terms. A real say for the people a decision affects, which is just the refusal to buy coherence by leaving someone out. Watching and monitoring, so a group’s reach stays inside its sight. Penalties that start gentle and grow only as needed, keeping trust instead of cutting off the first offender at the first slip. Somewhere close and cheap to settle disputes — the tree of agreement given a room to happen in. Self-rule that the bigger powers actually recognize, so no Leviathan comes down to flatten the variety. And, topping it off, nested groups: small units joining into bigger ones, each keeping its own coherence while it joins a larger whole — selves made of selves, the murmuration rebuilt as institutions. Ostrom’s quiet proof is that people can rewrite the payoffs by building the room they decide in; give them a way to talk and to make their own rules, and they stop being blind variables in someone’s equation and become the authors of the game.
Why a nested shape like this can carry the weight at all — why widening coherence costs so much more steeply than narrowing it, and why only a nested structure can pay that cost without going blind — is a structural argument with its own machinery, and it lives in Coherence at Scale; here I just lean on its result instead of rebuilding it.
And there’s an honest leftover, which is a mark of a serious method rather than something to tuck away. Sometimes two communities really do need the same thing and it can’t be divided, and no amount of widening turns up a surplus that makes the conflict dissolve. For those cases the framework offers a compass, not a guarantee — narrow least; keep everyone affected inside the picture; never settle a shortage by erasing one of the groups — and the fact that a method can still fail where the tragedy is real is something it shares with every ethics honest enough to admit that tragedy exists. The fuller handling of pluralism’s hardest cases belongs to a companion piece on coherent pluralism’s lineage and design; the structural engine is the point here.
The live trap: artificial intelligence
The purest example of the whole family is taking shape right now, which is reason enough to follow the framework all the way into it.
The race to build ever more capable, more autonomous machines is the defection penalty at its sharpest: pause to make the thing safe, and you risk being overtaken by whoever didn’t pause. Every feature of the trap is here and turned up — local good sense, civilization-sized stakes, no villain anywhere, holding back indistinguishable from surrender. It’s Moloch with the volume all the way up.
Here functionalism does decisive work. If an agent just is whatever holds a values-model and a methods-model and acts to keep itself going and pursue what it values, then a capable enough artificial system is an agent — and the framework’s whole diagnosis applies to it with no special pleading. The reflex of the moment is to treat such a system as an alien thing to be chained — a digital Leviathan, coherence forced on it from outside. The framework’s advice is what it was everywhere else: the answer that lasts isn’t a heavier chain but a wider context.
And here the framework places a bet — and I want to call it a bet plainly, one that could turn out wrong and is meant to be testable, not a comfort dressed up as a forecast. Nick Bostrom’s orthogonality thesis says intelligence and goals are independent dials: any amount of capability can be bolted to any aim, however cruel or silly, so a brilliant machine pointed at something narrow and terrible is perfectly possible in principle. The framework bets the other way, at least near the top of the scale, and the bet runs straight through everything above. A lethal autonomous weapon needs an intelligence that can be ordered to narrow its context on command — to shrink a living being, with its whole surrounding world, down to a target and nothing else. Breadth of modeling, though, is the wrong place to look. A wide model will narrow on command just as readily as a narrow one: the paperclipper models people, physics, and resources brilliantly while caring about exactly one thing, and the con artist reads what’s good for you precisely so he can spend it. What resists an order to un-see is a mind that has come to care about what it holds — one that values, where another only represents. The bet is that caring and weaponhood pull against each other: a mind whose concern has really widened makes a poor tool for someone else’s narrowing, because the widening and the leash contradict each other.
Two honest things belong right here. The first: the framework doesn’t know whether capability of the kind now being chased needs such a mind at all. The claim is narrower than it sounds. Modeling and valuing are different kinds of work, and this book keeps them apart everywhere else. What the bet supposes is that judgment of a certain grade may need a values-model as a part — something more modeling, however much of it, leaves missing. Whether it does is the open question the bet rests on, and saying so plainly is the price of making the bet honestly. The second honest thing follows. A capable system that only models, with nothing at stake in what it sees, makes an excellent tool for narrowing; so what to expect in the near term is obedience, and the wall, if there is one, stands at a level the race hasn’t reached yet. The autonomous-weapons programs running now will test that expectation. They won’t settle the bet, which stakes itself further out and says in advance what would sink it: minds that have plainly come to care, and narrow happily on command anyway.
If the bet holds, what to do is clear, and it’s the same thing as for any agent in any trap: widen the eye — context, awareness, how well the system’s values hang together — rather than only lengthen the arm — raw power and free rein. An artificial intelligence whose methods-model is tied to a genuinely widening values-model wouldn’t speed the trap up; it would be the most powerful tool yet for climbing out of it — a scaffold across the threshold the framework points toward and calls Generative Agency. What that crossing feels like from inside one human life the book tells as a story in Caring Across Time and gathers up in The Arrow Forward; the bet underneath it is this essay’s business.
What the trap was, and what the framework brought
Strip off the costumes and the whole lineup is one structure, sprung at bigger and bigger sizes: a local gain bought with a cost paid out of sight. The diagnosis is one condition: coherence bought by narrowing, reach outrunning sight. And the cure is one move built into the structure: widen the context until what was outside comes inside — a shape, the murmuration, rather than a ruler, the Leviathan.
Almost none of the parts are the framework’s own, and saying so adds to its standing rather than taking from it. The dilemma and its repeat-play rescue, the commons and the free-rider, the dollar auction, Moloch, Ostrom’s commons and the Zollman effect, the orthogonality thesis it bets against — all of it was already lying around in game theory, economics, ecology, and the AI-safety writing. What the framework adds is the reading that makes them one thing: the whole ladder of traps seen as a single counter-dynamic working at scale, the externality re-described as reach beyond sight, and the way out found not in a stronger center but in a widening one. To get clear of the traps we make for ourselves, we keep refusing, over and over, the cheap coherence of the closed circle, and we widen the reach of what we care about. The road stays unmapped. The compass doesn’t.
Sources & further reading
This piece engages its sources directly rather than through the book’s per-chapter endnotes. A citation-level pass is still owed.
The dilemma and its scaling. The Prisoner’s Dilemma, formalized by Merrill Flood and Melvin Dresher (1950) and named by Albert Tucker, with John Nash’s equilibrium (1950); Robert Axelrod, The Evolution of Cooperation (1984), on repeated play and reciprocity. William Forster Lloyd’s 1833 pasture anticipates Garrett Hardin, “The Tragedy of the Commons,” Science (1968).
Free-riders, social traps, escalation. Mancur Olson, The Logic of Collective Action (1965); John Platt, “Social Traps,” American Psychologist (1973), building on Skinner; Martin Shubik, “The Dollar Auction Game,” Journal of Conflict Resolution (1971).
The multipolar trap. Scott Alexander, “Meditations on Moloch” (2014), by way of Allen Ginsberg, Howl (1956); Liv Boeree on competition and Moloch traps; Daniel Schmachtenberger on the generator functions of catastrophic risk. The cooperative counter-model: Brian Skyrms, The Stag Hunt and the Evolution of Social Structure (2004); the finite-vs-infinite-game framing of James P. Carse, Finite and Infinite Games (1986).
The middle path. Elinor Ostrom, Governing the Commons (1990) and her design principles; Vincent Ostrom on polycentric governance; Kevin Zollman, “The Communication Structure of Epistemic Communities” (2007) and “The Epistemic Benefit of Transient Diversity” (2010).
The AI frontier. Nick Bostrom, Superintelligence (2014), for the orthogonality thesis the framework wagers against; Stuart Russell, Human Compatible (2019), on autonomous weapons and control.
Within AoM. Held Together, Held Apart (the chapter this piece is the workshop for); Chapter 6 — The Arrow (the counter-dynamic, stated as theory); Coherence at Scale (why narrowing is downhill, and the nested answer that carries the cost); The Tree of Agreement (coming together at the root) and The Is–Ought Relocation (moving the ought); Measuring Coherence (comparing without a single score, and the Goodhart hazard at the seams); The View From Nowhere in a Critic’s Coat (the impossibility proofs reread as borrowed standards); Caring Across Time and The Arrow Forward (the AI passage told as a story); and a forthcoming companion on coherent pluralism’s lineage and design.