Most things you measure cluster around an average. Height, exam scores, apple weights. But earthquakes, wildfires, city sizes, book sales, war deaths and startup returns are ruled by rare, enormous events. Knowing which kind of world you're in changes how you should plan, invest and build.
Measure the height of a thousand adults and you'll get a tidy cluster around the average. Measure the wealth of a thousand adults and you'll get something very different: a crowd of ordinary numbers and one or two giants who swamp everything else.
Height, exam scores and the weight of apples behave one way. Earthquakes, wildfires, city sizes, book sales, war deaths and startup returns behave another way. They are ruled by rare, enormous events, and knowing which kind of world you're in changes how you should plan, invest and build.
The world where averages work
Height is shaped by many small, independent influences: hundreds of genes, plus nutrition, illness, sleep and more. These push and pull in different directions and mostly cancel out. When many small effects add together, you get the bell curve (the normal distribution). This is the Central Limit Theorem at work.
The bell curve has a built-in ceiling. The tallest person in recorded history was Robert Wadlow, at about 2.72 meters. No one will ever be 27 meters tall. In this world the average is a good summary, and outliers are mild.
A simple game shows the same thing. Flip a coin 100 times and win $1 for each head. You'll land near $50 almost every time. Play a hundred times and the results all look alike. You can price this game precisely.
What changes when effects multiply
Now change one rule. You start with $1. Each flip, heads multiplies your money by 1.1 (a 10% gain) and tails multiplies it by 0.9 (a 10% loss). The average of the two factors is exactly 1, so it seems fair.
It isn't, for most players. A head followed by a tail leaves you with 1.1 × 0.9 = 0.99, a 1% loss. After 100 flips, a typical player has about 50 heads and 50 tails and ends with roughly 60 cents.
Yet the average across all players is still exactly $1. Why? A tiny number of lucky players string together long runs of heads. One who flipped 100 heads in a row would hold about $13,780. Those rare winners pull the average up while the typical player slowly shrinks.
When effects multiply, you get a lopsided curve with a long right tail (a lognormal distribution). The mean and the median drift apart. The average of the system stops describing the experience of the typical member.
Power laws: no typical size
A power law goes further. The rule is simple: each time you double the size of something, it becomes rarer by a fixed factor. In symbols, the chance of seeing something of size at least x falls like x to a negative power.
On an ordinary graph, this looks like a steep cliff with a very long tail. On a log-log graph (where each gridline means "times ten"), it becomes a straight line. The slope of that line is the exponent that controls how heavy the tail is.
The important consequence is that there is no typical scale. In a bell curve, the average and the spread tell you what's normal. In a power law, that yardstick disappears. If the exponent is low enough, the math gets stranger still:
Variance can be infinite, so the "spread" has no stable meaning. With a still lower exponent, even the mean is infinite, and the sample average keeps jumping upward each time a monster arrives.
Put a billionaire in a stadium of 50,000 people and the average net worth of the room jumps by millions. It describes no one in the room. In a heavy-tailed world, the biggest observation can be most of the total.
Pareto's discovery
In the 1890s, the Italian economist Vilfredo Pareto studied income and land records and found a recurring shape: a large number of small holdings and a small number of huge ones, with the upper tail decaying as a power law rather than a bell curve. He is also credited with noticing that about 80% of Italy's land belonged to about 20% of its people, the seed of the famous "80/20 rule."
Two honest caveats. Pareto's exponents differed between places and periods, so this isn't one universal number. And the power-law shape usually describes the upper tail, the rich, rather than the whole population.
The same skewed pattern shows up in many places: city populations, word frequencies, book and music sales, citations of scientific papers, and the size of wars. In venture investing, a small share of deals typically produces most of the returns. The exact percentages vary a lot by dataset, so treat any tidy "X% of Y" statistic with suspicion unless it comes with a source.
The St. Petersburg paradox: why the mean can mislead
In 1738, Daniel Bernoulli described a game. A pot starts at $2. Flip a coin until it lands heads. If heads comes on flip n, you win 2ⁿ dollars.
Half the time you win $2. A quarter of the time you win $4. An eighth of the time, $8. Each possibility contributes exactly $1 to the expected value, and there are infinitely many possibilities, so the expected value is infinite.
Yet almost no one would pay more than a few dollars to play. They're right. Three quarters of the time you win $4 or less, and about 99.9% of the time you win under roughly $1,000. The infinite average comes from outcomes so rare you will essentially never see them, and even if you did, the casino could never pay them. Payouts grow as fast as probabilities shrink, which is exactly the recipe for a power law.
The lesson is that an average driven by rare extremes is a poor guide to a single life. You live one path, not the average of all possible paths.
Where power laws come from
There's no single cause, but a few mechanisms recur.
Two exponentials that cancel. Earthquakes follow the Gutenberg–Richter law. For each step up in magnitude, there are about ten times fewer quakes, but each releases about 32 times more energy. A steady drop in frequency against a steeper rise in energy yields a power law in the energy released. The same fault process that produces countless tiny tremors also produces the rare giant. There is no special "big quake" mechanism.
Rich-get-richer. In growing networks, new nodes tend to link to already well-connected ones. This preferential attachment (Barabási and Albert, 1999) produces a few giant hubs and a long tail of small ones. It's a good model for parts of the web and citation networks, though how universal such "scale-free" networks are remains debated.
Multiplicative growth. Compounding, as in our coin game, naturally produces heavy right tails, though often lognormal rather than strictly power-law.
Criticality. This is the most surprising mechanism, so it gets its own section.
Criticality: when everything is connected
Heat a piece of iron. Below about 770°C (its Curie temperature), its atomic magnets line up and it behaves as a magnet. Above that, the alignment is destroyed and the spins point randomly.
Right at the transition, something strange happens. Clusters of aligned atoms appear at every size at once, from tiny to enormous, and a magnified picture looks statistically like the unmagnified one. The distance over which one atom influences another becomes effectively unbounded. A small nudge might die out immediately or ripple across the whole material, and you can't tell in advance which.
Physicists also found universality: very different systems, such as a liquid-gas transition and a simple magnet, can share the same critical exponents because the details wash out.
Self-organized criticality: systems that tune themselves
For ordinary phase transitions, an experimenter has to carefully set the temperature to reach the critical point. In 1987, Per Bak, Chao Tang and Kurt Wiesenfeld asked whether some systems might reach criticality on their own.
Their answer was the sandpile model. Drop grains onto a pile one at a time. The pile steepens until it hovers at a critical slope. From then on, one more grain might do nothing, or it might trigger an avalanche of any size. Avalanche sizes follow a power law, and there's nothing special about the grain that starts a big one. The state of the pile is what matters.
A related toy model of forest fires (Drossel and Schwabl, 1992) works similarly. Trees grow slowly, lightning strikes randomly, and fires spread to neighbors. The forest fills until fuel connects across large areas, and then a random spark can burn one tree or a vast region.
Real forests are messier. Fire-size records are indeed heavy-tailed, and decades of fire suppression in many dry western U.S. forests let fuel build up, making later fires larger and hotter. But big fires also depend heavily on drought, heat and wind, so SOC is a useful lens, not a full explanation. The same caution applies to SOC as a whole: it's an elegant idea, and how many real systems truly follow it is still argued over.
The moral of the model is still useful: in a system near criticality, the trigger is trivial and the condition of the system is what counts. After a disaster, "what was the spark?" is often the less important question. "How did the system get so primed?" is the better one.
What to do with this
-
Don't confuse calm with safety. Bertrand Russell's chicken, retold by Nassim Taleb as the "turkey problem," is fed every day and grows more confident until the day it isn't. A long quiet stretch may mean nothing has gone wrong yet, not that nothing can. A business or portfolio that shows smooth profits may be quietly selling insurance against rare disasters.
-
First, figure out which world you're in. Some domains are bell-curve-like: manufacturing tolerances, flight schedules, restaurant table turnover. Here consistency and efficiency win, and the average is a fair guide. Others are heavy-tailed: venture investing, publishing, media, open-source adoption, drug discovery. Here a few outliers dominate the results.
-
In heavy-tailed domains, survive and keep exposure to upside. Avoid any bet that can wipe you out permanently, because you can't collect a rare jackpot if you're out of the game. Within that limit, prefer bets with a small, known downside and a large open-ended upside: writing, publishing a tool, shipping a product, running many cheap experiments. Expect most to do little, and accept that.
-
Look beyond the mean. For any number that varies wildly (revenue per customer, traffic per page, returns per deal), also check the median and the top percentiles. If the mean is far above the median, a few extremes are doing the work. A log-log plot can help, but a straight-looking line is not proof. Many other distributions look straight over a limited range, so use proper statistical fitting before claiming a power law.
-
Don't over-suppress small shocks. Frequent minor failures and small fires can release pressure that would otherwise build toward a catastrophe. This doesn't mean courting disaster. It means being wary of systems, in engineering, teams or finance, that promise to eliminate all volatility.
-
Remember survivorship bias. Heavy tails mean the winners look spectacular, and we hear mostly from them. Being in a power-law domain makes success more skewed, not more likely. Count the failures too.
The takeaway
There are two broad kinds of randomness. In one, many small effects add up, and averages, consistency and optimization serve you well. In the other, effects multiply and cascade, and a few rare events account for most of what happens.
You can't predict the timing or size of the next avalanche. But you can learn to recognize which kind of game you're playing, and then choose your strategy, your risks and your expectations to match.