Power Laws
The average is often a decoy: in a power-law world, the largest few observations can matter more than the next million. A power law is a distribution where frequency falls as size rises: roughly, P(x) ∝ x^-α. Heights do not behave this way; cities, fortunes, links, citations, earthquakes, and startup returns often do.
The case
Bell curves train the eye to look for the middle. Power laws train the eye to look for the tail.
In a normal distribution, a 7-foot adult is rare and a 10-foot adult is biologically absurd. In a power-law distribution, the equivalent jump is normal enough to build institutions around. New York City has about 8.3 million people; the smallest incorporated places in the United States can have fewer than 10. The largest YouTube channels have hundreds of millions of subscribers; most channels have almost none. A paper with 50,000 citations and a paper with 5 citations can live in the same scientific database.
The exponent matters. In many empirical systems, α sits somewhere between 1.5 and 3, but the exact number is the argument, not decoration. If the tail is fat enough, variance may be unstable or undefined. That is why a single observation can rewrite the dataset: one fund return, one earthquake, one pandemic cluster, one bestselling book.
Where it shows up
Zipf’s law is the cleanest doorway. In language, the most common word appears about twice as often as the second, three times as often as the third, and so on. In English, “the” is usually around 5-7% of ordinary text. The tail contains words a reader may see once in a decade.
Named cases keep repeating across domains:
| Domain | Pattern | Named reference |
|---|---|---|
| City sizes | rank roughly predicts population | George Zipf, 1949 |
| Earthquakes | many small shocks, few huge ones | Gutenberg-Richter law |
| Web links | few pages attract huge inbound links | Barabási-Albert network model |
| Scientific citations | a small share of papers get most citations | Derek de Solla Price, 1965 |
| Wealth | upper tail is far from bell-shaped | Vilfredo Pareto, 1890s |
| Venture capital | one investment can return more than the fund | power-law portfolio logic |
The common mistake is to treat “rare” as “minor.” In a thin-tailed domain, rare events can be ignored most of the time. In a fat-tailed domain, rare events may be the domain.
Why they form
Preferential attachment is the simple engine: things with more attention get more chances to earn attention. A cited paper is easier to find, so it gets cited again. A large city has more jobs, so it pulls more migrants. A rich investor gets access to deals that a smaller investor never sees.
Multiplicative growth is the second engine. If one actor grows by 10%, then 10%, then 10%, while another grows by 2%, then 2%, then 2%, the gap does not add; it compounds. concept compounding is the motion. Power laws are one shape left behind by that motion.
Self-organized criticality is the physics version. Sandpile avalanches, earthquakes, solar flares, and some market moves show event sizes spread across many scales. The small event and the large event may come from the same system, not from separate categories.
What's contested
Not every straight line on a log-log chart is a power law. Cosma Shalizi, Aaron Clauset, and Mark Newman showed in 2009 that many claimed power laws fit lognormal or stretched-exponential distributions just as well. The visual test is too forgiving.
The second argument is causal. A dataset can have a power-law-looking tail without preferential attachment being the cause. Mechanism matters because the intervention changes: regulating financial leverage is not the same as redesigning citation incentives or earthquake codes.
Why this has to do with other realms
Power laws are where concept monetary debasement meets concept fermi paradox. Money systems concentrate claims when returns compound unevenly; cosmic risk concentrates attention because one asteroid, one gamma-ray burst, or one failed biosphere transition can dominate the survival ledger. The shared lesson is uncomfortable: some systems are not governed by the typical case.
They also explain why a wiki grows unevenly. A few pages become hubs, not because the rest are useless, but because some ideas touch more edges. concept first principles is a method page; person naval ravikant is a human node; both become link magnets if they help readers reframe other pages.
An open question
If power laws punish average-case thinking, what should replace the average in domains where the largest observation has not happened yet?
Key Sources
- The Long Tail by Chris Anderson (2006) — popular account of tail economics in media and retail.
- “Power-law distributions in empirical data” by Aaron Clauset, Cosma Rohilla Shalizi, and M. E. J. Newman (SIAM Review, 2009) — the load-bearing warning against lazy log-log fitting.
- Linked: The New Science of Networks by Albert-László Barabási (2002) — accessible account of preferential attachment and scale-free networks.
- “Networks of Scientific Papers” by Derek J. de Solla Price (Science, 1965) — early citation-network evidence.
- Cours d'économie politique by Vilfredo Pareto (1896-1897) — source of the Pareto distribution in income and wealth analysis.
Further Reading
- concept compounding — the engine that turns small rate differences into giant outcome gaps.
- Scale by Geoffrey West — why cities, organisms, and companies show scaling laws with different exponents.
- The Black Swan by Nassim Nicholas Taleb — polemical, useful for thinking about fat tails and broken averages.
- Santa Fe Institute lectures on scaling laws — good doorway into the physics version of the idea.
Abhishek's take
I watch power laws play out in the vendor base every season. The top 10% of suppliers often deliver 60-70% of the range’s revenue, while the long tail of 200 smaller vendors—each critical for niche fills—might together match just one of those top players. The tools I wrote now flag when a buy plan assumes a bell curve in supplier performance; reality is closer to Pareto, and the floor adjusts allocations before the first PO cuts. The mistake isn’t betting big on the head—it’s pretending the tail behaves like the mean.
See Also
- concept compounding
- concept monetary debasement
- concept fermi paradox
- concept first principles
- person naval ravikant