
The Law of Small Numbers: Why Your Brain Fails at Statistics
I teach behavioral design for a living. I’ve spent two decades studying what makes people do things. And I still catch myself falling for the same statistical traps that Daniel Kahneman describes in Thinking, Fast and Slow. Not because I haven’t read the research. Because my brain, like yours, is a story-construction machine. It would rather build a convincing narrative from three data points than sit with uncertainty long enough to ask whether those three data points mean anything at all.
This is my deep-dive into Chapters 9 and 10 of Kahneman’s book, where he covers the Substitution Effect and the Law of Small Numbers. But I’m not just summarizing. I’m going to tell you where I think Kahneman is right, where I think he’s incomplete, and what the Octalysis Framework reveals about why we fall for these traps in the first place. If you want the broader behavioral-design context around these mental shortcuts, pair this with my cognitive biases guide.
Speed Run Notes
The core warning: Your brain would rather tell a clean story from a tiny sample than admit it still does not know.
The 3 patterns: the Substitution Effect, the Law of Small Numbers, and what I call the Certainty Rush.
The Octalysis layer: premature certainty usually feels good because Core Drive 8 hates ambiguity, Core Drive 5 loves social proof, and Core Drive 2 rewards the feeling of having figured it out.
The practical fix: use small samples for big swings, demand bigger evidence for subtle claims, and separate what people do from why they are doing it.
The design lens: if a product or a conversation is manufacturing urgency too fast, check whether Black Hat motivation is hijacking the reader before understanding catches up.
- Pattern 1: The Substitution Effect
- Pattern 2: The Law of Small Numbers
- The Gates Foundation’s $1.7 Billion Mistake
- The “Hot Hand” Myth (And Why It Doesn’t Matter)
- Pattern 3: The Certainty Rush
- Where I Think Kahneman Gets It Wrong
- The Data-Driven Design Trap
- What to Actually Do Differently
- The A/B Testing Truth Nobody Talks About
- The Three Patterns at a Glance
- Frequently Asked Questions
Author Credibility: Yu-kai Chou

Yu-kai Chou created the Octalysis Framework after studying gamification since 2003 — years before the term entered mainstream vocabulary. As a Human-Systems Architect & Behavioral Designer, his framework has been applied by LEGO, Microsoft, Porsche, Coca-Cola, Salesforce, and MrBeast, impacting over 1.5 Billion Users.
Chou has taught the Octalysis methodology at Harvard, Stanford, Yale, Tesla, Google, BCG, and IDEO.
His work has been cited by Harvard, Stanford, MIT, Forbes, Wall Street Journal, Wired, US Department of Energy, NIST, NSF, NCBI, US Department of Education, ClinicalTrials.gov, and Google Scholar — with 3,700+ more academic publications. Explore his books here.
Pattern 1: The Substitution Effect
When your brain encounters a hard question it can’t answer intuitively, it does something sneaky. It silently swaps in an easier question and answers that instead. You never notice the swap happened.
Here’s the experiment that proves it. Researchers asked German students two questions: “How happy are you?” and “How many dates did you have last month?”
When “How happy are you?” came first, there was zero correlation between dating frequency and reported happiness. Students considered their whole life when evaluating happiness.
When “How many dates last month?” came first, there was a strong correlation. Students who had more dates reported being much happier. The dating question became the lens, doing the same mechanical work as an anchor — the first piece of information silently rewriting the answer to whatever question came after. “How happy am I?” got silently replaced by “How’s my dating life going?”
This works with anything. Whatever question comes first hijacks the second one. Ask about finances first, and “How happy are you?” becomes “How’s my bank account?” Ask about family first, and happiness becomes “How are things with my kids?”
The substitution effect explains why entrepreneurs are dangerously overconfident. Ask a founder “What’s the probability your startup will succeed?” and their brain silently substitutes: “How passionately excited am I about this startup?” High excitement produces “80 to 90 percent chance!” with zero supporting data. That is the same mistake behind a lot of bad cognitive-bias blind spots in strategy rooms. Ask “How much work can I get done this month?” and the brain substitutes: “How productive do I feel right now?” Feeling energized at 9 AM on Monday produces wildly optimistic forecasts that ignore the 15 meetings already on the calendar.
The substitution doesn’t announce itself. You don’t feel your brain making the swap. You feel like you answered the actual question.
Pattern 2: The Law of Small Numbers
Our brains believe the Law of Large Numbers applies to small samples — what Tversky and Kahneman first described in 1971 as “belief in the law of small numbers.” We expect a group of 20 people to be statistically representative when it’s actually just noise.
Kahneman demonstrates this with kidney cancer data. The counties with the lowest kidney cancer rates in the United States are rural, sparsely populated, mostly in the Midwest and South. Your brain immediately constructs a story: clean air, fresh food, less stress.
Now here’s the twist. The counties with the highest kidney cancer rates are also rural, sparsely populated, mostly in the Midwest and South.
The same brain that just told a “clean living” story now generates a “poor healthcare, high-fat diet” story. It never notices the contradiction. Low rates and high rates appearing in the same demographic can’t both be caused by that demographic’s lifestyle. The real explanation is simple: small populations produce extreme results in both directions. A county with 100 people where 3 get cancer looks catastrophic. A county with 100 people where nobody gets cancer looks miraculous. Neither result tells you anything. It’s statistical noise, and your brain dressed it up as a causal narrative.
The Gates Foundation’s $1.7 Billion Mistake
The bias has a body count. Kahneman uses the Gates small-schools episode — the case statistician Howard Wainer dissected as “The Most Dangerous Equation” — as a cautionary example of what happens when decision-makers confuse an extreme sample with a reliable pattern.
The seductive version of the story was simple: small schools seemed to overperform, so leaders concluded that smaller automatically meant better. The real lesson is not the exact multiplier. It’s that people saw the right tail and ignored the equally noisy left tail.
The problem: small schools also produce some of the worst outcomes. When you only have 30 students in a graduating class, one exceptional student (or one dropout) skews the entire percentage dramatically. Big schools cluster toward the mean because the Law of Large Numbers is actually working. Small schools scatter to the extremes because it isn’t.
Now, here’s where I push back on Kahneman’s presentation. If someone showed me data that small schools have BOTH the best AND the worst outcomes, I’d be convinced it’s statistical noise. But Kahneman only presented the “small schools are best” finding before revealing the flaw. He didn’t present the full picture. That’s a bit ironic for someone writing about how we jump to conclusions from partial data.
I also suspect there’s a genuine goldilocks effect at work. Smaller classrooms probably do help up to a point, because each student gets more attention. Malcolm Gladwell makes this case with his inverted U-curve in David and Goliath: quality improves as classes shrink, until they get too small and lose the benefits of peer dynamics and diverse interaction. The statistical noise explanation is correct but probably incomplete.
The “Hot Hand” Myth (And Why It Doesn’t Matter)
Basketball researchers studied whether the “hot hand” is real. After a player scores three or four shots in a row, is their next shot more likely to go in?
The data has generally been much less supportive than fans assume. A scoring streak feels meaningful, but repeated analysis has often found far less evidence for a stable hot-hand effect than commentators want to believe.
The coach of the Boston Celtics heard the finding and responded: “Who is this guy? I couldn’t care less.”
I’m actually with the coach on this one. Here’s why. Even if the “hot hand” doesn’t exist statistically, the player who just scored four in a row almost certainly has a high average shooting percentage. You should still give them the ball. The practical strategy doesn’t change whether or not the hot hand is real.
This is the question I always ask after learning about a cognitive bias: does knowing about this bias actually change what I should do? If the answer is “not really,” the insight is intellectually interesting but practically useless. If the answer is “yes, dramatically,” then it’s worth restructuring your decision-making process around it.
The Law of Small Numbers passes this test. Knowing about it should fundamentally change how you evaluate data, hire people, and interpret results. The hot hand? It’s a fun fact that changes nothing about basketball strategy.
Pattern 3: The Certainty Rush (The Three Core Drives Behind Bad Conclusions)
Here’s something Kahneman doesn’t explain but the Octalysis Framework makes clear: there’s a three-step motivational sequence driving our rush to premature conclusions.
Step 1: Core Drive 8 (Loss and Avoidance). Uncertainty feels like loss. Not knowing the answer to something creates real psychological discomfort. Your brain wants to resolve that discomfort the way it wants to resolve any threat: immediately.
Step 2: Core Drive 5 (Social Influence and Relatedness). To settle the discomfort, we look for a credible source. If someone we trust says “that study is legit,” we immediately drop all doubt. It doesn’t matter if their authority is irrelevant to the topic. A doctor’s opinion on economics carries the same settling power as an economist’s, because the brain isn’t evaluating expertise. It’s seeking relief from uncertainty.
Step 3: Core Drive 2 (Development and Accomplishment). Once we’ve concluded, we feel smart. “I learned something today.” This feeling of mastery rewards the conclusion itself — and the Peak-End Rule ensures the satisfaction of resolution gets remembered as the dominant emotional beat — making us actively resist reopening the question. Saying “I don’t know” would mean giving up that satisfying feeling of understanding.
This three-drive sequence, Core Drive 8 (CD8): Loss and Avoidance, Core Drive 5 (CD5): Social Influence and Relatedness, and Core Drive 2 (CD2): Development and Accomplishment, explains why misinformation spreads so effectively. Someone shares a fabricated claim. Your brain feels uncertain (CD8). A credible-seeming source validates it (CD5). You now “know” something and feel smart about it (CD2). Even when the claim is debunked, the emotional associations persist because System 1 already coded it as true.
| Stage | What your brain feels | Why it becomes dangerous |
|---|---|---|
| CD8: Loss and Avoidance | Uncertainty feels like a threat that needs immediate closure. | You stop asking whether the data is sufficient and start hunting for relief. |
| CD5: Social Influence | A trusted voice makes the answer feel settled. | Borrowed confidence gets mistaken for evidence. |
| CD2: Accomplishment | Conclusion feels like mastery. | Changing your mind now feels like losing status and progress. |
Here’s how I demonstrate this. I’ll tell someone: “Drinking water right after eating watermelon is bad for you.” Complete fabrication. But it sounds plausible, it has a specific-enough detail (the timing, the combination), and if a health-conscious friend shares it, the three-drive sequence fires instantly. Even after I reveal it’s made up, the next time that person eats watermelon and reaches for water, a tiny hesitation will be there. System 1 doesn’t un-learn easily.
Where I Think Kahneman Gets It Wrong
I have enormous respect for Kahneman’s work. But I disagree with him on one significant point.
Kahneman says our tendency to recognize patterns in random data is an evolutionary error. I think that’s only half right.
Imagine you’re an early human and lions have appeared near your camp three times in five months instead of the usual once. Is the increased frequency random? Probably. Should you become more alert anyway? Absolutely. The cost of being more alert when there’s no actual threat is low. The cost of ignoring a real pattern is death. Our brains are calibrated for asymmetric risk, not statistical accuracy.
Here’s the deeper point: we only recognize a handful of broad pattern shapes while everything else lumps together as “random.” So when we spot a recognizable pattern, our instinct is to over-weight it just because it feels familiar. Our brains aren’t wrong that a recognizable pattern deserves attention. They’re wrong about how much certainty to assign to it.
The evolutionary error isn’t pattern recognition. It’s certainty inflation. We’re right to notice the pattern. We’re wrong to be sure about what it means.
The Data-Driven Design Trap
This connects directly to a problem I see constantly in product design and gamification consulting.
“Data-driven design” has become the gold standard for decision-making. Test everything. Measure everything. Let the data tell you what works. In principle, this is sound. In practice, it produces a specific failure mode: you optimize for the metrics without asking what’s producing them.
If your app shows users coming back 10 times a day, the data says they love it. But when you dig into which Core Drives are producing those visits, you might find it’s Core Drive 6: Scarcity and Impatience and Core Drive 8: Loss and Avoidance. People come back 10 times because they’re afraid of missing something, not because they enjoy the experience. That is why I always separate visible behavior from motivational structure and from the cocktail effects between Core Drives.
The behavioral data looks identical. The motivational reality is completely different. One pattern produces long-term engagement. The other produces burnout and resentment.
You have to be multidimensional: behavioral data tells you what people DO, motivational analysis tells you why they do it, and long-term experience quality tells you whether they’ll keep doing it. Data-driven design without motivational analysis is like driving with a speedometer but no map.
What to Actually Do Differently
Knowing about these biases is only useful if it changes your behavior. Here’s what I’ve changed in my own practice.
Test dramatically different ideas first. Small samples are actually fine when you’re testing big differences. If you open a store and punch every customer in the face, and 10 out of 10 leave angry, you don’t need 10,000 data points to conclude that punching customers is bad. When the effect size is dramatic, small samples work. Small samples only fail when you’re testing subtle variations, like whether a blue button or green button converts better. So test the bold ideas first (where small samples are reliable), then optimize the details (where you need large samples).
Always ask: “Does this change my strategy?” Not every insight is actionable. The hot hand myth is fascinating but changes nothing about basketball. The Law of Small Numbers should change everything about how you evaluate data. Filter your cognitive bias knowledge through this question before reorganizing your life around it.
Notice when you feel certain. Certainty is the feeling, not the evidence. When you catch yourself feeling sure about a conclusion you just reached, that’s the Core Drive 8 → Core Drive 5 → Core Drive 2 sequence completing. Pause. Ask: how much data am I actually basing this on? Would I bet my salary on it? The gap between your confidence and your willingness to bet is the gap between System 1’s story and statistical reality.
Distrust narratives from small samples. The next time someone tells you a story about how a small group proves a trend (“Our pilot program with 15 users showed amazing results!”), remember the kidney cancer counties. Small samples produce extreme results in both directions. The story is almost certainly incomplete.
The A/B Testing Truth Nobody Talks About
Most A/B tests in the tech industry run with insufficient sample sizes. Teams ship the “winning” variation after a few hundred conversions and call it data-driven.
Here’s the uncomfortable truth: tiny lifts on tiny samples are usually storytelling bait, not reliable learning. If you see a subtle difference after a small burst of traffic, treat it as a prompt for more testing, not as permission to declare victory.
But when two designs are far apart — one makes the offer dramatically clearer or removes a major friction point — even a smaller sample can tell you something useful. The effect size matters as much as the raw count.
| Test shape | What a small sample can tell you | What it cannot tell you yet |
|---|---|---|
| Huge design difference | Whether one concept is obviously healthier than the other. | The precise long-term lift once novelty wears off. |
| Small UI tweak | Almost nothing reliable beyond “keep collecting data.” | Whether a 1 to 3 percent movement is real or just noise. |
| Mid-sized behavior change | A directional read if the mechanism is strong and the context is clean. | A precise forecast you should bet the roadmap on. |
The practical rule: don’t A/B test small variations with small samples. Either get enough traffic to test subtle differences reliably, or test dramatic differences where small samples are informative. The middle ground, subtle differences with small samples, is where the Law of Small Numbers eats your data budget for breakfast.
The Three Patterns at a Glance
The Substitution Effect: Your brain swaps hard questions for easy ones without telling you. Whatever topic is primed first becomes the lens for everything after. When making important decisions, deliberately reframe the question from multiple angles before accepting your first answer.
The Law of Small Numbers: Small samples produce extreme results in both directions. Your brain interprets these extremes as meaningful patterns. When evaluating data from small groups, remember: the more dramatic the result, the more likely it’s noise.
The Certainty Rush: Uncertainty triggers Core Drive 8 discomfort, which drives you to seek Core Drive 5 social proof, which delivers Core Drive 2 mastery satisfaction. This three-drive sequence makes conclusions feel rewarding, which makes changing your mind feel like loss. Treat certainty as a warning sign, not a destination.
If you want to keep going, the Octalysis Framework is the system I use to decide which Core Drive sequence is firing in any product, decision, or argument — including the Certainty Rush itself. For the full applied treatment with a decade of case studies, the canonical reference is my book Actionable Gamification: Beyond Points, Badges, and Leaderboards. And if your team is shipping A/B tests where small samples keep producing confident wrong answers, reach out — I take a small number of advisory engagements each quarter.
Frequently Asked Questions
How does the Law of Small Numbers differ from confirmation bias?
Confirmation bias is about seeking evidence that supports what you already believe. The Law of Small Numbers is about treating insufficient evidence as conclusive regardless of what you believe. You can fall for the Law of Small Numbers without any prior belief at all. The kidney cancer example works precisely because people had no prior theory about cancer rates in rural counties. They constructed a narrative from scratch based on a small, noisy dataset.
Can the Substitution Effect be used intentionally in design?
Yes, and it routinely is. Surveys that ask “How satisfied are you with our customer service?” before “Would you recommend us to a friend?” are using substitution deliberately. The specific satisfaction question primes the general recommendation. Designers should understand this mechanism and use it ethically, because it’s happening whether you design for it or not.
How much data is “enough” to avoid the Law of Small Numbers?
It depends on the effect size you’re trying to detect. My rule of thumb is simple: the subtler the difference, the more patient you need to be. Huge differences can show up early. Mid-sized differences need more disciplined repetition. Tiny differences demand a lot more traffic than most teams want to admit. The critical error is treating all tests as if they need the same sample size. Match your sample discipline to the boldness of the change.
How does this connect to the Octalysis Framework?
The Certainty Rush maps directly to a three Core Drive sequence: Core Drive 8 (CD8): Loss and Avoidance creates discomfort with uncertainty, Core Drive 5 (CD5): Social Influence provides a shortcut to resolution, and Core Drive 2 (CD2): Development and Accomplishment rewards the conclusion. Understanding this sequence helps designers build systems that encourage thoughtful analysis rather than premature certainty.
Does Kahneman address how to fix these biases?
Kahneman is more diagnostic than prescriptive. He’s excellent at documenting the failures of System 1 but offers limited practical guidance for overcoming them. That’s the gap I try to fill: once you know the bias exists, what do you actually change about your decision-making process, your data analysis workflow, or your design practice? The practical rules in this article are my answer to that gap.

