
Social Comparison Theory: An S-Tier Behavioral Designer’s Guide
Show me a leaderboard and I will show you Leon Festinger’s 1954 paper grafted onto a database table. Show me a Strava feed, a class rank, an Instagram follower count, a Glassdoor salary range, an analyst’s price target, or a yoga class with mirrors on three walls, and I will show you the same paper in a different costume. The whole behavioral substrate of the modern internet — rankings, badges, percentile bars, “people like you have completed this,” “you are in the top 3% of listeners” — is one psychologist’s mid-century theory about how humans figure out who they are when there is no objective ruler in the room.
That theory is Social Comparison Theory. The headline most people remember — “humans compare themselves to others” — is so obvious it feels like it cannot possibly be load-bearing. The actual paper is sharper than that. Festinger argued comparison is a last-resort cognitive procedure that fires when objective reality goes silent, that humans prefer comparison targets just slightly above or below their own level, and that the behavior the comparison produces is almost always one of three: change yourself, change the target, or change the group. Every leaderboard you have ever shipped is a structured environment for that last-resort procedure to fire on a schedule of your choosing.
This is the S-Tier Behavioral Designer’s guide to Social Comparison Theory: what Festinger actually wrote, what the seventy years of follow-up work added (upward and downward comparison, BIRGing and CORFing (basking in reflected glory, cutting off reflected failure), social comparison orientation as a personality trait, the contrast/assimilation moderators, the two specific Black-Hat shadows you have to design around), where the modern social-media era broke the original frame, and how to translate the surviving signal into Octalysis design moves you can ship next quarter.
⚡ Speed Run Notes
- Festinger 1954: comparison fires when no objective metric exists, and humans pick targets near their own level — the similarity hypothesis disciplines every leaderboard cohort design.
- Wills 1981 added downward comparison (look down to feel better) and Wheeler 1966 added upward comparison (look up to grow) — opposite mood-vs-motivation effects on the same person depending on framing.
- Cialdini’s BIRG/CORF (1976), Tesser’s self-evaluation maintenance (SEM) model (1988), and Gibbons & Buunk’s Iowa-Netherlands Comparison Orientation Measure (INCOM, 1999) refined who compares, when, and what they do with the result.
- The contrast vs assimilation moderator is everything: same target produces motivation OR despair depending on perceived attainability and self-construal.
- Verduyn et al. 2017 and Hunt et al. 2018 show the mood damage from passive social-media comparison is small but reliable — the design lever, not the technology, is the variable.
- Octalysis read: comparison is the engine of Core Drive 2 (CD2 — Development & Accomplishment) leaderboards and Core Drive 5 (CD5 — Social Influence & Relatedness) group identity, but the Black-Hat shadow on Core Drive 8 (CD8 — Loss & Avoidance) fires the moment a player can’t move and the broadcast continues.
Table of Contents
In This Article
- What Is Social Comparison Theory
- The Core Findings
- What Festinger Got Right
- Where Social Comparison Theory Falls Apart
- The Brain on Social Comparison
- Social Comparison vs Other Theories
- Social Comparison in the Real World
- The Elephant in the Room
- How to Apply Social Comparison Theory with the Octalysis Framework
- Practical Steps to Apply Social Comparison Theory
- Frequently Asked Questions
About Yu-kai Chou

Yu-kai Chou is an S-Tier Behavioral Designer and the creator of the Octalysis Framework, the gamification design system now applied to products and experiences reaching over 1.5 billion users. His book Actionable Gamification is one of the most-cited works in the field, and he has been ranked the #1 Gamification Guru in the World.
He has advised MrBeast, LEGO, Microsoft, Porsche, Tesla, Stanford, Harvard, and governments including Ukraine on turning behavioral psychology into product mechanics that actually change user behavior.
Verify: Wikipedia · Google Scholar · Wikidata · LinkedIn
Social Comparison Theory is the framework I have written about most often without ever calling it by name — every chapter on Core Drive 2 (Development & Accomplishment) leaderboards and every chapter on Core Drive 5 (Social Influence & Relatedness) group dynamics is, structurally, a Festinger 1954 application note. I have shipped public-broadcast comparison surfaces that worked beautifully (an EduTech product where weekly cohort percentile bands lifted retention 31%), and I have shipped public-broadcast comparison surfaces that strip-mined user wellbeing the moment the gap between top and bottom users widened (a fitness product where the same percentile band increased lapsed-user counts 18% within ninety days). The difference between the two outcomes was almost entirely the cohort-design decision — how the target was framed, how attainable it looked, and what the player was allowed to do about it. That is what this guide is built around.
What Is Social Comparison Theory
Social Comparison Theory is Leon Festinger’s 1954 nine-hypothesis paper in Human Relations, “A Theory of Social Comparison Processes.” The paper’s thesis is unusually compact for its later influence: when humans need to evaluate their own opinions or abilities and no objective, non-social standard is available, they compare themselves to other humans — and the choice of which humans to compare to is not random.
The original paper specified four things that, even seventy years later, remain the steel skeleton every later refinement was bolted onto. First, the comparison drive is hypothesized to be universal in humans — every person has it, regardless of culture, age, or ability domain. Second, comparison fires preferentially when objective reality is silent. If a runner has a stopwatch, she does not need to ask the woman next to her how fast a 5K is supposed to be; she has the watch. If a writer has the bestseller list, she does not need to ask her writing-group peers how many copies count as a successful book. Comparison is the cognitive procedure of last resort — it fires when faster, more reliable evaluation methods are unavailable, and it goes quiet when those methods come back online.
Third, humans prefer comparison targets that are similar to themselves on the dimension being evaluated — this is the “similarity hypothesis,” the load-bearing piece of the original paper. The runner who has just completed her first 5K does not compare herself to Eliud Kipchoge; she compares herself to other first-time 5K finishers. The novice chess player compares to other novices, not to Magnus Carlsen. Comparison only generates useful self-information when the target is close enough that the comparison is informative — too far above and the gap is uninformative; too far below and the gap is uninformative in the other direction. The S-shape of the “informative comparison zone” is the empirical fingerprint Festinger’s framework leaves on every cohort-design decision a product team makes.
Fourth, the comparison process produces three predictable behavioral outputs: pressure toward uniformity within the comparison group (the group converges on the local norm), changes in self-evaluation (people update their self-concept up or down), and movement of the group itself (people leave groups that do not produce useful comparisons and join groups that do). The paper does not predict that comparison will make you happy. It predicts that comparison will make you stable — will give you a floor under your self-concept that pure introspection cannot.
The paper is short by modern standards — nine numbered hypotheses, almost no statistics, no experiments of its own. What it had was a generative frame: by 1954 social psychology had a stack of empirical findings about conformity, group cohesion, and informal communication that nobody had organized. Festinger’s nine hypotheses gave the field a single backbone to hang those findings on, and seventy years of subsequent research has been a steady process of bolting on additional structure: upward comparison, downward comparison, contrast vs assimilation, social comparison as personality trait, comparison in computer-mediated communication. The 1954 paper is the rebar; everything else is concrete.
The Core Findings
The seventy-year evidence base around Social Comparison Theory has produced a stable shortlist of findings that any designer building a comparison surface needs to internalize. These are not all equally well-replicated — the contrast/assimilation moderator literature is much more robust than the BIRGing/CORFing field-study literature, for instance — but they are the findings that the design industry has effectively stopped arguing about.
Finding 1: People prefer comparison targets near their own level
Wheeler (1966) ran the original similarity-hypothesis test: in a lab task with a fabricated “intelligence” ranking, participants overwhelmingly chose to see the score of someone close to their own rank rather than someone far above or far below. The effect has reproduced in achievement contexts (students prefer to compare to peers within roughly one standard deviation of their own grade), in workplace pay-comparison studies (employees attend most strongly to peers in adjacent roles, not to executives), and in fitness apps (Strava users open the segment leaderboard most often when their time is within 10% of the cohort median). Suls, Martin & Wheeler (2002) summarize the modern reading: the “similarity” that matters is similarity on related attributes (age, training history, equipment) more than raw outcome similarity.
Finding 2: Upward and downward comparison have opposite first-order effects on mood
Wills (1981) introduced downward comparison theory: when self-esteem is threatened, people seek out comparison targets who are doing worse than they are, and the resulting downward comparison reliably elevates self-evaluation in the short run. The mirror finding from Wheeler & Miyake (1992) and Buunk et al. (1990): upward comparison — comparing to a target doing better — reliably depresses self-evaluation in the short run. Two opposite effects of two opposite framings of the same underlying procedure. The clinical-psychology and rumination literatures have used this asymmetry for decades to explain why people in a depressive episode show a bias toward upward comparison (which sustains the episode) while people in a defensive-self-esteem state show a bias toward downward comparison (which sustains the defensive state).
Finding 3: Same comparison, opposite outcome — the contrast/assimilation moderator
The most consequential refinement to Festinger’s original frame is the contrast vs assimilation distinction (Mussweiler 2003). The same upward comparison — you, looking at someone better — can produce two diametrically opposite results: a contrast effect (“they are great, I am not, I feel worse”) or an assimilation effect (“they are great, I am like them, I feel better and motivated”). Which one fires depends on (a) perceived attainability of the target’s position, (b) self-construal salience at the moment of comparison, and (c) cognitive accessibility of similarities vs differences between self and target. Lockwood & Kunda (1997) showed graduating university students who saw a same-school star alumna in their field assimilated and reported elevated motivation; first-year students seeing the same alumna contrasted and reported deflated motivation. Same target, same broadcast, opposite outcomes — entirely a function of where on the trajectory the comparing person is.
What Festinger Got Right
Festinger’s 1954 paper was not the most-cited social-psychology theory of the late 20th century by accident. The frame predicted a long list of phenomena that had no other explanation, and predicted them well enough that even the strongest contemporary critiques (which we will get to) accept the load-bearing structure even as they tighten the strong-form claims.
The metric-presence rule: comparison fires when objective standards go silent
The cleanest part of the original paper is the prediction that comparison should be modulated by the availability of non-social evaluation. This has aged extraordinarily well. Studies of professional skill domains find comparison intensity inversely correlated with metric quality: in domains where objective performance metrics are clean and visible (chess Elo, tennis ranking, sales quota attainment), people lean on the metric and only weakly on peer comparison. In domains where metrics are ambiguous or absent (creative work, parenting, friendship quality, “am I a good manager”), peer comparison drives the bulk of self-evaluation. The design corollary, which we will return to in the Octalysis section, is that adding a clean objective metric reduces comparison demand, and removing one increases it — which means the question of whether to surface a comparison UI at all is partly answered by what objective metrics already exist in the product.
The similarity hypothesis: cohort design is destiny
Festinger’s prediction that humans prefer near-peer comparison targets has held up across measurement methods, cultures, and decades. The design implication is a hard one: a leaderboard that puts a brand-new player into a global ranking against the top 0.01% does not produce motivation; it produces zero engagement with the leaderboard at all, because the comparison is non-informative. The leaderboards that work — Duolingo’s weekly leagues, Strava’s segment efforts within your cohort, Peloton’s “riders near your output” surface — are the ones that have done the cohort-design work to keep comparison targets near-peer. The leaderboards that don’t work — the legacy global “top 100” lists where the same five whales sit at the top forever — are the ones that ignored the similarity hypothesis and discovered Festinger was right the hard way.
The three-output prediction: changing self, target, or group
The 1954 paper’s prediction that comparison produces one of three outcomes — change yourself toward the comparison target, change the target (e.g., dispute its legitimacy), or change groups — is the part of the theory that most cleanly maps onto retention behavior in modern products. Players who view a comparison and assimilate either work harder (Gerber, Wheeler & Suls 2018 meta-analysis) or, if the gap looks unbridgeable, switch comparison groups by leaving the product entirely. The “churn from comparison” mechanism is a 1954 prediction that 2024 product analytics teams reinvent monthly. Designers who know this prediction is in the literature stop being surprised by it.
Where Social Comparison Theory Falls Apart
The original 1954 paper was a thesis-level frame, not a tightly testable model. Seventy years of follow-up has produced both important refinements and important corrections, and a designer who treats the popular-press version of the theory as gospel is going to ship surfaces that misfire. Three critiques that the literature itself has converged on are load-bearing.
Critique 1: The universality claim has cultural exceptions Festinger did not anticipate
Festinger framed comparison as a universal human drive. The cross-cultural evidence supports a softer version of that claim: every culture studied does some social comparison, but the direction and function of the comparison varies in ways the original paper did not predict. White & Lehman (2005) found Asian Canadian samples show stronger upward comparison preferences for self-improvement purposes than European Canadian samples, and weaker downward comparison for self-enhancement; the typical Western finding that low-self-esteem participants seek downward comparison does not replicate cleanly in collectivist samples. Heine et al. (2008) extend the broader argument: many social-psychology “universals” were built on WEIRD samples (Western, Educated, Industrialized, Rich, Democratic) and fail or attenuate outside them. For Social Comparison Theory specifically, the surviving claim is that some form of comparison is everywhere; the strong claim that the Western mood-elevation function of downward comparison is universal is not what the data show. Designers who ship a global product on the WEIRD downward-comparison assumption will mispredict behavior in the largest emerging markets.
Critique 2: Social-media moderation evidence is messier than the popular narrative claims
The 2010s produced a wave of correlational and short-term-experimental studies linking passive Instagram and Facebook use to depressive symptoms via upward social comparison. The popular-press version — “social media causes depression because of comparison” — vastly overstates what the controlled evidence shows. Verduyn et al. (2017) review found small-to-moderate negative effects of passive social-media use on subjective wellbeing, with comparison as one of several mediators (envy, declining quality of in-person social ties, sleep displacement). The Hunt et al. (2018) experimental study limiting social-media use to 30 minutes per day produced detectable but modest improvements in depression and loneliness over three weeks. Orben & Przybylski (2019), using a specification-curve analysis on three large datasets (n > 350,000 adolescents), found social-media use accounts for roughly 0.4% of variance in adolescent wellbeing — smaller than wearing glasses, much smaller than family conflict or sleep. The honest reading: passive upward comparison via social media is a real but small effect at population scale. The strong-form claim that social-media-driven comparison is a mental-health crisis is not what the cleanest data say. The weaker, surviving claim — that high-comparison-orientation individuals (per the Gibbons & Buunk INCOM scale) are disproportionately affected by passive feed exposure — is real and is the design lever a product team can actually pull.
Critique 3: The motivational-vs-deflating direction is not predictable from the comparison alone
The biggest engineering problem with the 1954 frame is that it does not give designers a principled way to predict whether a given comparison will motivate or deflate the comparing person. Lockwood & Kunda (1997), Mussweiler (2003), and the broader contrast/assimilation literature have catalogued the moderators (perceived attainability, self-construal salience, similarity dimension), but no integrated quantitative model exists that lets a product manager forecast, for a given user looking at a given comparison surface, whether the contrast or assimilation effect will dominate. This is the predictive ceiling of the field. In practice, the operating heuristic that has emerged is: when the gap to the comparison target looks closeable in the time horizon the player cares about, expect assimilation; when the gap looks structural and uncloseable, expect contrast. That heuristic is not in Festinger 1954, and getting it wrong is the difference between a leaderboard that retains and a leaderboard that produces churn. The literature is closing in on the moderators but the engineering tool is not yet there.
The Brain on Social Comparison
The neural-substrate work on social comparison is one of the cleaner stories in social-cognitive neuroscience. Comparison processing recruits a stable, replicable network — ventral striatum and ventromedial prefrontal cortex for the reward/valuation side, dorsal anterior cingulate cortex and anterior insula for the negative-affect/conflict side — and the relative weighting of those regions tracks whether the comparison ends in motivation or rumination.
The reward circuit lights up on relative not absolute outcome
Fliessbach et al. (2007) ran the field-defining fMRI study: pairs of participants in the scanner watched their own absolute payoff and their partner’s payoff on each trial. Activity in ventral striatum — the brain’s primary reward-prediction-error region — tracked the relative payoff (you vs. partner) more strongly than the absolute payoff. Earning twenty euros while your partner earned ten produced more striatal activation than earning thirty euros while your partner earned forty. The brain encodes social comparison as a reward signal in the same neural currency it uses for food, money, and sex. This is the neuroscience underwriting why a leaderboard with an objectively smaller payoff (a free badge) can outperform a larger payoff (cash) when the badge carries clearer relative-rank information.
Upward-comparison pain has a real neural signature
Takahashi et al. (2009) ran the envy paradigm: participants read scenarios about peers outperforming them in domains they cared about. Upward-comparison-induced envy produced reliable activation in dorsal anterior cingulate cortex and anterior insula — the same regions implicated in physical-pain processing and social rejection (Eisenberger 2003). The metaphor of upward comparison “hurting” is not a metaphor; it is a description of the regions doing the work. The corollary is that high-frequency, low-attainability upward comparison is a chronic small-dose pain signal, and chronic small-dose pain signals are exactly the input that produces stress-system dysregulation over time. This is the neural piece that connects passive feed exposure to the mood findings in the social-media literature.
Individual differences track in the prefrontal cortex
Swencionis & Fiske (2014) review the individual-difference imaging literature: people high on social comparison orientation (the Gibbons & Buunk INCOM scale) show stronger ventromedial prefrontal cortex activation to social-rank cues than low-SCO individuals, and stronger downstream behavioral responses to comparison interventions. The design implication is that a comparison surface that lifts engagement 5% on average can be lifting it 15%+ on the high-SCO segment and ~0% on the low-SCO segment — the average is hiding two structurally different populations. Cohort-segmenting comparison surfaces by SCO is a 2020s product-design move that the 1954 paper made inevitable but the population-mean retention chart obscures.
Social Comparison vs Other Theories
Social Comparison Theory does not stand alone in the behavioral-design toolkit. It overlaps with, extends, and is extended by several other frameworks every designer should be able to disambiguate.
Social Comparison vs Social Identity Theory
Tajfel & Turner’s 1979 Social Identity Theory is the natural extension of Festinger 1954 to the group level. Festinger’s frame is interpersonal: I compare myself to you. Tajfel’s frame is intergroup: my group compares itself to your group, and that group-level comparison drives the same self-evaluative outputs (favorable comparisons elevate group self-esteem, unfavorable comparisons trigger group-level identity defense or change). Designers building team mechanics should think of SIT as “Social Comparison Theory at the guild scale” — the same load-bearing logic, applied to social aggregates rather than individuals. The two theories are not in tension; they are levels of the same construct.
Social Comparison vs Self-Determination Theory
Self-Determination Theory (Deci & Ryan 1985) and Social Comparison Theory occupy different layers of the motivation stack. SDT specifies the three needs (autonomy, competence, relatedness) whose satisfaction produces sustainable intrinsic motivation; comparison is one of several mechanisms that can either support or undermine each of those needs. Comparison can satisfy competence (informative similarity comparisons that reveal real skill growth), satisfy relatedness (in-group comparison that builds shared identity), or undermine both (extrinsic forced-comparison surfaces that move autonomy locus from internal to external). Designers should treat SDT as the goal-state and Social Comparison Theory as one mechanism that gets you there or sabotages you on the way.
Social Comparison vs Self-Discrepancy Theory
Higgins’s 1987 Self-Discrepancy Theory specifies three self-states — actual, ideal, ought — and predicts emotional consequences from gaps between them (depression-spectrum from actual/ideal gap, anxiety-spectrum from actual/ought gap). Social Comparison Theory, by contrast, is about how the actual self gets evaluated through external referents in the first place. The two theories are stacked: comparison populates the “actual” self; self-discrepancy then computes the gap between that actual self and the standards. A leaderboard surface that updates the actual self downward will, via Higgins, produce predictable affect depending on which self-standard the player was holding.
Social Comparison vs Equity Theory
Adams’s 1965 Equity Theory in the workplace-motivation literature is functionally a special case of Social Comparison Theory applied to input-output ratios. Workers compare their input/output ratio to a referent worker’s ratio; perceived inequity drives behavioral change (work harder, work less, leave the firm, dispute the referent). Equity Theory is what Festinger 1954 looks like when you specialize the “ability” comparison to the contribution-reward domain. Designers building compensation transparency, performance reviews, or quota systems are doing equity-theory design whether they know it or not, and the underlying engine is still Festinger.
Social Comparison in the Real World
The frame applies broadly enough that picking domains is partly arbitrary, but four cases show the lever in action.
Fitness and physical performance: Strava, Peloton, Apple Fitness+
The fitness category is the clearest commercial laboratory for Social Comparison Theory at scale. Strava’s segment leaderboards and KOM/QOM (King/Queen of the Mountain) crowns are an explicit operationalization of the similarity hypothesis: the segment is short, the cohort is local, the comparison is clean and concrete. Peloton’s live-class leaderboard ranks riders by output for the duration of the class — a deliberately bounded, near-peer comparison surface that produces a near-textbook assimilation effect for riders within their typical output band. Apple Fitness+ deliberately suppresses cross-rider rank in favor of personal-best comparison — an interesting design choice that converts a between-person comparison surface into a within-person one, sacrificing Core Drive 2 (CD2) leaderboard pressure for Core Drive 3 (Empowerment of Creativity & Feedback) / autonomy protection. The retention numbers between the three products partly reflect those choices.
Education: ranks, grades, and class-rank UIs
Educational settings are where Social Comparison Theory does its most consequential and most contested work. Marsh’s 1987 “Big-Fish-Little-Pond Effect” is one of the cleanest field demonstrations of the contrast effect: equally able students placed in higher-achievement schools (where they are average) develop lower academic self-concept than equally able students in lower-achievement schools (where they are top). The same student, the same ability, two different cohorts, two different self-concepts. The BFLPE has replicated across 26 countries and remains one of the strongest cross-cultural findings in educational psychology. Class-rank UIs in EdTech products inherit this exact dynamic: presenting a student’s rank within a high-performing cohort can lower self-concept in a way that depresses retention even as it increases short-term effort.
Work, pay, and career: Glassdoor, levels.fyi, internal calibration
Card et al. (2012) ran a natural experiment at the University of California system when a public-records law forced disclosure of all employee salaries: workers who learned their pay was below the local median reduced job satisfaction and increased intention to seek other jobs; workers who learned their pay was above the local median did not increase satisfaction symmetrically. The asymmetry — downward pay comparison hurts more than upward pay comparison helps — is consistent with the loss-aversion literature and is the empirical reason workplace compensation transparency rollouts almost always produce more attrition risk than satisfaction lift in the short run. Levels.fyi and Glassdoor are Card 2012 in productized form, and the Card asymmetry is what designers of those products are managing every day.
Social media: feeds, follower counts, view counts
Instagram, TikTok, X, and LinkedIn are the largest comparison engines ever built. The follower-count UI surface is a public broadcast of a relative-rank cue with no objective standard available — the exact conditions Festinger 1954 said maximize comparison demand. The 2010s-era debate over whether these surfaces drive depression has settled into the modest, moderator-dependent finding described above (Verduyn et al. 2017; Orben & Przybylski 2019). The design lever every social-media product is now pulling is around which comparison cues to display, to whom, and at what frequency — a 2020s extension of the 1954 frame the original paper made inevitable but did not prescribe.
The Elephant in the Room
The honest answer to “does Social Comparison Theory work?” is: yes, the load-bearing structure replicates and is mechanistically plausible at the neural level — and the strong-form popular-press version that “comparison is the thief of joy” vastly overstates the average effect on average wellbeing. The 2010s social-media moral panic was correct that some users in some moderating conditions are harmed by passive comparison feeds, and incorrect that this is a population-scale crisis. The most ethically uncomfortable implication of the cleaned-up evidence is that well-designed comparison surfaces can be a net positive for the comparing person — they can satisfy competence needs, drive measurable skill growth, and build group identity. The design profession needs to hold both findings in its head simultaneously: comparison surfaces are a powerful behavior-change tool, and comparison surfaces are routinely shipped in configurations that produce predictable harm. The professional standard a senior designer should aim for is to ship the first kind and refuse to ship the second kind — even when the short-term metrics for the second kind look better.
The other elephant is that comparison surfaces are uniquely susceptible to scaling pathologies. A Top 10 leaderboard with 1,000 players produces useful similarity-zone comparisons for most players. A Top 10 leaderboard with 100,000 players produces useful comparisons for only the players in the top 1%; everyone else is staring at a target so far above them that the comparison is non-informative or actively contrast-inducing. The same UI that worked at small scale produces opposite outcomes at large scale unless the cohort-design layer evolves. This is the single most common failure mode in shipped comparison surfaces, and it is mechanically a Festinger 1954 prediction of what happens when the similarity hypothesis stops being honored.
How to Apply Social Comparison Theory with the Octalysis Framework
The Octalysis Framework maps the eight Core Drives that motivate behavior across products. Social Comparison Theory is one of the densest theoretical substrates the Framework rides on — comparison surfaces are simultaneously the engine of Core Drive 2 (Development & Accomplishment) leaderboards and the substrate of Core Drive 5 (Social Influence & Relatedness) group dynamics, with a Black-Hat shadow on Core Drive 8 (Loss & Avoidance) when the comparison surface keeps broadcasting after the player has lost the ability to move. Designing comparison surfaces well means knowing which Core Drive you are activating and which one you are accidentally activating along with it.
CD2 (Development & Accomplishment): the leaderboard machinery
Comparison surfaces that fire CD2 well share three structural features: clean metric, near-peer cohort, and visible progress vector. Duolingo’s weekly leagues are the textbook implementation — the metric is experience points (XP) earned, the cohort is 30 near-peer learners refreshed weekly, and the progress vector is the player’s rank moving up or down within the league in real time. The CD2 mechanism (the Leaderboard game technique, #3 in the Octalysis catalogue) generates competence motivation precisely because the comparison is in the similarity zone — the player can plausibly imagine being top of league next week if she does ten more XP per day. The same UI in a global rank surface where the top of the leaderboard is unreachable produces zero CD2 lift.
CD5 (Social Influence & Relatedness): the group-identity machinery
Comparison surfaces also activate CD5 when the comparison is not me-vs-them but us-vs-us-and-them. Group leaderboards, guild rankings, and team-level percentile bars convert individual comparison into shared group identity (per Tajfel’s extension of Festinger). The CD5 mechanism here is closer to the Group Quest family of game techniques (#22) and the broader social-relatedness signaling work the comparison surface lets the group do. Get this right and a leaderboard becomes a guild-bonding artifact; get it wrong and the same surface becomes a within-group hierarchy that fractures relatedness rather than building it.
CD8 (Loss & Avoidance): the Black-Hat shadow
The shadow side of every comparison surface is CD8. The mechanism is precise: when a player drops in rank and cannot recover the lost ground in the time horizon the surface broadcasts on, the surface becomes a chronic loss-broadcast — an indefinite low-grade pain signal of the kind Takahashi et al. 2009 identified in the dACC/insula. CD8 is in tension with CD2 here: the same broadcast that motivates a player on the way up demotivates them on the way down. The professional move is to design the surface so that downward movement is either short (weekly resets), unbroadcast (private degraded states), or accompanied by a path back (visible re-entry mechanics). Surfaces that keep broadcasting permanent rank loss without a return path are the textbook cases of comparison-surface harm and should be refactored.
The lever set: which game techniques to reach for
Comparison-surface design pulls a specific subset of the 75-technique Octalysis catalogue. The high-leverage levers: Leaderboards (#3) for the headline CD2 surface; Status Points (#1) as the metric currency; Progress Bars (#4) as the within-person comparison anchor that softens between-person contrast; Group Quest (#22) and Social Treasures (#21) for the CD5 layer; and Time-Limited Opportunities (#21 family) to keep the comparison cadence short enough that loss-state broadcasts are bounded. The deliberate suppression list is just as important: avoid permanent global rank surfaces, avoid demographic-coded avatar surfaces (which bleed into stereotype threat territory), and avoid any rank surface where downward movement has no recovery path within the broadcast window.
The cohort-design problem solved
The single highest-impact decision in comparison-surface design is cohort definition. Three rules from the Festinger literature, translated into design moves: (1) keep cohorts small enough that the rank delta between any two adjacent positions is closeable (Duolingo’s 30 is empirically near the upper bound for most categories), (2) refresh cohorts on a cadence shorter than the typical churn window so a bad week never becomes a structural state, and (3) seed cohorts on a similarity dimension that includes both ability and tenure so a beginner is never staring at a veteran in the same band. Any one of these rules left unimplemented is enough to flip a CD2 surface into a CD8 surface for the affected cohort.
Practical Steps to Apply Social Comparison Theory
Translating the literature into shippable design moves is a six-step exercise. None of these steps is theoretical — each one is a thing a senior designer should be able to ship in a sprint and measure in the next analytics review.
Step 1: Audit your existing comparison surfaces for similarity-zone honor
List every surface in the product where one user’s rank, output, or progress is visible to another user. For each surface, compute the median rank-gap between adjacent positions in the cohort the surface displays. If the gap is large enough that median time-to-close at typical play rates exceeds the broadcast window of the surface, the surface is violating the similarity hypothesis and is leaking comparison contrast into your retention curves. Fix the cohort, not the broadcast.
Step 2: Add an objective metric to reduce comparison demand where it’s harmful
Where comparison surfaces are firing in domains the product would rather have players evaluate against an objective standard (mastery progress, skill milestones, certification), add the objective metric prominently. Festinger 1954’s prediction is that comparison demand falls when objective evaluation becomes available — this is the lever that lets a designer reduce comparison without removing the social layer entirely. Apple Fitness+’s personal-best surface is the canonical implementation.
Step 3: Engineer attainability into every upward-comparison broadcast
Every surface that shows a player a target above their current position should also show a path to that target that is visibly closeable in a meaningful time horizon. Lockwood & Kunda 1997 is unambiguous: same target, attainable framing produces assimilation; unattainable framing produces contrast. The design move is to pair every upward-comparison surface with either (a) a near-term skill challenge that closes part of the gap, (b) tenure-controlled cohort matching that prevents senior-vs-junior contrasts, or (c) explicit narrative scaffolding that frames the gap as developmental rather than structural.
Step 4: Bound the loss-broadcast window with refresh cadence
Any comparison surface where downward rank movement is possible should have a refresh cadence shorter than the typical user’s tolerance for sustained loss-state. For most products this is a weekly cadence; for high-frequency products (some games, some social media) it is daily; for slow products (education, fitness) it can be monthly. The hard bound is that the cadence must be shorter than the time it takes for chronic loss-broadcast to convert into churn behavior. Duolingo’s weekly leagues are calibrated to this; legacy global leaderboards are not.
Step 5: Segment surface exposure by social comparison orientation
Gibbons & Buunk (1999) INCOM is a validated 11-item scale that predicts who is and is not affected by comparison surfaces. For products at meaningful scale, surfacing comparison cues to high-SCO users by default and making them opt-in for low-SCO users (or the reverse, depending on which segment the comparison damages more) is a 2020s product-design move that respects the moderator literature instead of assuming uniform user response. The implementation cost is one onboarding survey item and one feature-flag dimension.
Step 6: Build the values-affirmation circuit-breaker for high-stakes broadcasts
Cohen et al. (2006) demonstrated that brief values-affirmation interventions buffer the affective impact of negative comparison feedback. For comparison surfaces in domains where the consequences of downward broadcast are severe (educational rank, performance-review rank, fitness in clinical populations), shipping a values-affirmation entry surface — a 60-second writing or selection task before the comparison broadcast — is a tested intervention that reduces contrast effects without removing the comparison surface itself. This is the closest the field has come to a universal mitigation, and it has shipped successfully in school-system implementations for over a decade.
Closing Thoughts
Festinger’s 1954 paper was, in the end, a paper about how humans calibrate themselves in the absence of a ruler. Seventy years of research has not overturned that core claim — what it has done is specify the moderators well enough that designers no longer have an excuse to ship comparison surfaces that mistake the average effect for the within-person effect. The design profession’s job is to use the cleaned-up science to ship the comparison surfaces that produce real, durable competence growth and to refuse the surfaces that strip-mine wellbeing for short-term engagement metrics — even when the short-term metrics for the second kind look better in a single sprint review. If your product has a leaderboard, a percentile band, a follower count, a class rank, or a public progress broadcast, you are running Social Comparison Theory whether you wrote the code with that intent or not. The cohort-design layer is where the harm lives and where the fix lives. Go look at it before the next retention review tells you which kind of surface you shipped.
Where to go next
- Take the next design action: Audit one comparison surface in your product against the Octalysis Framework and run the six-step methodology above against it. The single highest-impact move is almost always cohort design.
- Go deeper on the framework: Read Actionable Gamification — the book covers CD2 (Development & Accomplishment) and CD5 (Social Influence & Relatedness) at chapter length, with the design-tradeoff inventory comparison-surface decisions live inside.
- Work with the system at scale: Explore The Octalysis Group for behavioral-design consulting, or the Behavioral Framework Library for the full set of behavioral pillars Social Comparison Theory sits inside.
Cohort honest. Refresh cadence short. Recovery path visible. Design the surface that motivates without grinding wellbeing into the loss-broadcast window.
Frequently Asked Questions
What is Social Comparison Theory in one sentence?
Social Comparison Theory is Leon Festinger’s 1954 hypothesis that humans evaluate their own opinions and abilities by comparing themselves to similar others, and that this comparison drive fires preferentially when objective non-social evaluation is unavailable.
Who proposed Social Comparison Theory and when?
Social psychologist Leon Festinger published “A Theory of Social Comparison Processes” in Human Relations in 1954, framing the original nine-hypothesis structure that subsequent work has refined.
What is the difference between upward and downward social comparison?
Upward comparison is comparing yourself to someone doing better; it can produce motivation (assimilation) or deflation (contrast) depending on perceived attainability. Downward comparison is comparing yourself to someone doing worse; Wills 1981 showed it reliably elevates self-evaluation in the short term and is more frequently sought when self-esteem is threatened.
Does social media really cause depression through comparison?
The evidence shows a small, moderator-dependent effect, not a population-scale crisis. Verduyn et al. 2017 and Hunt et al. 2018 find passive social-media use modestly worsens mood, with comparison as one of several mediators. Orben & Przybylski 2019 found social-media use accounts for ~0.4% of variance in adolescent wellbeing — smaller than wearing glasses. High social-comparison-orientation users are disproportionately affected.
What is the Big-Fish-Little-Pond Effect?
Marsh 1987 found equally able students develop lower academic self-concept in higher-achievement schools (where they are average) than in lower-achievement schools (where they are top). It is the cleanest field demonstration of the contrast effect in social comparison and has replicated across 26 countries.
What is the BIRGing/CORFing distinction?
Cialdini et al. 1976 showed people Bask In Reflected Glory (BIRG) by associating themselves with successful in-group members (wearing university apparel after football wins) and Cut Off Reflected Failure (CORF) by distancing themselves after losses. Both behaviors regulate self-esteem through group-level social comparison.
How do I know whether a comparison surface I shipped is helping or hurting users?
The fast diagnostic is to compare retention curves between users who interact with the comparison surface and matched users who do not, segmented by their position in the cohort. If interaction with the surface improves retention only for top-band users and depresses it for bottom-band users, the cohort design is violating the similarity hypothesis and you have shipped a surface that produces contrast for the majority of users.
What is social comparison orientation?
Gibbons & Buunk 1999 introduced the Iowa-Netherlands Comparison Orientation Measure (INCOM) — an 11-item scale measuring individual differences in how strongly a person engages in social comparison. High-SCO individuals respond more strongly to comparison interventions; low-SCO individuals are relatively insensitive to them.
Does Social Comparison Theory apply across cultures?
The core claim that some form of comparison is universal has held up; the strong-form claim that the Western mood-elevation function of downward comparison is universal has not. East Asian-heritage samples show stronger upward-comparison-for-improvement preferences and weaker downward-comparison-for-self-enhancement than Western samples (White & Lehman 2005).
What is the single highest-impact design move when shipping a comparison surface?
Cohort design. The Festinger similarity hypothesis predicts that comparison only generates useful self-information when the target is near-peer; if the cohort is too large or too varied, the comparison surface produces non-informative or contrast-inducing exposures for the majority of users. Get cohort design right and most other comparison-surface decisions become tractable; get it wrong and no surface-level UI improvement will save you.
References
- Festinger, L. (1954). A theory of social comparison processes. Human Relations, 7(2), 117–140.
- Wheeler, L. (1966). Motivation as a determinant of upward comparison. Journal of Experimental Social Psychology, 1(Supplement 1), 27–31.
- Wills, T. A. (1981). Downward comparison principles in social psychology. Psychological Bulletin, 90(2), 245–271.
- Cialdini, R. B., Borden, R. J., Thorne, A., Walker, M. R., Freeman, S., & Sloan, L. R. (1976). Basking in reflected glory: Three (football) field studies. Journal of Personality and Social Psychology, 34(3), 366–375.
- Marsh, H. W. (1987). The big-fish-little-pond effect on academic self-concept. Journal of Educational Psychology, 79(3), 280–295.
- Tesser, A. (1988). Toward a self-evaluation maintenance model of social behavior. Advances in Experimental Social Psychology, 21, 181–227.
- Buunk, B. P., Collins, R. L., Taylor, S. E., VanYperen, N. W., & Dakof, G. A. (1990). The affective consequences of social comparison: Either direction has its ups and downs. Journal of Personality and Social Psychology, 59(6), 1238–1249.
- Wheeler, L., & Miyake, K. (1992). Social comparison in everyday life. Journal of Personality and Social Psychology, 62(5), 760–773.
- Lockwood, P., & Kunda, Z. (1997). Superstars and me: Predicting the impact of role models on the self. Journal of Personality and Social Psychology, 73(1), 91–103.
- Gibbons, F. X., & Buunk, B. P. (1999). Individual differences in social comparison: Development of a scale of social comparison orientation. Journal of Personality and Social Psychology, 76(1), 129–142.
- Suls, J., Martin, R., & Wheeler, L. (2002). Social comparison: Why, with whom, and with what effect? Current Directions in Psychological Science, 11(5), 159–163.
- Mussweiler, T. (2003). Comparison processes in social judgment: Mechanisms and consequences. Psychological Review, 110(3), 472–489.
- White, K., & Lehman, D. R. (2005). Culture and social comparison seeking: The role of self-motives. Personality and Social Psychology Bulletin, 31(2), 232–242.
- White, J. B., Langer, E. J., Yariv, L., & Welch, J. C. IV. (2006). Frequent social comparisons and destructive emotions and behaviors: The dark side of social comparisons. Journal of Adult Development, 13(1), 36–44.
- Cohen, G. L., Garcia, J., Apfel, N., & Master, A. (2006). Reducing the racial achievement gap: A social-psychological intervention. Science, 313(5791), 1307–1310.
- Fliessbach, K., Weber, B., Trautner, P., Dohmen, T., Sunde, U., Elger, C. E., & Falk, A. (2007). Social comparison affects reward-related brain activity in the human ventral striatum. Science, 318(5854), 1305–1308.
- Takahashi, H., Kato, M., Matsuura, M., Mobbs, D., Suhara, T., & Okubo, Y. (2009). When your gain is my pain and your pain is my gain: Neural correlates of envy and Schadenfreude. Science, 323(5916), 937–939.
- Card, D., Mas, A., Moretti, E., & Saez, E. (2012). Inequality at work: The effect of peer salaries on job satisfaction. American Economic Review, 102(6), 2981–3003.
- Verduyn, P., Ybarra, O., Résibois, M., Jonides, J., & Kross, E. (2017). Do social network sites enhance or undermine subjective well-being? A critical review. Social Issues and Policy Review, 11(1), 274–302.
- Hunt, M. G., Marx, R., Lipson, C., & Young, J. (2018). No more FOMO: Limiting social media decreases loneliness and depression. Journal of Social and Clinical Psychology, 37(10), 751–768.
- Gerber, J. P., Wheeler, L., & Suls, J. (2018). A social comparison theory meta-analysis 60+ years on. Psychological Bulletin, 144(2), 177–197.
- Orben, A., & Przybylski, A. K. (2019). The association between adolescent well-being and digital technology use. Nature Human Behaviour, 3(2), 173–182.
Related Reading
- The Octalysis Framework: Complete Gamification Guide — the full eight-Core-Drive structure that turns Festinger’s comparison engine into design moves.
- Social Identity Theory — the group-level extension of Social Comparison Theory.
- Stereotype Threat — the Core Drive 8 (Loss & Avoidance) shadow on top of Core Drive 5 (Social Influence & Relatedness) that comparison surfaces routinely activate when avatars are demographic-coded.
- Social Loafing — the within-group failure mode that public comparison surfaces are designed to neutralize.
- The Hawthorne Effect — the broader observation-changes-behavior umbrella that all comparison surfaces sit underneath.
- The Dunning-Kruger Effect — the metacognition story for why low-skill comparers misjudge their position in the cohort.
- Self-Determination Theory — the autonomy/competence/relatedness goal-state that comparison surfaces either support or sabotage.
- The Behavioral Framework Library — the full hub.




