Short answer: Cognitive load theory says working memory can only hold a few new items at once while long-term memory is unlimited, so instruction works when it respects that limit while learners build knowledge structures called schemas. John Sweller published the foundational paper in 1988.
The theory names three types of load. Intrinsic load is the built-in difficulty of the material, extraneous load is wasted effort created by poor presentation, and germane load is the productive effort that builds schemas.
The design rule that follows is to eliminate extraneous load, manage intrinsic load by sequencing, and protect room for germane load. Support that helps novices, like worked examples, can start to harm learners as their expertise grows, which researchers call the expertise reversal effect.
Two designers ship an onboarding flow for the same product in the same week. One funnel is gorgeous: big friendly buttons, one field per screen, a progress ring, twelve taps to “done.” Completion rate hits 94 percent and everyone celebrates. The other funnel looks busier, asks the user to actually do the core action once, with a worked example beside it, and finishes at 71 percent. Six weeks later the team looks at who is still active and who can actually use the product. The “worse” funnel wins by a mile. The pretty one created a stream of users who tapped Next twelve times and learned nothing.
That gap is the whole subject of cognitive load theory. Built by John Sweller in 1988 out of a frustrating finding about math problems, it is the most useful working model we have of a hard fact: human working memory is tiny, long-term memory is effectively infinite, and almost everything that goes wrong in teaching, onboarding, tutorials, and interface design is a failure to respect the gap between the two. Sweller did something rare for a learning theory. He turned the vague worry about “overload” into specific, testable instructional effects, and most of them have held up for nearly forty years.
I have spent over twenty years studying what makes people invest effort, and I built the Octalysis Framework to map the drives behind that effort. So this is not another “chunk your content and reduce clutter” summary. I am going to show you how working memory actually fails, the three kinds of load Sweller separated, the one of those three that quietly broke and had to be redefined, and then the layer the theory never built: cognitive load tells you how big the mental bill is, but it is silent on whether the learner will ever choose to pay it. That second question is the one that decides whether anybody learns anything, and it is the one Octalysis was built to answer.
Speed Run Notes
- Working memory holds only about four chunks at once and dumps them in seconds. Long-term memory is unlimited. Learning is the act of moving knowledge from the tiny store to the huge one by building schemas, and instruction fails when it overloads the tiny store.
- There are three kinds of load. Intrinsic load is the built-in difficulty of the material. Extraneous load is wasted effort created by bad presentation. Germane load is the productive effort that actually builds schemas. The three share one fixed budget.
- The design job is simple to state: cut extraneous load to near zero, manage intrinsic load by sequencing, and protect the freed budget for germane work. Most teams do only the first step and stop.
- The theory’s best finding, the expertise reversal effect, also undercuts it: the worked example that helps a novice actively burdens an expert. There is no good design in the abstract, only good design for one learner’s current schema.
- Germane load is the one load type nobody can impose. It is effort the learner has to choose to spend. Cognitive load theory makes room for that effort and then assumes it happens. It does not.
- The Octalysis move: stop trying to make learning effortless and start making the effort worth spending. Each load type maps to a Core Drive, and the S-Tier choice is to free the budget, then motivate the learner to fill it with schema-building, not chrome.
In This Article
- What Is Cognitive Load Theory?
- The Working-Memory Bottleneck and the Long-Term Escape Hatch
- The Three Types of Cognitive Load
- The Effects That Made It a Science
- What Sweller Got Right
- Where Cognitive Load Theory Falls Apart
- What’s Really Happening Inside the Brain
- Cognitive Load Theory vs Other Frameworks
- Cognitive Load Theory in the Real World
- The Elephant in the Room
- How to Apply It with the Octalysis Framework
Author Credibility: Yu-kai Chou

Yu-kai Chou created the Octalysis Framework after studying gamification since 2003 — years before the term entered mainstream vocabulary. As a Human-Systems Architect & Behavioral Designer, his framework has been applied by LEGO, Microsoft, Porsche, Coca-Cola, Salesforce, and MrBeast, impacting over 1.5 Billion Users.
Chou has taught the Octalysis methodology at Harvard, Stanford, Yale, Tesla, Google, BCG, and IDEO.
His work has been cited by Harvard, Stanford, MIT, Forbes, Wall Street Journal, Wired, US Department of Energy, NIST, NSF, NCBI, US Department of Education, ClinicalTrials.gov, and Google Scholar — with 3,700+ more academic publications. Explore his books here.
What Is Cognitive Load Theory?
Cognitive load theory is a theory of how the architecture of human memory shapes what we can learn and how we should teach. Its central claim is blunt: working memory, the mental space where conscious thinking happens, can hold only a handful of new items at once and loses them within seconds unless they are rehearsed. Long-term memory, by contrast, has no known limit. Learning is the process of building organized knowledge structures, called schemas, in long-term memory, and then automating them so they run without conscious effort. Instruction succeeds or fails based on one thing: whether it respects the bottleneck of working memory while that schema-building is going on.
John Sweller arrived at this from an odd direction. In the early 1980s he was studying how people solve math and physics problems, and he kept finding something that should not have happened. Students who solved lots of problems got good at finding answers but did not get better at understanding the underlying structure. The act of hunting for a solution, it turned out, consumed so much working memory that almost none was left over to notice the pattern that would have made the next problem easy. The problem-solving was crowding out the learning. That single observation, published as “Cognitive load during problem solving” in 1988, grew into a framework that now governs how serious instructional designers think about everything from algebra worksheets to surgical simulators.
The reason it matters far beyond the classroom is that every interface is a teacher whether it intends to be or not. A first-time user opening your product is a novice trying to build a schema of how the thing works, under exactly the working-memory limits Sweller described. The checkout flow, the settings panel, the new-feature tooltip, the API documentation: each one is an instructional event, and each one either respects the bottleneck or jams it. Cognitive load theory is the closest thing we have to a physics of that bottleneck.
The Working-Memory Bottleneck and the Long-Term Escape Hatch
To understand the theory you have to hold two facts side by side, because the whole framework lives in the tension between them.
The first fact is that working memory is humiliatingly small. George Miller’s famous 1956 paper put the limit at seven items, plus or minus two. Later work by Nelson Cowan tightened that to roughly four chunks when rehearsal is blocked. Either way, the number is tiny, and it gets worse: those few items decay in about twenty seconds without active maintenance, and the more the items have to interact with each other, the fewer you can juggle at all. Hold a phone number in mind while someone asks you a question and watch the number evaporate. That is not a personal failing. That is the hardware.
The second fact is the escape hatch, and it is the part most summaries skip. Long-term memory is not just a storage closet for facts you have finished learning. It actively reaches back into working memory and changes what counts as a single item. A chess master glances at a mid-game board for five seconds and reproduces all twenty-five pieces perfectly, while a novice manages six or seven. The master does not have a bigger working memory. The pioneering studies by de Groot and later Chase and Simon showed the master sees a handful of familiar patterns where the novice sees twenty-five separate pieces. Years of schema-building turned what was twenty-five items into four. The bottleneck did not widen. The things passing through it got bigger.
This is the move that makes cognitive load theory more than a warning about clutter. The limit on working memory is real and permanent, but it is a limit on novel, unorganized information. Anything you have built a schema for slips through almost for free, because the schema is treated as one chunk no matter how much detail it actually contains. So the goal of instruction is not merely to avoid overloading working memory in the moment. It is to use those scarce working-memory slots to build the schemas that will make future processing effortless. You spend the bottleneck now to widen it forever. Miss that and you will design experiences that feel light and teach nothing, because feeling light and building schemas are not the same thing.
The Three Types of Cognitive Load
Sweller’s most influential contribution was to split the load on working memory into three sources, because the three call for opposite responses. Lumping them together as “cognitive load” is the mistake that leads designers to strip out the wrong things. Two of these you want to cut, one you want to manage, and one you actually want to increase. Confuse them and you will sand the learning right out of your experience while congratulating yourself for simplifying it.

Intrinsic Load: The Difficulty Baked Into the Material
Intrinsic load is the inherent complexity of what is being learned, and it is set by something Sweller called element interactivity: how many pieces of information have to be held in mind and related to each other at the same time to make sense of the thing. Learning the vocabulary word for “cat” in a new language has near-zero element interactivity. You can learn each word in isolation. Learning to balance a chemical equation has high element interactivity, because the coefficients, the atoms on each side, and the conservation rule all have to move together. Change one number and three others shift.
The crucial property of intrinsic load is that you cannot reduce it without reducing what is actually being learned. The difficulty is in the material, not in your slides. What you can do is manage it. Sequencing simple before complex, segmenting a process into stages, and pre-training the names and parts before showing the whole all lower the number of elements a learner has to juggle at any one moment. A good tutorial does not make a hard system easy. It feeds the hard system to you one interacting piece at a time, so that the high total interactivity is never all present in working memory at once. This is why the best teachers of genuinely difficult skills are not the ones who oversimplify. They are the ones who order the difficulty.
Extraneous Load: The Effort You Waste on Bad Design
Extraneous load is the processing imposed not by the material but by the way it is presented. It is pure waste, and it is the load every designer is responsible for. When a diagram sits on page 2 and its explanation sits on page 5, the learner burns working memory holding the diagram in mind while hunting for the words, and that effort builds nothing. That is the split-attention effect. When a narrator reads aloud the exact text already printed on the screen, the learner’s mind tries to reconcile two identical streams and wastes capacity doing it. That is the redundancy effect, and it surprises people: more information made learning worse.
Extraneous load is where almost all the easy wins live, because unlike intrinsic load you can take it to near zero without losing any content. Integrate the label into the diagram. Cut the redundant narration. Remove the decorative carousel that animates while the user is trying to read. Every unit of extraneous load you remove is a unit of working-memory budget handed back to the learner to spend on something that matters. The trap is believing that removing extraneous load is the whole job. It is only the half that clears the table. It does not put any food on it.
Germane Load: The Effort That Actually Builds Schemas
Germane load is the working-memory effort devoted to the real work of learning: noticing patterns, connecting new information to existing schemas, and constructing the mental model that will make all of this automatic later. It is the only load that produces durable learning. Reappraising a worked example to ask “why does this step come before that one,” explaining a concept back to yourself, retrieving an answer from memory instead of rereading it: all of these impose load, and all of that load is the good kind.
Here is the property that almost every treatment of cognitive load theory glides past, and that the rest of this article is built on. Intrinsic load is imposed by the material. Extraneous load is imposed by the design. Germane load is imposed by no one. It is effort the learner has to choose to invest. You can clear the table of extraneous load and serve the material in perfect order, and a disengaged learner will still spend the freed capacity on daydreaming, on the notification that just buzzed, on getting to the end of the module as fast as possible. The theory makes room for germane processing and then quietly assumes it will happen. That assumption is the seam in the whole framework, and it is exactly where motivation, the thing cognitive load theory does not model, decides everything.
The three loads share one fixed budget, the size of working memory. They add up: intrinsic plus extraneous plus germane cannot exceed capacity. If extraneous load eats most of the budget, there is nothing left for germane work, and no learning happens no matter how motivated the person is. That is the cognitive constraint, and it is real. But once you have cleared the extraneous load and made room, the remaining question is no longer cognitive. It is whether the learner spends the freed budget on building a schema or on getting away. Cognitive load theory owns the first question completely and is silent on the second.
The Effects That Made It a Science
A lot of learning theories are just vocabulary: a new set of words for things teachers already sensed. What separates cognitive load theory is that it generated specific, counterintuitive, testable predictions, and then survived the tests. These are called cognitive load effects, and they are the reason the theory earns respect from people who are otherwise skeptical of educational psychology.
The Worked-Example Effect
For novices, studying a fully worked-out example produces better learning than solving the equivalent problem yourself, even though solving feels more active and more rigorous. Sweller and Cooper showed this with algebra in 1985: students who studied worked examples outperformed students who solved problems, and they did it in less time. The reason follows directly from the architecture. A novice solving a problem spends nearly all of working memory on the search for an answer, leaving nothing for the schema-building that actually transfers. A worked example removes the search and lets the learner spend the whole budget on understanding the structure. The hard-won lesson for product and curriculum design: making people struggle through a task before they have a schema does not build the schema. It just exhausts the budget.
The Split-Attention and Modality Effects
When information from two sources must be integrated to be understood, separating them in space or time forces the learner to hold one while searching for the other, and that holding is wasted load. Put the label on the part of the diagram it describes, not in a legend below. The related modality effect goes further: because working memory has partly separate channels for visual and auditory information, presenting a diagram with spoken narration can carry more total information than the same diagram with printed text, since the printed text competes with the diagram for the single visual channel while speech uses the idle auditory one. This is the empirical backbone under every “narrate the animation, do not caption it” multimedia guideline.
The Redundancy Effect
Adding information that repeats what the learner already has hurts learning rather than helping it. Narration identical to on-screen text, a diagram that is fully self-explanatory accompanied by a paragraph re-explaining it, “helpful” captions on a clear visual: each forces the mind to process a second copy and reconcile it with the first. Redundancy is the effect that most offends common sense, because we assume more explanation is generous. In working-memory terms it is a tax. The kindest thing you can do with redundant material is delete it.
The Expertise Reversal Effect
This is the deepest and most disruptive of the effects, and it deserves more attention than it usually gets, because it eventually turns around and bites the theory itself. The worked examples, integrated diagrams, and step-by-step scaffolds that help novices stop helping as expertise grows, and past a certain point they actively hurt. Kalyuga and colleagues documented this clearly in 2003. Once a learner has built the relevant schema, the guidance that used to fill the gap becomes redundant information the learner now has to process and reconcile against what they already know. The scaffold that was a lifeline becomes a weight. The novice needs the worked example. The expert needs to be handed the problem and left alone.
Sit with what this implies. There is no such thing as low-cognitive-load design in the abstract. The same screen, the same lesson, the same tutorial can be beautifully optimized for one learner and actively damaging for another, and the only variable that flipped is what the learner already knows. Good design is not a property of the artifact. It is a relationship between the artifact and a specific mind at a specific stage. That single finding, which is cognitive load theory’s own crown jewel, quietly dissolves the dream of a universal set of design rules, and we will come back to it when we talk about where the theory falls apart.
What Sweller Got Right
Before the critiques, credit where an enormous amount is due, because cognitive load theory got more right than almost any rival in instructional psychology.
It took memory architecture seriously as the binding constraint on learning, instead of treating the mind as an infinitely flexible sponge. For decades, education swung between fashions, discovery learning, learning styles, technology for its own sake, most of which ignored the fixed limits of the machine doing the learning. Sweller started from the hardware and reasoned forward, and that is why his predictions held while the fashions faded.
It produced falsifiable, replicated effects. The worked-example effect, split-attention, redundancy, and modality effects have been reproduced across subjects, age groups, and cultures. In a field justly criticized for results that evaporate on replication, cognitive load theory’s core effects are among the sturdier findings we have. That is not nothing. That is the difference between science and branding.
And it correctly diagnosed the most expensive failure in instruction: the confusion between performance and learning. Sweller’s original insight, that solving problems can make you better at solving the next identical problem while teaching you nothing transferable, is the same disease that shows up in the frictionless onboarding flow with the 94 percent completion rate. Activity is not learning. Completion is not competence. Smoothness is not understanding. Cognitive load theory was the first framework to give a mechanical reason why, and forty years later that reason still holds.
Where Cognitive Load Theory Falls Apart
A theory this durable deserves serious criticism rather than the polite kind. Cognitive load theory has three genuine weak points, and the deepest one is not a measurement quibble. It is a hole in the model exactly where human motivation should be.
Germane Load Was So Slippery It Had to Be Redefined
For nearly two decades, germane load was treated as a third independent source alongside intrinsic and extraneous. The problem is that nobody could measure it independently of the thing it was supposed to cause. Germane load was “the load that produces learning,” and it was inferred from… learning. That is a circle. If a lesson worked, germane load must have been high. If germane load was high, the lesson worked. The construct could explain any result after the fact and predict none in advance, which is the signature of an idea that is not pulling its weight.
To his credit, Sweller saw this and fixed it. In a 2010 paper he reframed germane load so that it is no longer a separate source of load at all. Instead, germane load is the working-memory resources redirected to deal with intrinsic load, the effort that goes toward the productive interactivity rather than the wasteful kind. It is a cleaner formulation. But it is also a quiet admission that one of the three pillars the framework was taught with for twenty years did not survive contact with measurement, and a lot of textbooks and training decks still teach the old three-source version as if nothing happened. A theory whose third construct had to be rebuilt mid-flight should make you hold its precision claims loosely.
It Measures a Hidden Bottleneck by Asking People How Hard It Felt
Cognitive load is supposed to be an objective property of how working memory is being taxed. In practice, the dominant way researchers measure it is a single self-report question, the Paas scale: “How much mental effort did you invest in that task,” answered on a nine-point scale. A theory about the precise, hidden limits of mental hardware is validated, most of the time, by asking people to rate their own effort on a slider. Physiological measures like pupil dilation and EEG exist, but they are noisy, expensive, and hard to attribute to one load type rather than another. The result is that the theory’s central quantity is far softer than its mechanical language implies. When a paper says intrinsic load was high and extraneous was low, that distinction often rests on study design and a one-item rating, not on a clean instrument reading.
It Treats Difficulty as the Enemy, and Difficulty Is Sometimes the Point
This is the critique that matters most for anyone designing real experiences. Cognitive load theory’s prescription, taken at face value, is to minimize load so the learner can build schemas efficiently. But a separate and equally well-supported line of research says that some difficulty is exactly what produces durable learning. Robert Bjork’s desirable difficulties, the testing effect, and Manu Kapur’s work on productive failure all show that conditions which feel harder and slower at the time, retrieving instead of rereading, struggling with a problem before being shown the solution, spacing practice out, lead to stronger long-term retention and transfer than the smooth, easy version.
The reconciliation matters, so do not let anyone tell you the two camps simply contradict. The difficulties Bjork wants you to add are germane: they force the schema-building work. The load Sweller wants you to cut is extraneous: it forces wasted work. A theory that only says “reduce load” without distinguishing which load gives cover to the worst instinct in modern design, which is to strip out every bit of effort until nothing is left to learn. Cognitive load theory, read carelessly, becomes the scientific justification for the frictionless tutorial that teaches nothing. And it reads carelessly to almost everyone, because the headline is “reduce cognitive load,” and that is the half people remember.
Underneath all three critiques is one absence. Cognitive load theory models the size of the budget and how it gets spent, but it has nothing to say about why a learner would choose to spend it on building a schema rather than on escaping the lesson. It assumes the motivation to learn is already there and the only question is whether the design lets that motivation through. Anyone who has watched a student tap through a perfectly designed module to get to the end, or a user skip a flawless onboarding flow, knows that assumption is the whole ballgame and the theory does not address it.
What’s Really Happening Inside the Brain
The model holds up against what neuroscience has learned about memory, which is part of why it has aged so well. Working memory is not a single place but a set of cooperating systems, mapped most influentially by Alan Baddeley and Graham Hitch in 1974. A central executive allocates attention; a phonological loop holds verbal and acoustic information for a few seconds; a visuospatial sketchpad does the same for images and spatial relationships; and an episodic buffer, added later, binds these streams into coherent moments. This architecture is largely seated in the prefrontal cortex, the most metabolically expensive real estate in the brain, which is part of why sustained mental effort feels tiring and why the capacity is so tightly rationed.
The modality effect makes physical sense in this picture: the phonological loop and the visuospatial sketchpad are partly independent resources, so spreading information across both raises the total a learner can hold without overflow. And schema automation has a real neural signature. As a skill is practiced, the heavy, effortful processing that initially recruits the prefrontal cortex shifts toward more consolidated, automatic circuits, and the task stops feeling like juggling and starts feeling like recognition. The chess master’s compressed perception is this consolidation made visible. Expertise does not expand the prefrontal bottleneck. It builds long-term structures that let more of the work bypass the bottleneck entirely.
This is the deeper reason the long-term-memory escape hatch is the most important part of the theory. The brain’s permanent fix for a tiny working memory is not a bigger working memory. It is a vast, well-organized long-term memory that does most of the heavy lifting before working memory is ever involved. Every schema you build is a piece of cognition you no longer have to pay full price for. Design that builds schemas is design that permanently expands what the learner can do, while design that merely keeps load low in the moment leaves the learner exactly as dependent on you as they were before.
Cognitive Load Theory vs Other Frameworks
Cognitive load theory is not a closed island. It sits inside a web of memory, learning, and motivation models, and seeing the joins is how you learn to use it without overreaching.
vs Miller’s Magic Number and Cowan’s Four
Miller’s seven plus or minus two and Cowan’s revised limit of about four chunks describe the size of the bottleneck. Cognitive load theory takes that capacity limit as its starting premise and adds the move that makes it useful for teaching: the escape through schema-building in long-term memory. The capacity researchers told us the room is small. Sweller told us how to get furniture through a small door by assembling it on the other side first.
vs Mayer’s Cognitive Theory of Multimedia Learning
Richard Mayer’s multimedia learning theory is the closest sibling, and the two are often used together. Mayer took the cognitive-load architecture and operationalized it into a tested set of multimedia principles, the coherence principle, the signaling principle, the redundancy principle, the modality principle, that tell you exactly how to lay out words and pictures. Where Sweller gives you the physics, Mayer gives you the engineering handbook for screens. If you design e-learning, you almost certainly want both, with Mayer’s principles as the concrete checklist and cognitive load theory as the reason they work.
vs Bloom’s Taxonomy
Bloom’s Taxonomy, the subject of a companion pillar in this library, tells you which level of cognitive objective to aim at, from Remember through Create. Cognitive load theory tells you how much you can load into working memory on the way to any of those levels. They compose cleanly. A Create-level objective has high element interactivity and therefore high intrinsic load, which means it demands more aggressive sequencing and more prior schema-building before a learner can attempt it. Bloom names the destination. Cognitive load theory governs the size of the steps you can take toward it without the learner falling off.
vs Flow Theory
Mihaly Csikszentmihalyi’s flow is the affective mirror of cognitive load theory’s cognitive account. Flow lives in the balance between challenge and skill: too much challenge for your skill produces anxiety, too little produces boredom, and the narrow channel between them produces total absorption. Translate that into load language and it is nearly the same dial. Challenge above capacity is overload. Challenge far below capacity is the underload of a task so easy that no germane work is happening. Flow describes how that balance feels from the inside; cognitive load theory describes what is happening to the working-memory budget that produces the feeling. Designing for flow and designing within the load budget are two readings of one instrument.
vs Desirable Difficulties
The sharpest tension, and the most productive once resolved. Bjork’s desirable difficulties say that making learning feel harder, through retrieval, spacing, and interleaving, builds more durable knowledge, which sounds like a direct contradiction of “reduce load.” It is not, once you keep the three load types separate. Bjork is adding germane load: difficulty that forces schema-building. Sweller is cutting extraneous load: difficulty that forces wasted work. The synthesis is the entire practical thesis of this article. Cut the difficulty that builds nothing, and deliberately keep, even add, the difficulty that builds schemas. The art is telling the two apart, and that is a judgment call the load numbers alone will never make for you.
Cognitive Load Theory in the Real World
The theory was born in math classrooms, but its reach is far wider, because the working-memory bottleneck does not care whether the learner is a student, a new user, a trainee surgeon, or a player. Anywhere a human meets unfamiliar complexity, these rules apply.
Onboarding and UX
A first-time user is a novice building a schema of your product under a four-chunk limit, which is why the single most common onboarding mistake is showing everything at once. The dashboard with eleven panels, the settings page with forty toggles, the empty state with no example: all of them present high element interactivity to a mind with no schema to compress it. The fix is pure cognitive load theory. Pre-train the parts before the whole. Show one worked example of the core action instead of a feature tour. Strip the extraneous chrome competing for the visual channel. But notice the limit of the fix. You can perfectly sequence an onboarding flow and a user with no reason to care will still bounce, because reducing load gives them the capacity to learn and supplies none of the motivation to use it.
E-Learning and Instructional Design
This is the home turf, and it is where Mayer’s multimedia principles turn the theory into a checklist: narrate animations rather than captioning them, put labels inside diagrams, cut the decorative stock photos and background music, and start novices on worked examples before live problems. Done well, the effect on real learning is large and repeatedly demonstrated. The failure mode is the corporate compliance module that mistakes “low load” for “click Next quickly,” reducing every screen until there is nothing left to build a schema from. Smooth, brief, and forgotten by lunch.
Game Tutorials
The best practitioners of intrinsic-load management on Earth are not professors. They are game designers, because a game that overloads a new player in the first ten minutes is a game that gets refunded. Watch how a well-made game teaches a system with brutal underlying complexity: it introduces exactly one mechanic per area, lets you practice it in a safe space with no penalty, then combines it with the next one only once the first has become automatic. That is pre-training, segmenting, and simple-to-complex sequencing executed at a level most e-learning never reaches. And games solve the half cognitive load theory cannot, because the player is choosing to spend germane effort the whole time. The challenge is voluntary, the progress is visible, and the mastery is the reward. Games are the proof that managing load and motivating effort are two different jobs, and that you need both.
Healthcare and High-Stakes Training
Patient instructions, medication schedules, and discharge paperwork are instructional events delivered to people whose working memory is already taxed by stress, illness, and fear, which collapses the available budget precisely when accuracy matters most. Cognitive load theory explains why the dense pamphlet fails and the single illustrated card with one instruction per step succeeds. In surgical and clinical training, worked-example sequences, segmented simulation, and the careful fading of guidance as competence grows are how high-interactivity skills get built without overwhelming the trainee. Here the expertise reversal effect is not academic. The scaffolding that protects the first-year resident must be removed for the senior one, or it becomes interference at exactly the moment hesitation is dangerous.
The Elephant in the Room
Here is what the instructional-design and UX industries will not put on the conference slide: they love the half of cognitive load theory that says “reduce load” and ignore the half that says learning requires effort the learner has to choose to spend. One half sells beautifully. “Reduce friction, simplify, make it effortless” is a pitch every executive understands and every design team can execute with a redesign. The other half, “and then you have to motivate the learner to invest the freed capacity in hard, generative work,” is the part nobody can solve with a cleaner interface, so it quietly falls off the roadmap.
The result is an entire generation of products and courses optimized to minimize all load, including the germane load that was the only kind doing any teaching. Beautiful onboarding that produces users who cannot use the product. Polished microlearning that produces employees who pass the quiz and retain nothing. Frictionless tutorials measured by completion rate, a metric that rewards exactly the behavior, tapping through fast, that guarantees no schema gets built. The industry took “reduce cognitive load” and heard “remove all effort,” because removing effort is the part you can ship.
The uncomfortable truth underneath is that you cannot cut your way to competence. A schema is built by effortful, motivated processing, and no amount of interface polish manufactures that effort. The reason teams stop at “minimize extraneous load” is not laziness. It is that the second half of the job lives in a domain interface design does not own: motivation. Cognitive load theory clears the table, which is real and necessary work. But a cleared table is not a meal, and the field has spent twenty years perfecting the clearing while pretending the meal would cook itself. It will not. Which is exactly where the Core Drives come in.
How to Apply Cognitive Load Theory with the Octalysis Framework
The Octalysis Framework identifies the eight Core Drives that make humans do anything, and it is the engine cognitive load theory is missing. Cognitive load theory governs the size of the working-memory bill and how to keep it from overflowing. Octalysis governs whether the learner is willing to pay that bill and keeps paying it across the long climb from novice to expert. Put them together and the three load types stop being a static budget and become a motivational design map.
Extraneous Load Is Friction, and Friction Is the Anti-Core-Drives
Every unit of extraneous load is a unit of friction, the exact quantity the conversion world subtracts from motivation in its equations. Split attention, redundant narration, confusing navigation, decorative noise: each one taxes working memory while supplying no drive to keep going. Cutting extraneous load is the cognitive form of cutting friction, and on this point cognitive load theory and Octalysis agree completely. But there is a subtlety that separates an S-Tier designer from a minimalist with a delete key. Not everything “extra” is extraneous. A well-built element of Core Drive 7 (Unpredictability & Curiosity) or Core Drive 3 (Empowerment of Creativity & Feedback) adds processing and adds value, because the processing points at the thing to be learned. The danger is the seductive-details effect, where fun-but-irrelevant facts get bolted on to make a lesson “engaging.” That is Black Hat Core Drive 7 spent on extraneous load: curiosity bait that steals the budget and aims it away from the schema. The rule is precise. Use curiosity that pulls the learner deeper into the material, never curiosity that pulls them sideways out of it.
Intrinsic-Load Management Is the Onboarding Ramp, and That Ramp Is Core Drive 2
Managing intrinsic load means sequencing simple before complex, segmenting a process, and pre-training the parts before the whole. That is, almost word for word, the description of a well-designed onboarding and scaffolding journey, the early Experience Phases in Octalysis design. And the drive that powers a learner up that ramp is Core Drive 2 (Development & Accomplishment): the sense of getting visibly better, one conquered step at a time. When a game introduces one mechanic per level, it is managing intrinsic load and feeding Core Drive 2 in the same motion, because each mastered piece is both a lowered element-interactivity load and a felt win. Sequence the difficulty so that every stage is winnable with the schema the learner already has, and you have managed intrinsic load and built an accomplishment ladder at once. The two designs are the same design.
Germane Load Is the Motivated Core, and It Lives Where Octalysis Lives
This is the crux of the entire crosswalk. Germane load is the one load type that cannot be imposed by the material or the designer. It is effort the learner chooses to spend, which means it is the load that runs entirely on motivation. Cognitive load theory’s instruction is to “make room for germane processing.” It has no mechanism to make the learner actually fill that room, because the mechanism is not cognitive. It is the Core Drives. A learner pours effort into building a schema when there is a reason: Core Drive 1 (Epic Meaning & Calling) tells them this matters beyond the task; Core Drive 2 (Development & Accomplishment) tells them they are getting better; Core Drive 3 (Empowerment of Creativity & Feedback) lets them make the material their own through generation and experiment; Core Drive 4 (Ownership & Possession) makes the growing skill feel like a thing they possess and want to grow. Germane load is precisely the seam where cognitive load theory goes silent and Octalysis takes over. The theory says clear the room. Octalysis says, and here is how you get them to walk into it and work.
The Inversion: Do Not Minimize Load, Spend the Budget Right
The default instinct, in instructional design and in frictionless UX alike, is to minimize total load and call the result good. That instinct is half right and dangerously incomplete, because minimizing load with no further thought minimizes germane load too, and germane load was the only kind that taught anyone anything. The S-Tier move is not a low bill. It is a budget spent entirely on schema-building. Ruthlessly cut extraneous load to zero. Manage intrinsic load with sequencing so the learner is never overwhelmed. Then deliberately fill the freed capacity with motivated, effortful, generative work, retrieval, self-explanation, creation, the desirable difficulties, and use the Core Drives to make the learner want to do that work. Easy is not the goal. A learner whose whole working-memory budget is aimed at building a schema, and who is motivated to keep it aimed there, is the goal. Reappraisal of the design question from “how do I make this effortless” to “how do I make the effort worth spending” is the entire shift.
The Warning Light
There is a specific dashboard pattern that tells you a design has fallen into the trap, and it is worth wiring an alarm to it. When completion rates and “smooth onboarding” satisfaction scores climb while actual competence, retention, and real-task success stay flat or fall, you have built a design that minimized all load including the germane kind. It feels effortless because nothing is being learned. The frictionless funnel with the 94 percent completion rate and the inert users six weeks later is this warning light in its purest form. High completion plus low competence is not a paradox. It is the predictable result of cutting germane load along with the extraneous, and it is the single most common way good intentions about “reducing cognitive load” produce experiences that teach nothing at all.
Practical Steps for Designing Within the Working-Memory Budget
Translating all of this into something you can do on Monday comes down to seven moves, in order. The order matters, because each step frees or protects the budget the next one needs.
- Audit for extraneous load first. Before adding anything, hunt for what to remove: split attention between text and visuals, narration that duplicates on-screen text, decorative animation, seductive details, and anything that imposes processing without building the target schema. This is the cheapest, highest-return work, and it is the only load you can cut to near zero without losing content.
- Size the intrinsic load by element interactivity. Ask how many pieces have to be held and related at once to understand the thing. High interactivity means you must sequence simple before complex, segment the process, and pre-train the names and parts before showing the whole. You are not making it easier. You are feeding the difficulty one interacting piece at a time.
- Use worked examples for novices, then fade them. Start beginners with fully worked examples rather than cold problems, then move through partially completed problems to full independence as their schema forms. Build the off-ramp deliberately, because the expertise reversal effect guarantees the scaffold that helped will eventually burden.
- Protect the freed budget for germane work. Once you have cleared extraneous load, resist the urge to refill the space with more chrome, more features, more decoration. Refill it with the effort that builds schemas: retrieval, self-explanation, generation, and the desirable difficulties that feel harder but teach more.
- Motivate the germane spend with Core Drives. Making room for effort does not produce effort. Give the work a reason with Core Drive 1 (Epic Meaning & Calling), a felt sense of progress with Core Drive 2 (Development & Accomplishment), and room to make it their own with Core Drive 3 (Empowerment of Creativity & Feedback) and Core Drive 4 (Ownership & Possession). This is the half cognitive load theory cannot do for you.
- Re-test as expertise changes. Good design is a relationship between the artifact and a specific mind at a specific stage, not a fixed property. What helps the novice burdens the expert, so build expert paths, skippable scaffolds, and faded guidance, and check that your “helpful” support has not become interference for people who have moved on.
- Watch the warning light. Track competence and retention, not just completion and satisfaction. If completion is climbing while real-task success is flat, you have cut germane load along with the extraneous, and you are shipping an experience that feels effortless because it teaches nothing. Add the productive difficulty back.
Cognitive Load Theory Was the Budget, Octalysis Is the Reason to Spend It
Cognitive load theory is one of the most useful things ever built in instructional psychology, and the reason is that it started from the hardware and refused to look away from a humbling fact: the conscious mind can hold almost nothing at once, and everything we know about teaching has to bend around that. Sweller gave us the budget, the three kinds of spending, and a set of effects sturdy enough to design by. That is a real gift, and most of the failures in onboarding, training, and education trace back to ignoring it.
But a budget is not a purpose. The theory tells you the size of the working-memory bill and how to keep it from overflowing, and then it stops, right at the edge of the only question that finally matters: will the learner choose to spend their freed capacity on building a schema, or on getting away? That choice is not cognitive. It is motivational, and it is governed by the Core Drives. The designer who clears extraneous load and walks away has set a beautiful table and served nothing. The one who clears the load and then makes the effort worth spending, who gives the work meaning, progress, and ownership, is the one whose users actually learn.
So the next time you find yourself reducing cognitive load, finish the sentence. Reduce it, yes, and then ask what you are freeing the budget for, and whether you have given the learner a single real reason to spend it. Clear the table. Then cook the meal.
Frequently Asked Questions
What is cognitive load theory in simple terms?
Cognitive load theory says that working memory, the space where conscious thinking happens, can only hold a few new items at once, while long-term memory is unlimited. Learning means building organized knowledge structures called schemas in long-term memory. Instruction works when it respects the small working-memory limit while those schemas are being built, and fails when it overloads working memory with too much at once.
What are the three types of cognitive load?
Intrinsic load is the built-in difficulty of the material, set by how many elements must be held and related at once. Extraneous load is wasted effort created by poor presentation, like split attention or redundant text. Germane load is the productive effort that actually builds schemas. The design goal is to cut extraneous load, manage intrinsic load by sequencing, and protect room for germane load.
Who created cognitive load theory?
Educational psychologist John Sweller developed cognitive load theory, with his foundational paper published in 1988. It grew out of his earlier research on problem solving, where he found that searching for solutions consumed so much working memory that little was left for learning the underlying patterns. The framework was later expanded with colleagues including Jeroen van Merriënboer and Fred Paas.
What is the difference between intrinsic and extraneous cognitive load?
Intrinsic load comes from the material itself and cannot be reduced without reducing what is learned, though it can be managed through sequencing and segmenting. Extraneous load comes from how the material is presented and is pure waste that can be cut to near zero without losing any content. The first you manage; the second you eliminate.
Why is germane load controversial?
For years germane load was treated as a separate source of load, but it could only be measured by the learning it was supposed to cause, which made it circular. In 2010 Sweller redefined it as the working-memory resources redirected to deal with intrinsic load, rather than an independent third source. Many textbooks still teach the older three-source version.
What is the expertise reversal effect?
The expertise reversal effect is the finding that instructional support which helps novices, like worked examples and detailed guidance, stops helping and eventually harms learners as their expertise grows. Once a learner has built the relevant schema, the extra guidance becomes redundant information they must process and reconcile. It means there is no universally good design, only design suited to a learner’s current level.
How is cognitive load theory used in UX and onboarding?
A first-time user is a novice building a schema under tight working-memory limits, so good onboarding pre-trains the parts before the whole, shows one worked example of the core action instead of a full feature tour, and strips out chrome that competes for attention. The limit is that reducing load gives users the capacity to learn but not the motivation to use it, which is where motivational design comes in.
Does reducing cognitive load always improve learning?
No. Reducing extraneous load helps, but reducing all load can strip out the germane effort that builds durable learning. Research on desirable difficulties and the testing effect shows that some struggle, like retrieving from memory or solving before being shown the answer, produces stronger long-term retention. The goal is to cut wasteful difficulty while keeping, or adding, the productive kind.
How does cognitive load theory relate to gamification?
Cognitive load theory governs how much you can load into working memory; gamification and the Octalysis Framework govern whether the learner is motivated to spend that capacity on building a schema. Germane load is the one load type that cannot be imposed, only chosen, which makes it a motivational problem. The Core Drives supply the reason to invest the effort that cognitive load theory makes room for but cannot create.
References
- Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285.
- Sweller, J., & Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction, 2(1), 59–89.
- Chandler, P., & Sweller, J. (1991). Cognitive load theory and the format of instruction. Cognition and Instruction, 8(4), 293–332.
- Sweller, J. (1994). Cognitive load theory, learning difficulty, and instructional design. Learning and Instruction, 4(4), 295–312.
- Sweller, J., van Merriënboer, J. J. G., & Paas, F. (1998). Cognitive architecture and instructional design. Educational Psychology Review, 10(3), 251–296.
- Sweller, J. (2010). Element interactivity and intrinsic, extraneous, and germane cognitive load. Educational Psychology Review, 22(2), 123–138.
- Sweller, J., van Merriënboer, J. J. G., & Paas, F. (2019). Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31, 261–292.
- Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31.
- Paas, F. (1992). Training strategies for attaining transfer of problem-solving skill in statistics: A cognitive-load approach. Journal of Educational Psychology, 84(4), 429–434.
- Paas, F., Tuovinen, J. E., Tabbers, H., & Van Gerven, P. W. M. (2003). Cognitive load measurement as a means to advance cognitive load theory. Educational Psychologist, 38(1), 63–71.
- Miller, G. A. (1956). The magical number seven, plus or minus two. Psychological Review, 63(2), 81–97.
- Cowan, N. (2001). The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences, 24(1), 87–114.
- Baddeley, A. D., & Hitch, G. (1974). Working memory. Psychology of Learning and Motivation, 8, 47–89.
- Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55–81.
- de Groot, A. D. (1965). Thought and Choice in Chess. Mouton.
- Mayer, R. E. (2009). Multimedia Learning (2nd ed.). Cambridge University Press.
- Mayer, R. E., & Moreno, R. (2003). Nine ways to reduce cognitive load in multimedia learning. Educational Psychologist, 38(1), 43–52.
- Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why minimal guidance during instruction does not work. Educational Psychologist, 41(2), 75–86.
- Kapur, M. (2008). Productive failure. Cognition and Instruction, 26(3), 379–424.
- Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. Shimamura (Eds.), Metacognition: Knowing About Knowing. MIT Press.
- Chou, Y. (2015). Actionable Gamification: Beyond Points, Badges, and Leaderboards. Octalysis Media.
Related Reading
- The Octalysis Framework: The 8 Core Drives of Gamification
- The Behavioral Framework Library
- Bloom’s Taxonomy: S-Tier Behavioral Designer’s Guide
- Emotion Regulation: S-Tier Behavioral Designer’s Guide



