Blog · Behavioral Analysis Contact Me
Baddeley’s Working Memory Model: S-Tier Behavioral Designer’s Guide
Behavioral Analysis

Baddeley’s Working Memory Model: S-Tier Behavioral Designer’s Guide

Try to hold these seven digits in your head while you keep reading: 4, 9, 1, 7, 2, 8, 5. Now imagine I interrupt you, ask you to spell your own surname backward, then tell you to recite the digits. Most of them are gone. You did not forget them in any normal sense, because you only met them four seconds ago. They evaporated because the place you were keeping them is almost absurdly small, and the moment something else needed that space, the digits were overwritten.

That tiny, fragile, constantly-overwritten space is working memory, and almost everything a person does inside your product, your classroom, your form, or your onboarding flow has to pass through it first. In 1974, Alan Baddeley and Graham Hitch looked inside that space and found it was not one thing. It was a small system with separate parts: a manager that allocates attention, one channel for words and sounds, a different channel for pictures and space, and later a fourth part that binds the fragments into something you can actually use. They called it the working memory model, and it is still the most useful map we have of the bottleneck every experience squeezes through.

I have spent more than twenty years studying what moves people to act, and I built the Octalysis Framework to map the eight drives underneath that action. So this is not a textbook tour of a famous memory model. I am going to show you how working memory is actually built, why “just simplify it” is only half the advice, what your brain is physically doing while it holds those digits, and then the part almost no designer works on by intention. Working memory is the scarcest resource in the room. Most people try only to spend less of it. The real move is to win it.

Speed Run Notes

  • Working memory is not one store. Baddeley and Hitch split it into a central executive (the attention manager), a phonological loop (verbal and acoustic), and a visuospatial sketchpad (visual and spatial), with an episodic buffer added later to bind it all together.
  • The verbal and visual channels are separate. You can run a word task and a picture task at the same time with little interference, but two word tasks collide. That single fact rewrites how you should lay out information.
  • Capacity is brutal. The real ceiling is closer to four chunks than the famous “seven,” and a chunk is whatever your long-term knowledge lets you treat as one unit. Expertise is mostly bigger chunks.
  • Anxiety, dread, and clutter spend working memory too. A frightened or distracted user has less capacity left for your interface, no matter how clean it is. You are competing for the space, not just filling it.
  • Cognitive Load Theory tells you how to stay under the ceiling. It is silent on why a user would spend any of their ceiling on you at all. That gap is exactly where motivation, and Octalysis, lives.
  • The S-Tier move: stop only subtracting load. Reduce what the experience costs in working memory, and separately earn the allocation by making your core value the thing the user’s attention manager actually wants to hold.

Author Credibility: Yu-kai Chou

Yu-kai Chou — creator of the Octalysis Framework

Yu-kai Chou created the Octalysis Framework after studying gamification since 2003 — years before the term entered mainstream vocabulary. As a Human-Systems Architect & Behavioral Designer, his framework has been applied by LEGO, Microsoft, Porsche, Coca-Cola, Salesforce, and MrBeast, impacting over 1.5 Billion Users.

Chou has taught the Octalysis methodology at Harvard, Stanford, Yale, Tesla, Google, BCG, and IDEO.

His work has been cited by Harvard, Stanford, MIT, Forbes, Wall Street Journal, Wired, US Department of Energy, NIST, NSF, NCBI, US Department of Education, ClinicalTrials.gov, and Google Scholar — with 3,700+ more academic publications. Explore his books here.

What Is the Working Memory Model?

Working memory is the mental workspace where you hold and manipulate information for the few seconds you are actively using it. It is the difference between storage and thinking. Long-term memory is the warehouse. Working memory is the workbench, and the workbench is tiny.

Before 1974, the dominant picture came from Richard Atkinson and Richard Shiffrin, who described memory as a chain of stores: information enters a sensory register, passes into a single short-term store, and with enough rehearsal moves into long-term memory. That short-term store was treated as one passive box that simply held a few items until they faded or transferred. It was a clean model, and it was wrong in an important way.

Alan Baddeley and Graham Hitch noticed that the single-box idea could not explain what people actually do. If short-term memory were one undifferentiated store, then filling it up with one task should wreck performance on any other task that needs it. Yet people can hold a string of numbers in mind and still read a sentence, reason through a problem, or follow directions, with surprisingly little collapse. A single passive store cannot do two things at once like that. Something with separate working parts can.

So they reframed short-term memory as an active system with specialized components, each with its own limited capacity, coordinated by a controller. They renamed it working memory to stress the point: this is not a holding pen, it is a processor. The name change carried the whole argument. You do not just store in here. You work in here, and the work is constrained by how the parts are built.

For anyone who designs anything a human has to think through, this is the single most consequential fact about the mind. Every instruction, every menu, every checkout, every lesson, every conversation is competing for room on a workbench that holds about four things and clears itself constantly. Get the load wrong and the most brilliant content in the world never reaches the person, because it never made it onto the bench.

The Four Components of Working Memory

The model began with three parts in 1974 and grew to four in 2000, when Baddeley added the episodic buffer to fix a problem the original version could not handle. Here is what each part does, and why each one matters the moment you try to design for a real human.

The Central Executive: The Attention Manager

The central executive is the boss of the system. It does not store much itself. Its job is to direct attention, decide which subsystem handles what, switch between tasks, suppress distractions, and pull relevant knowledge out of long-term memory. Think of it as an air-traffic controller with a very small tower and far too many planes.

This is the part that decides whether the user’s mental processing goes to your interface or to the argument they had this morning. And here is the uncomfortable truth designers skip past: the central executive has limited capacity and it is always already busy. It is not sitting idle waiting to engage with your onboarding. It is fielding worries, hunger, the notification that just buzzed, and a dozen half-finished thoughts. You are not filling an empty manager. You are competing for an overbooked one.

Baddeley was always honest that the central executive was the least understood and most overworked part of his own model. He once described it as the component carrying the heaviest load of “residual ignorance.” We will come back to that, because it is also the model’s biggest weakness.

The Phonological Loop: The Inner Voice

The phonological loop handles verbal and acoustic information: words, numbers, language, anything you can say or hear in your head. It has two pieces. The first is the phonological store, an “inner ear” that holds sound-based traces for roughly one and a half to two seconds before they decay. The second is the articulatory rehearsal process, an “inner voice” that refreshes those traces by silently repeating them, the way you rehearse a phone number while you walk to find a pen.

That two-second limit is not a metaphor. It is measurable, and it produces one of the most reliable findings in psychology. You can hold roughly as many words as you can pronounce in about two seconds. Short words win. Long words lose. This is why a phone number gets chunked into small groups and why a password rule that demands a long unpronounceable string is fighting the brain’s hardware, not just the user’s patience.

The Visuospatial Sketchpad: The Inner Eye

The visuospatial sketchpad handles visual and spatial information: what things look like, where they are, how they move and relate in space. It is the “inner eye” you use to picture your kitchen and count the windows, or to rotate a shape in your mind, or to recall the rough layout of a page you scrolled past. Robert Logie later argued it splits further into a visual cache, holding form and color, and an inner scribe, handling spatial and movement information.

The decisive point is that this channel is separate from the phonological loop. They draw on different resources. That separation is not a footnote. It is the most practically useful fact in the entire model, and we will build a design principle on it shortly.

The Episodic Buffer: The Binder

By 2000, Baddeley faced a problem. The original three components could not explain how information from different channels gets combined, or how working memory talks to long-term memory, or how people hold far more in a meaningful sentence than the two-second loop should allow. So he added a fourth part: the episodic buffer.

The episodic buffer is a limited-capacity store that binds information from the phonological loop, the visuospatial sketchpad, and long-term memory into single, integrated episodes, into chunks. It is how the sound of a word, the sight of an object, and what you already know about it fuse into one unit you can hold and act on. It is also the reason expertise feels like a superpower. A chess master holds an entire board position as a few meaningful chunks, while a novice sees thirty-two separate pieces and drowns. The board did not change. The chunking did.

This binding function is where capacity quietly expands. The raw channels are tiny and fixed. But what counts as one item on the workbench depends on the knowledge you bring, and the episodic buffer is where that knowledge does its work.

The Experiments That Built the Model

The working memory model is not a tidy idea someone reasoned into existence. It is the residue of a long series of experiments where people tried to break short-term memory in specific ways and watched which parts gave out. A few of those studies carry most of the weight.

The dual-task experiments are the foundation. Baddeley and Hitch had people hold a load of digits in mind while simultaneously reasoning, comprehending, or learning. If working memory were one store, a heavy digit load should cripple the second task. It barely dented it. The reasoning slowed a little and stayed accurate. That dissociation, doing two demanding things at once without total collapse, is the empirical fingerprint of a system with separate parts rather than one shared pool.

The word-length effect, shown by Baddeley, Thomson, and Buchanan in 1975, pinned down the phonological loop. People recalled more short words than long words, and the cutoff tracked pronunciation time, not the number of words. The loop holds about two seconds of speech, full stop.

The phonological similarity effect, from Conrad and Hull, showed that letters which sound alike (B, C, T, P, V) are harder to hold in order than letters that look alike, proving the loop codes by sound rather than appearance. And articulatory suppression sealed it: if you force people to mutter “the, the, the” out loud, you occupy the inner voice, the rehearsal process stops, and the word-length effect vanishes. Block the mechanism and the signature disappears exactly as the model predicts.

A later generation of studies added a second kind of measurement. Simple span tasks, like repeating a string of digits, capture mostly the passive holding capacity of the loop. But Meredyth Daneman and Patricia Carpenter built complex span tasks that force holding and processing at the same time, such as reading a series of sentences while keeping the last word of each one in mind. Performance on those tasks predicts reading comprehension, reasoning, and following directions far better than simple span does, because real thinking is never pure storage. It is storage under interference, which is exactly what the central executive is for. The n-back task, where you decide whether the current item matches the one from a few steps back, became the other workhorse, precisely because it loads holding and updating together.

Finally, the neuropsychology. The patient known as KF, studied by Tim Shallice and Elizabeth Warrington, had a severely impaired short-term memory yet a working long-term memory, the opposite of what the Atkinson-Shiffrin chain predicted. If everything had to pass through a single short-term store on the way to long-term storage, a broken short-term store should block new long-term learning. It did not. A model with separate, dissociable parts explains KF easily. A single-store model cannot.

What Baddeley Got Right

The model has lasted more than fifty years, which in psychology is close to immortality. It earned that durability for concrete reasons.

First, it replaced a passive store with an active processor, and that reframe turned out to be correct and generative. Treating the mind’s workspace as something that manipulates rather than merely holds opened decades of productive research and gave clinicians, educators, and designers a working vocabulary.

Second, the separation of verbal and visual channels has held up under enormous testing. The prediction is specific and falsifiable: tasks that share a channel should interfere, tasks that use different channels should coexist. Try to hold a sentence in mind while someone reads you a list and both suffer. Try to hold a sentence while you track a moving dot and you manage both. That pattern repeats across hundreds of studies, and it is the basis of dual-coding theory in education and of every well-built interface that pairs a short label with a clear picture instead of stacking two paragraphs.

Third, the model explained the broken cases. KF and patients like him made sense for the first time. A theory that accounts for the clean lab data and the messy clinical data is doing real work, not just describing the average undergraduate in a quiet room.

Fourth, it gave the world a usable budget. The practical heuristics that flow from it, chunk information, do not split attention across two verbal streams, pair words with images, keep the active set small, are reliable precisely because they trace back to how the components are actually built. Baddeley did not just describe the limit. He explained its shape, and a limit with a shape is a limit you can design around.

There is a fifth reason the construct earned its keep, and it is the one that should make every designer pay attention. Working memory capacity is not a trivia score. It predicts outcomes that matter: reading comprehension, problem solving, the ability to follow multi-step instructions, even general fluid intelligence, the raw capacity to reason about novel problems. Tracy Alloway and others have shown that a child’s working memory in the early years forecasts later academic achievement at least as well as IQ, sometimes better, and often independent of it. When you overload a user’s working memory, you are not just making them mildly uncomfortable. You are temporarily handing a capable person the cognitive profile of someone who cannot keep up. The most expensive thing you can do to a smart user is exceed their workbench, because from the inside it feels like being made stupid by your product.

Where the Working Memory Model Falls Apart

A model this useful invites the temptation to treat it as settled fact. It is not, and the honest practitioner needs to know exactly where it cracks.

The Central Executive Is a Homunculus

The deepest problem is the boss. Saying “the central executive allocates attention and decides what to do” can explain almost any result after the fact, which is another way of saying it explains nothing in advance. It is a little person inside the head doing the smart parts, and a theory that hides its intelligence inside an unexplained manager has just relabeled the mystery. Baddeley knew this and said so. Decades of work have tried to break the executive into specific functions, such as updating, shifting, and inhibition, with real progress, but the original “executive does the clever bits” formulation is closer to a placeholder than an explanation.

The Episodic Buffer Was a Patch

The fourth component arrived because the first three could not account for binding, for the interaction with long-term memory, or for how a meaningful sentence carries far more than the loop should allow. Adding a buffer solved those problems by naming a part that does exactly what was missing. That is a legitimate scientific move, but it has the flavor of patching a hole rather than predicting one. How the buffer actually binds, and whether it is a distinct store or just the central executive wearing a different hat, remains underspecified.

The Boxes May Not Be Real

The model draws clean rectangles, but the brain does not. Competing theories argue the boxes are convenient fictions. Nelson Cowan describes working memory not as separate stores but as the currently activated portion of long-term memory, with a narrow focus of attention holding about four items. Randall Engle frames individual differences in working memory as differences in controlled attention, not buffer size. Anders Ericsson and Walter Kintsch showed that experts use a “long-term working memory” that blows past the supposed limits within their domain. Each of these explains some data better than the boxes do. The model is a map, and like every map it is wrong in the places where the territory is more tangled than the drawing.

One more limit matters enormously for anyone hoping to dodge the problem: you mostly cannot train your way out of it. A wave of brain-training products promised that practicing n-back drills would expand working memory and, with it, intelligence. Susanne Jaeggi’s early studies sparked the excitement, but the careful meta-analyses that followed, led by Monica Melby-Lervåg and Charles Hulme, found the same disappointing pattern again and again. People get better at the training task. The gains almost never transfer to anything else. Working memory capacity behaves more like height than like a muscle. You do not get to assume the user will arrive with a bigger workbench because they practiced. The capacity in front of you is the capacity you have to design for, which is exactly why design, not user self-improvement, carries the load.

What’s Really Happening Inside the Brain

The components are functional fictions, useful labels for jobs the brain performs. So what is the physical brain doing while you hold those seven digits?

The classic answer comes from Patricia Goldman-Rakic, who recorded from the prefrontal cortex of monkeys during delay tasks. When an animal had to hold a location in mind across a few seconds with nothing on the screen, specific neurons in the dorsolateral prefrontal cortex kept firing through the entire delay, then stopped once the response was made. That persistent firing, neurons staying active to keep information alive when the world has gone quiet, looked like the physical act of holding something in mind. Working memory, in this view, is information kept warm by sustained neural activity in a prefrontal-parietal network.

The picture has since grown more contested and more interesting. Some researchers argue that working memory is not always held by continuous firing. Information can be stored “activity-silent,” carried in rapid, temporary changes to the strengths of synapses, and read back out by a quick burst of activity when needed. On this account the memory can go briefly dark and still be there, more like a charge held in the wiring than a light left on. The debate between persistent-activity and activity-silent models is one of the live frontiers of neuroscience.

The functional channels map loosely onto anatomy too. Verbal rehearsal leans on left-hemisphere language regions around Broca’s area and the supramarginal gyrus. Visuospatial holding leans more on right-hemisphere and posterior parietal and occipital regions. Dopamine appears to act as a gate, helping decide what gets into the prefrontal store and what is kept out, which is part of why motivation and reward are not separate from working memory but woven through it. The thing that decides what you hold is chemically tied to the thing that decides what you care about. Hold that idea. It is the bridge to everything that follows.

Working Memory vs Other Models

Baddeley’s model does not stand alone. It sits inside a family of theories that agree on the bottleneck and argue about its shape. Knowing the neighbors keeps you from treating one map as the whole territory.

vs the Atkinson-Shiffrin Multi-Store Model

This is the model Baddeley replaced. Atkinson and Shiffrin pictured a single passive short-term store that fed long-term memory through rehearsal. Baddeley kept the basic insight that short-term and long-term memory differ, then shattered the single store into active, specialized parts. The upgrade matters for design: the old model implies “keep it short,” while the new one implies the far richer “keep it short in each channel, and use the channels you are not already loading.”

vs Cowan’s Embedded-Processes Model

Nelson Cowan rejects the separate-boxes architecture. In his view there are no dedicated buffers; working memory is simply the part of long-term memory that is currently activated, with a spotlight of attention illuminating about four chunks at a time. Where Baddeley draws structures, Cowan draws a state of activation. For practitioners the two converge on the same hard number: the usable focus is around four items, not the folkloric seven. They disagree about the machinery and agree about the ceiling, which is the part you have to design under either way.

vs Miller’s Magical Number and Cognitive Load Theory

George Miller’s 1956 paper gave the world “seven, plus or minus two,” and that number escaped into business culture and never left. The harder, more recent estimate from Cowan is closer to four chunks when you strip out rehearsal and grouping tricks. The practical descendant of all this is John Sweller’s Cognitive Load Theory, which takes the working memory limit as its starting axiom and builds an entire instructional method on top of it. If the working memory model is the anatomy of the bottleneck, Cognitive Load Theory is the applied engineering for working within it. Read them together and you have both the why and the how.

Working Memory in the Real World

The model stops being academic the instant you watch it govern whether a real person succeeds or gives up. Four arenas show it plainly.

Onboarding and the Workplace

Most onboarding fails not because the steps are hard but because they arrive all at once. A new hire or a new user is handed a flood of names, tools, rules, and goals, and the workbench overflows in the first ten minutes. The fix the model prescribes is not “explain better.” It is “stage the load.” Reveal one meaningful chunk, let it settle into long-term memory where it stops consuming working memory, then add the next. Progressive disclosure is not a design fashion. It is the working memory model turned into a sequence.

Education and Multimedia Learning

Richard Mayer built an entire science of multimedia learning on the separation of the two channels. His findings read like direct corollaries of Baddeley: present a diagram with spoken narration and people learn well, because the picture loads the sketchpad while the words load the loop. Present the same diagram with a wall of on-screen text and learning collapses, because now the eyes are forced to process both the picture and the text through the visual channel while the words have to be silently read into the verbal one. Same information, double the load, half the learning. The channel you choose is not cosmetic. It is the budget.

Interfaces, Forms, and Checkout

Every abandoned checkout is partly a working memory story. A form that makes you hold a confirmation code in your head while you scroll, a price that only adds up across three screens, a multi-step flow that never shows you where you are, all of it spends the user’s tiny budget on bookkeeping instead of deciding to buy. The cure is to offload memory onto the screen. Recognition is cheap; recall is expensive. Show the running total, keep the entered data visible, mark the step the user is on. Every item the interface remembers for the user is an item their workbench is free to spend on saying yes.

The one-time passcode is the perfect miniature of all this. You receive a six-digit code in one app and have to carry it into another. Six digits is right at the edge of the loop, so the better systems chunk the code into two groups of three on the screen, which is the same trick that turned ungrokkable phone numbers into memorable ones for a century. The best systems remove the working memory demand entirely by autofilling the code, so it never has to touch the loop at all. Watch how a single design decision moves a task from “hold six volatile items across an app switch” to “hold nothing.” That delta, measured in dropped sign-ins, is the working memory model showing up directly on a revenue chart.

Healthcare and Instructions

The stakes climb fast in medicine. Discharge instructions, medication schedules, and consent information routinely exceed the working memory of a frightened, tired, or older patient, and the failure shows up as missed doses and repeat admissions. Anxiety makes it worse, because fear itself consumes working memory, leaving even less room for the instructions. Good clinical communication treats this as a capacity problem: one instruction at a time, paired with a written or visual aid that lives outside the head, confirmed by teach-back rather than a hopeful “any questions?”

The Elephant in the Room

Here is what the entire user-experience industry has built a comfortable religion around, and the part it leaves off the slide. The religion is load reduction. Simplify. Minimize. “Don’t make me think.” Progressive disclosure, fewer fields, cleaner screens. All of it is correct, all of it traces straight back to the working memory model, and all of it is only half the truth.

The unsold half is this: you can build the most frictionless, lowest-load interface ever shipped and still lose the user’s working memory completely. Because the central executive is not choosing between your clean screen and your cluttered screen. It is choosing between your screen and the doomscroll, the worry, the rival app, the text they are dying to answer. Reducing your load lowers the price of admission. It does nothing to make the user want to buy the ticket. A boring effortless experience and an exciting effortless experience cost the same in working memory. One of them gets none of it anyway.

The industry sells load reduction so hard because subtracting is safe, measurable, and uncontroversial. You can A/B test removing a field. It is much harder to put “we also have to be worth thinking about” on a roadmap, because that is not a subtraction, it is a creative act, and it cannot be reduced to a checklist. So the comfortable half gets all the attention and the decisive half gets hand-waved as “engagement,” someone else’s problem on a different team.

Working memory is a scarce resource, and scarce resources are not merely conserved. They are competed for. The designers who only subtract are playing defense in a game that is won on offense.

How to Apply It with the Octalysis Framework

The working memory model tells you the size and shape of the bottleneck. It tells you the workbench holds about four chunks, that the verbal and visual channels are separate, that anxiety eats capacity, and that the central executive allocates the rest. What it cannot tell you is why a user would spend that scarce capacity on your experience instead of the thousand other things pulling at it. That is a motivation question, and the Octalysis Framework is built to answer it. Below, each part of the working memory system maps to specific Core Drives, so you can see exactly where “reduce the load” ends and “earn the allocation” begins.

The Octalysis Framework with the 8 Core Drives and game techniques — Yu-kai Chou

The Central Executive Is the Attention You Have to Win: Core Drive 1 and Core Drive 2

The central executive decides where the scarce processing goes, and it does not allocate to your task just because your task is present. It allocates to whatever feels most worth the spend. Two Core Drives are how you earn that spend honestly. Core Drive 1 (Epic Meaning & Calling) gives the executive a reason that is bigger than the screen: when a person believes the task matters, that it is part of something they care about, their attention manager defends the allocation against distraction instead of leaking it. Core Drive 2 (Development & Accomplishment) supplies the other reason: visible progress and earned competence make the task feel like it is paying out, so the executive keeps investing. The White Hat use is to give a genuinely valuable task a meaning and a sense of progress strong enough to hold attention. The failure mode here is not a Black Hat trick at all. It is the boring, meaningless task that the executive quietly starves, no matter how clean you made it.

The Two Channels Are the Budget, So Do Not Bankrupt It: Core Drive 6 and Core Drive 8

The phonological loop and the visuospatial sketchpad are tiny, separate, and easy to overload. This is where working memory behaves exactly like the scarcest resource in the room. Core Drive 6 (Scarcity & Impatience) is what working memory IS to your design: a hard limit on supply, where every element you add spends a budget you cannot refill. The discipline is to treat each item on the workbench as expensive, to split information across the two channels instead of stacking it in one, and to offload everything you can onto the screen so the head is free. Core Drive 8 (Loss & Avoidance) is the way the budget gets bankrupted: fear, dread, and threat consume working memory directly, which is why an anxious user processes less of a perfectly clean interface than a calm one does of a messy one. The White Hat use of this knowledge is to remove the anxiety and clutter that drain capacity. The Black Hat use, and you see it constantly, is to overload working memory on purpose: confusion pricing, deliberately convoluted cancellation flows, terms buried in a wall of text built to exceed the budget so the user gives up and clicks accept.

The Episodic Buffer Is Where Meaning Becomes Capacity: Core Drive 3 and Core Drive 4

The episodic buffer binds fragments into chunks using what the user already knows, and the size of a chunk is the closest thing to a capacity cheat code that exists. A novice holds thirty-two chess pieces and drowns; a master holds four meaningful patterns and sees the whole board. Two Core Drives expand the buffer’s reach. Core Drive 3 (Empowerment of Creativity & Feedback) builds richer chunks by letting users actively combine elements and see the result, because knowledge you construct and test yourself binds into denser, more retrievable units than knowledge you were handed. Core Drive 4 (Ownership & Possession) does the same through belonging: what a user builds, names, and owns becomes a single chunk they carry effortlessly, the way you hold the layout of your own home as one thing. And Core Drive 7 (Unpredictability & Curiosity) sits alongside them, because an open curiosity loop is held in this buffer until it closes, which is why a well-placed question keeps a user mentally engaged across a gap that empty content could never bridge. The S-Tier move is to scaffold meaning so that novices chunk like experts, turning your scarce four slots into four rich units instead of four lonely facts.

Practical Steps for Designing Within Working Memory

Theory earns its keep when it changes what you do on Monday. Here is the working memory model as a sequence of concrete moves.

  1. Count the load on every screen. Walk through each step and tally the independent items the user must hold in their head to proceed. If the count climbs past four, you have a redesign, not a copy edit.
  2. Split across channels, do not stack in one. Pair a short verbal label with a clear visual instead of two blocks of text or two streams of speech. You get two budgets working in parallel rather than one budget overdrawn.
  3. Offload memory onto the interface. Show the running total, keep entered data visible, mark the current step, surface the confirmation code where it is needed. Recognition is cheap and recall is expensive, so make the screen remember and let the user decide.
  4. Chunk for the novice, not the expert. Group, label, and scaffold so a beginner can hold your information in expert-sized units. Do not assume the chunks that are obvious to you exist yet in their head.
  5. Protect the executive from anxiety. Remove dread, interruption, and clutter, because fear and distraction spend the very capacity your experience needs. A calmer user has a bigger workbench.
  6. Earn the allocation, do not just lower the cost. Give the core task genuine meaning and visible progress so the attention manager wants to hold it. Load reduction gets you in the door. Motivation is why the user walks through.
  7. Test under realistic load. Evaluate the experience with a second demand running, the way life actually delivers it, not in a silent room with a focused tester. The quiet-lab version always looks easier than it is.

Working Memory Was the Map, Octalysis Is the Move

Baddeley and Hitch gave us something rare and durable: an accurate map of the smallest, most contested space in the human mind. The workbench holds about four things, clears itself in seconds, runs words and pictures on separate tracks, and answers to a manager that is always already overbooked. Every designer should know that map by heart, because nothing you build reaches a person without crossing it first.

But a map of a scarce resource is only half a strategy. The other half is the question the map cannot answer: out of everything competing for those four precious slots, why would the user give any of them to you? Reduce the load until the price of entry is near zero, yes. Then go further than the comfortable half of the field ever goes, and make your core value the thing the user’s attention actually wants to hold. The mind that decides what to keep is the same mind that decides what to care about. Win the second and you have already won the first.

Frequently Asked Questions

What is Baddeley’s working memory model in simple terms?

It is a description of the mental workspace where you hold and manipulate information for a few seconds. Instead of one passive short-term store, Baddeley and Hitch proposed a system with separate parts: a central executive that directs attention, a phonological loop for words and sounds, a visuospatial sketchpad for images and space, and a later-added episodic buffer that binds everything into usable chunks.

What are the four components of working memory?

The central executive (the attention manager that allocates resources and switches between tasks), the phonological loop (the verbal and acoustic channel, an inner voice rehearsing sound-based information), the visuospatial sketchpad (the visual and spatial channel, an inner eye), and the episodic buffer (a binder that integrates information from the other components and from long-term memory into single chunks).

How is working memory different from short-term memory?

Short-term memory, in the older view, was a single passive store that simply held a few items for a short time. Working memory is an active system that holds and manipulates information, with specialized components and a controller. The shift in name marks the shift in idea: it is not a holding pen, it is a processor you think inside of.

How many things can working memory hold?

The famous figure is “seven, plus or minus two,” from George Miller. The more careful modern estimate, from Nelson Cowan, is closer to four chunks once you remove rehearsal and grouping tricks. Crucially, a chunk is whatever your knowledge lets you treat as one unit, so an expert effectively holds far more than a novice by packing more into each slot.

What is the difference between the phonological loop and the visuospatial sketchpad?

The phonological loop handles verbal and sound-based information, words, numbers, and language, by silently rehearsing them. The visuospatial sketchpad handles visual and spatial information, what things look like and where they are. The key practical point is that they are separate, so you can load both at once with much less interference than loading either one twice.

Why was the episodic buffer added to the model?

The original three components could not explain how information from different channels gets combined, how working memory interacts with long-term memory, or how people hold far more in a meaningful sentence than the two-second loop should allow. Baddeley added the episodic buffer in 2000 as a limited-capacity store that binds those fragments and links to long-term knowledge.

What is the strongest criticism of the working memory model?

The central executive is underspecified. Describing it as the part that “allocates attention and decides what to do” can explain almost any result after the fact, which means it risks being a placeholder rather than an explanation. Competing models from Cowan, Engle, and others reframe working memory as activated long-term memory or controlled attention, and explain some data without the separate boxes.

How does the working memory model apply to design and gamification?

The model defines the bottleneck every experience passes through: a workbench of about four chunks, two separate channels, and a manager that anxiety can drain. Good design respects that budget by chunking, splitting across channels, and offloading to the screen. But respecting the limit only lowers the cost of attention. The Octalysis Framework supplies the other half, the motivation that makes a user’s attention manager actually want to spend its scarce capacity on your experience.

References

Baddeley, A. D., & Hitch, G. (1974). Working memory. In G. H. Bower (Ed.), The Psychology of Learning and Motivation (Vol. 8, pp. 47-89). Academic Press.

Atkinson, R. C., & Shiffrin, R. M. (1968). Human memory: A proposed system and its control processes. In K. W. Spence & J. T. Spence (Eds.), The Psychology of Learning and Motivation (Vol. 2, pp. 89-195). Academic Press.

Baddeley, A. D., Thomson, N., & Buchanan, M. (1975). Word length and the structure of short-term memory. Journal of Verbal Learning and Verbal Behavior, 14(6), 575-589.

Conrad, R., & Hull, A. J. (1964). Information, acoustic confusion and memory span. British Journal of Psychology, 55(4), 429-432.

Baddeley, A. D. (2000). The episodic buffer: A new component of working memory? Trends in Cognitive Sciences, 4(11), 417-423.

Baddeley, A. D. (2003). Working memory: Looking back and looking forward. Nature Reviews Neuroscience, 4(10), 829-839.

Shallice, T., & Warrington, E. K. (1970). Independent functioning of verbal memory stores: A neuropsychological study. Quarterly Journal of Experimental Psychology, 22(2), 261-273.

Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81-97.

Cowan, N. (2001). The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences, 24(1), 87-114.

Logie, R. H. (1995). Visuo-spatial Working Memory. Lawrence Erlbaum Associates.

Engle, R. W. (2002). Working memory capacity as executive attention. Current Directions in Psychological Science, 11(1), 19-23.

Ericsson, K. A., & Kintsch, W. (1995). Long-term working memory. Psychological Review, 102(2), 211-245.

Daneman, M., & Carpenter, P. A. (1980). Individual differences in working memory and reading. Journal of Verbal Learning and Verbal Behavior, 19(4), 450-466.

Alloway, T. P., & Alloway, R. G. (2010). Investigating the predictive roles of working memory and IQ in academic attainment. Journal of Experimental Child Psychology, 106(1), 20-29.

Jaeggi, S. M., Buschkuehl, M., Jonides, J., & Perrig, W. J. (2008). Improving fluid intelligence with training on working memory. Proceedings of the National Academy of Sciences, 105(19), 6829-6833.

Melby-Lervåg, M., & Hulme, C. (2013). Is working memory training effective? A meta-analytic review. Developmental Psychology, 49(2), 270-291.

Goldman-Rakic, P. S. (1995). Cellular basis of working memory. Neuron, 14(3), 477-485.

Stokes, M. G. (2015). “Activity-silent” working memory in prefrontal cortex: A dynamic coding framework. Trends in Cognitive Sciences, 19(7), 394-405.

Mayer, R. E. (2009). Multimedia Learning (2nd ed.). Cambridge University Press.

Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257-285.

Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55-81.

Chou, Y. (2015). Actionable Gamification: Beyond Points, Badges, and Leaderboards. Octalysis Media.



WOULD YOU LIKE YU-KAI CHOU TO WORK WITH YOUR ORGANIZATION?

Yukaichou.com Main Contact Form

Continue your training

Reading is XP. Now test what drives you — or pick a quest path.

Keep exploring

Related articles