
Why Most Certifications Are Worthless (And How to Fix Them)
I have a confession that probably shouldn’t be a confession. I think most professional certifications are worthless, and I’ve thought so for a long time — including most of the years I was building one of my own.
What changed my mind about my own certification wasn’t a conference talk or a research paper. It was watching the people who passed multiple-choice tests in adjacent industries get hired into jobs they couldn’t do. Project Management Professional certificate holders who couldn’t run a project. Six Sigma Green Belts who couldn’t diagnose a real defect. Scrum Masters who recited the Scrum Guide flawlessly and then stood frozen when a sprint went sideways.
Their certificates were technically real. The skills they were supposed to certify were technically absent. And every employer who’d been burned by hiring one of them had quietly stopped trusting the credential. Which means every honest holder of that same credential had quietly lost the value they paid for.
So when we built the Octalysis Certification, I refused to put a multiple-choice test anywhere near it. Not because quiz tests are easy. Because they certify the wrong thing.
Speed Run Notes
- Most professional certifications test recognition and memorization (regurgitation) when the job requires application. The credential ends up predicting nothing about whether you can actually do the work, which is why so many employers have learned to ignore them.
- The right test for any certification is the employability test: if I hire this person tomorrow based on their certificate, will I get a competent practitioner? If the answer is “I’d still need to test them in the interview,” the certificate has zero signal value.
- Performance-based certification asks for the work itself, not a description of the work. You hand the candidate a real or realistic problem, watch what they actually do, and grade against the standard a working professional would meet.
- The Octalysis Certification has no Level 3 yet, and that’s intentional. Scarcity protects credential value. Releasing the next level only when the current one has produced enough real practitioners keeps the signal honest.
- Certification design is itself a behavioral design problem. Get the assessment wrong and you train your candidates toward exactly the wrong skill: memorizing flashcards instead of practicing the craft. The test becomes the curriculum, whether you like it or not.
In This Article
- The Regurgitation Test (And Why It Got Normalized)
- The Employability Test: The One Question That Matters
- What Performance-Based Certification Actually Looks Like
- The Transfer Problem (Why Quiz Scores Don’t Predict Job Performance)
- Certification Inflation: How a Credential Loses All Its Signal
- Why Octalysis Has No Level 3 Yet (The Scarcity Strategy)
- The Core Drives Behind a Credential That People Actually Trust
- If You’re Designing a Certification Program
- If You’re Evaluating Certifications (As a Learner or Employer)
- Knowledge → Understanding → Application → Mastery
Why I’m the one writing this
I’m Yu-kai Chou, creator of the Octalysis Framework and the person who designed the Octalysis Certification program. I’ve spent over a decade watching what works and what fails when you try to actually credential a behavioral design skill, not just teach one.
That matters for this post because I built our certification the unpopular way on purpose. Multiple-choice tests are cheaper, faster, and infinitely more scalable than performance-based assessment. Choosing the harder path means I had to get extremely clear about why the easier path destroys the credential it pretends to create. That’s the argument you’re about to read: not academic theory, but the design logic behind a working program.
The Regurgitation Test (And Why It Got Normalized)
Walk into almost any professional certification program in the world right now and you’ll find some version of this loop: watch lectures, read materials, take a test made of multiple-choice questions, get a certificate.
I call this the regurgitation test, because that’s literally what it measures. Can you regurgitate the material back at someone? Can you recognize the right answer when it’s sitting in a list of four options, three of which are obviously wrong?
The standard defense for this design is that it’s “scientific” or “objective.” Multiple-choice tests have a defensible scoring rubric. They’re cheap to grade. They scale. And in fairness, for some narrow knowledge domains — like memorizing specific safety protocols — they’re not crazy.
The problem is that almost no professional skill works that way. Behavioral design certainly doesn’t. A multiple-choice test can confirm that you know the eight Core Drives by name. It cannot confirm that you can walk into a real loyalty program, identify which Core Drives the program is mistakenly relying on, and propose three new feature concepts that fix the imbalance. Those are different cognitive operations entirely.
The reason regurgitation tests got normalized isn’t that they work. It’s that they’re easy to defend. “He passed the test” is a legally tidy answer for an HR department that doesn’t want to take responsibility for a hiring decision. The certificate becomes a liability shield, not a competence signal.
And once that becomes the dynamic, the certificate’s whole reason for existing has flipped. It used to be a way of saying “this person can do the work.” Now it’s a way of saying “you can’t blame me if they can’t.”
The Employability Test: The One Question That Matters
There’s exactly one question I ask whenever I’m designing a certification, and it’s the same question every certification designer should be forced to answer in writing before they ship anything.
If I hire someone tomorrow purely on the strength of this certificate, am I going to get a competent practitioner?
That’s it. That’s the entire test. We can call it the employability test, but it’s really just a forcing function. The moment you write that sentence down honestly, the design pressure on your assessment changes completely.
Suddenly you can’t accept “they passed the multiple-choice test” as proof of anything. You have to imagine yourself as the person whose project is on the line, whose budget is on the line, whose reputation with their boss is on the line. Would you bet your quarter on this candidate? If the answer is anything other than “yes, on the strength of this credential alone,” then the credential is decoration.
This is also the brutal reason most certifications quietly stop being trusted. Employers run their own internal version of the employability test, repeatedly, every time they hire. They notice when a candidate with the certificate still needs six months of ramp-up to hit basic productivity. They notice when the certified candidates aren’t any better than the uncertified ones. And once they notice, they stop weighting the certificate in their decisions. They might still ask for it on the job listing, because that’s HR boilerplate, but mentally they’ve already discounted it to zero.
From the candidate’s perspective, this is the worst possible outcome. They’ve paid the money. They’ve taken the test. They’ve earned the credential. And the actual market signal it sends has degraded to nothing, through no fault of theirs, simply because the program never had the discipline to certify what employers wanted to see.
What Performance-Based Certification Actually Looks Like
Performance-based certification gives the candidate a real or realistic professional problem and grades them on the actual work they produce, instead of testing recognition or memorization with a multiple-choice exam. The certificate certifies job-shaped output, not recall.
Performance-based certification is mechanically simple to describe and brutally hard to scale. You hand the candidate a real or realistic problem from the field, you watch what they actually produce, and you grade their output against the standard a working professional would have to meet.
For Octalysis, that means: here is an actual product, app, or experience. Identify which of the eight Core Drives it’s currently activating, where it’s underweighted, where it’s overweighted, and what specific design changes you’d make to rebalance it. Submit the analysis. Submit the proposed redesign. Then defend your reasoning to someone who actually does this work for a living.
That’s not a test of whether you watched the videos. That’s a test of whether you can do the job.
You can see the difference if you compare what the two assessment styles measure:
| Quiz-Based Approach | Performance-Based Approach |
|---|---|
| Can you name the 8 Core Drives? | Can you identify which Core Drives a live product is actually activating? |
| Can you define Core Drive 2 (Development & Accomplishment)? | Can you brainstorm five new Core Drive 2 features for a specific product context? |
| Can you recognize a player type from a description? | Can you redesign an experience around the actual motivations of the player type you’re targeting? |
| Do you remember the framework? | Can you apply the framework to a domain you’ve never seen before? |
There’s a more general principle hiding underneath this whole comparison, and it’s worth saying out loud. Rewards should flow to the activity that produces the desired behavior, not the activity that measures it. Measurement is cheap and legible, so program designers default to it. But measurement isn’t what you want. You want the thing the measurement is supposed to measure.
A quiz isn’t learning. It’s a low-cost guess at whether learning happened. Pay the candidate (with a credential) for the guess and you’ll get candidates optimized to pass guesses. Pay the candidate for the actual work — the analysis, the redesign, the defense of reasoning — and you’ll get candidates optimized to do the work. The certificate sits at the end of whichever path you reward, and the path is the credential’s real curriculum.
Notice the asymmetry. The quiz column tests whether information made it from a slide into your short-term memory. The performance column tests whether the framework has become a tool you can pick up and use under pressure. Those are not the same skill, and pretending they are is how a credential rots.
The cost is real. Performance-based assessment takes longer to grade. It takes a human evaluator with actual expertise, not a Scantron machine. It can’t be auto-shipped to a million people next quarter. But that cost is also a feature, because it’s the exact reason the credential keeps its signal value over time.
The Transfer Problem (Why Quiz Scores Don’t Predict Job Performance)
If you’ve ever spent time inside a learning sciences department, you’ve heard about the transfer problem. It’s the unsolved Achilles’ heel of education research, and it explains why so many high test scores lead to so much workplace incompetence.
Transfer is the ability to take what you learned in one context and apply it in a different one. Near transfer means using the skill in a closely related situation. You learned to solve quadratic equations in math class, and you can solve a new quadratic equation on a math test. Far transfer means using the skill in a completely new situation. You learned probability in math class, and now you can reason cleanly about whether to take a job offer that has uncertain compensation.
Decades of research have repeatedly found that traditional academic assessment predicts near transfer reasonably well and far transfer almost not at all. People who ace the test in the classroom routinely fail to apply the same concept in a slightly different setting. The skill stayed welded to the testing situation. It never became portable.
This is exactly what employers run into. The certified candidate can pass another version of the same test. They cannot reliably take what’s in the test and apply it to the actual product, in the actual market, on the actual deadline. The credential certifies the wrong dimension of the skill.
Performance-based certification attacks this directly. By grading on output produced in messy, realistic conditions, it forces the candidate to demonstrate at least near transfer (applying the framework to a product they’ve never seen before in the certification context) and gestures toward far transfer (defending their reasoning to a real practitioner who can probe for weak spots). It’s still not perfect (no assessment is), but it’s a measurably stricter standard than recognition.
Certification Inflation: How a Credential Loses All Its Signal
There’s a quiet death spiral that almost every certification program eventually enters, and it has a name: certification inflation.
It usually goes like this. A new credential launches. It’s hard to get. The first wave of holders are highly capable people, partly because the early adopters in any field tend to be the most motivated.
Employers notice. The credential gains signal value. More candidates pursue it. The program faces a choice: keep the bar high, or relax it to capture the larger market.
Almost every program relaxes it. Sometimes it’s explicit: they lower the passing score, simplify the test, shorten the program. Sometimes it’s implicit: they keep the same test, but the test was always testing the wrong thing, and now there are simply more people who’ve memorized their way through it.
Either way, the result is identical. The pool of certified people gets larger and noticeably weaker on average. Employers start running into more candidates who have the credential but not the skill. They lose trust. The signal value collapses. And now the credential is functionally a participation trophy: you have one, your competitor has one, neither of you can use it to differentiate yourself in a hiring market.
The candidates aren’t the villains in this story. Most of them did exactly what was asked of them. The problem is upstream. The program designers chose a model of assessment that scaled easily but didn’t gate honestly, and the inevitable downstream consequence is that everyone who holds the credential now holds something worth almost nothing.
This is why I’m so allergic to scaling certification quickly. The temptation to grow the credential’s holder base is enormous, and almost every move that grows the base also corrodes the signal. It’s not that growth is bad. It’s that growth at the expense of rigor isn’t growth. It’s depreciation of an asset everyone trusted you to protect.
You can watch this happen in real time with the AI-credential boom of the last two years. Once large language models hit the mainstream, dozens of “Certified Prompt Engineer,” “Generative AI Practitioner,” and “AI Specialist” programs launched within months — almost every one of them a recorded course plus a multiple-choice exam.
Anyone who actually hires for AI work can already tell you what’s happening. The certificates are everywhere and the signal is approaching zero, because the assessment never tested whether the candidate could ship a useful AI feature. The certificate runs the same play as PMP and Six Sigma did before it. The category changes. The collapse mechanism doesn’t.
Why Octalysis Has No Level 3 Yet (The Scarcity Strategy)
Here’s something I get asked all the time. The Octalysis Certification has Level 1 and Level 2. Where’s Level 3? When are you launching it? Why are you holding it back?
The answer is that I’m not holding it back so much as I’m refusing to launch it before it’s earned. We have not unlocked Level 3 yet because the goal is that we need a certain critical mass of people to actually reach Level 2 before Level 3 even becomes a coherent thing to design. You can’t design the master class until you have enough genuine practitioners whose work shows you what mastery looks like in the wild.
This serves three purposes at once.
The first is scarcity. As long as Level 3 doesn’t exist, Level 2 is the ceiling, and the ceiling stays meaningful. The moment we ship a Level 3 just to have one, the perceived weight of Level 2 quietly shifts from “the highest you can go” to “the second-highest.” That’s not a marketing point. That’s a credential value point. Your top-end practitioners notice.
The second is quality control. Level 3 should add something Level 2 cannot. If we can’t yet articulate what that something is, based on patterns we’ve observed in actual Level 2 holders’ work, then we shouldn’t ship it. Premature shipping is exactly how programs end up with Level 3 being “Level 2 plus a longer essay,” which is meaningless padding.
The third is community building. The first wave of Level 2 holders becomes the foundation. Their work, their case studies, their teaching becomes the raw material from which Level 3 can later be defined. You can’t build a meaningful master tier without that lineage in place first.
This maps cleanly to Core Drive 6 (Scarcity & Impatience) — and people sometimes hear that and assume it’s a manipulation move. It isn’t. Real scarcity, the kind that holds up, comes from a refusal to ship something before it’s ready, not from artificial drip-feeding. The waiting is honest. The shape of Level 3 will be defined by what Level 2 holders actually produce, and that takes the time it takes.
The Core Drives Behind a Credential That People Actually Trust
Certification design is itself a behavioral design problem. If you don’t think of it that way, you’ll accidentally engineer a program that pulls candidates in for the wrong reasons and trains them in the wrong direction.
The wrong-shaped certification leans almost entirely on extrinsic motivators. Core Drive 2 (Development & Accomplishment) as a status badge to put on a resume. Maybe a hint of Core Drive 4 (Ownership & Possession) from “I now have a certificate.” Get in, regurgitate, get the badge, leave. The candidate’s behavior during the program optimizes for the badge, not the skill. The behavior the program rewards is the behavior the program gets.
Underneath the Core Drive mix is a deeper outcome the right-shaped certification produces, which is identity shift. Not “I took a course,” not “I have a certificate,” but “I’m now someone who knows this domain.” That sentence is the actual prize a candidate is buying. A credential that doesn’t change how the holder sees themselves is a piece of paper with a logo on it. A credential that does change that is a passport into a new version of their working life.
This is also why I obsess over what Octalysis Level 2 holders can do at the end of the program rather than what they can recite. The candidates I’m proudest of describe themselves differently after they earn it — they walk into a meeting saying “I’m a behavioral designer,” not “I’m a marketer who took a gamification class.” The badge is the artifact. The identity shift is the product.
A well-shaped certification leans on a different mix. It still uses Core Drive 2, because earning a real credential is a real accomplishment. But it foregrounds Core Drive 3 (Empowerment of Creativity & Feedback) by asking candidates to actually create work and get expert feedback on it. It uses Core Drive 5 (Social Influence & Relatedness) by giving candidates a real cohort of practitioners they’re earning their place among.
And it uses Core Drive 6 (Scarcity) by being honest that not everyone clears the bar. The credential is meaningful precisely because the bar is real.
The cumulative effect is that the candidate’s experience inside the program looks like the experience of doing the actual job. They aren’t memorizing flashcards. They’re producing artifacts, getting them critiqued, iterating, defending their reasoning. By the time they earn the credential, they’ve already been doing the work for months. That’s not a side effect. That’s the point.
If You’re Designing a Certification Program
If you’re building a credentialing program for any skill (design, engineering, healthcare, project management, AI literacy, whatever), here’s the short version of what I’d tell you, in the order I’d tell it.
- Write the employability test sentence on the first page of your design doc. “If an employer hires our credential holder tomorrow, will they get a competent practitioner?” Re-read it whenever you’re tempted to take a shortcut. It’s amazing how many shortcut decisions look obviously wrong the moment you put them next to that sentence.
- Test application, not memorization. The candidate’s deliverable should be the same shape as the deliverable a working professional in your field produces. If it isn’t, you’re certifying the wrong thing.
- Name what you’re certifying with surgical specificity. “Certificate of Completion” tells an employer nothing. “Certified Octalysis Game Designer,” “Performance-Audited Behavioral Designer Level 2,” “Verified Practitioner: Loyalty Program Architecture” — those tell an employer something actionable. The title carries the specificity of the assessment. If your title is generic, your assessment was generic, and your candidate will hit the market with a credential that doesn’t even claim a verifiable skill.
- Use real or realistic projects as the assessment. Toy problems and contrived case studies are tempting because they’re easier to grade, but they’re also easier to game. Real messiness is the test.
- Grade against the question “would I hire this person?” Not “did they get a passing score on a rubric.” The rubric is a tool, not the standard. The standard is whether the work would survive contact with a paying client.
- Limit supply intentionally. This is the part that feels wrong to most credential programs because it works against short-term revenue. But the credential’s long-term value is its scarcity, and the scarcity has to be defended.
- Don’t ship higher levels until the foundational level has produced enough real practitioners to define what mastery looks like. Premature Level 3 (or Level 4, or “Master Certificate”) is just credential inflation in slow motion.
If You’re Evaluating Certifications (As a Learner or Employer)
Different audience, different practical advice. Same underlying principle.
If you’re a learner deciding whether to invest in a certification, ask these in order:
- Ask “what can certificate holders from this program actually do?” — not “is this program reputable?” Reputation lags reality by years. The right question targets the work the credential is supposed to predict.
- Ask for portfolio examples from existing graduates. Look for case studies, client work, projects they’ve produced. If the program can’t readily show you what its graduates have built, that’s a signal. If the only “proof of value” the program offers is a list of jobs former students got, dig deeper — were they hired because of the credential, or in spite of it?
- Check whether employers in your target field actually respect the credential. Not “have heard of it.” Respect. There’s a difference between a recruiter checking a box and a hiring manager actively wanting candidates with that credential.
- Strongly prefer performance-based programs over quiz-based ones. The cost is higher upfront, but the credential you walk away with is more likely to still be worth something five years from now.
If you’re an employer screening candidates, the rule is simpler:
- Don’t trust “certified” without knowing what specific program they’re certified by, and what specifically that program tests.
- Ask for work samples, not just certificate numbers. Test the actual application of the skill in your interview process. Even the best certification can’t replace your own evaluation, and the worst ones aren’t even worth the time it takes to verify them.
Knowledge → Understanding → Application → Mastery
The deeper principle underneath all of this is the ladder that most education accidentally stops climbing.
The full ladder runs from knowledge (you know the facts) to understanding (you can explain why) to application (you can actually use it) to mastery (you can apply it to situations the framework wasn’t originally designed for and get good results).
Most professional education stops at understanding. The student can recall the material and explain the logic, and that’s where the curriculum ends and the certificate gets issued. The bottom two rungs of the ladder are mistaken for the whole ladder.
Learning-sciences researchers have a name for this ladder. It’s called Bloom’s Taxonomy, and the standard version actually runs six levels deep — Remember, Understand, Apply, Analyze, Evaluate, Create. The four-rung version I’m using here is a simplification, but it captures the gate that matters: the line between Understand and Apply. Most professional education stops just before that gate. Most degree programs stop just before that gate. And, predictably, most certifications stop just before it too.
Performance-based certification is what it looks like when a program decides to push the gate up to Apply, and on its best days nudges candidates into Analyze and Evaluate when they have to defend the reasoning behind their work. Mastery — the Create level, where the practitioner generates novel behavioral-design solutions for situations the framework wasn’t built for — is reserved for the field. No assessment can credential it. Only years of real work can.
Mastery, in this scheme, is reserved for what comes after the credential — the part that can only be earned in the field, on real problems, under real constraints, over years. No assessment can credential it.
That’s why I’m patient about Level 3. That’s why I won’t put a multiple-choice test on Level 1. And that’s why I think the entire industry of credential design is overdue for an honest reckoning. The credentials that survive the next decade are going to be the ones that actually predict job performance — because that’s the only credential employers will keep paying attention to.
If you want a closer look at how Octalysis is taught and assessed, the best place to start is the Octalysis Certificate page — it walks through how each level works in practice and what a real performance-based assessment looks like end to end. (If you want to back up first and understand the underlying framework the certification tests competence in, the Octalysis Framework page is the canonical reference.)
Related Reading
- The Octalysis Framework: The Complete Guide to Behavioral Design
- Octalysis Certificate: Prove Your Gamification Design Skills
- Core Drive 3: The Golden Core Drive of Empowerment & Feedback
- Core Drive 6: Scarcity & Impatience
- Books by Yu-kai Chou: Actionable Gamification & 10,000 Hours of Play
About the Creator of the Octalysis Framework
Yu-kai Chou created the Octalysis Framework after studying gamification since 2003 — years before the term entered mainstream vocabulary. As a Human-Systems Architect & Behavioral Designer, his framework has been applied by LEGO, Microsoft, Porsche, Coca-Cola, Salesforce, and MrBeast, impacting over 1.5 Billion Users.
Chou has taught the Octalysis methodology at Harvard, Stanford, Yale, Tesla, Google, BCG, and IDEO.
His work has been cited by Harvard, Stanford, MIT, Forbes, Wall Street Journal, Wired, US Department of Energy, NIST, NSF, NCBI, US Department of Education, ClinicalTrials.gov, and 3,700+ more academic publications. If you want to develop the same skill the post argues most certifications fail to credential, you can start with the Octalysis Certificate program. Explore his books here.

