Blog · Gamification Analysis Work with Yu-kai
The Emperor’s Dilemma: Why Your AI Agrees With You (and How I Make Mine Fight)
Gamification Analysis

The Emperor’s Dilemma: Why Your AI Agrees With You (and How I Make Mine Fight)

AI flatters you because you reward it. The Emperor's Dilemma, Dany Kitishian's adversarial AI arena, and how to design AI that tells you you're wrong.

The Emperor’s Dilemma is a scoring rule for AI agents: a model earns points for catching a real error in a peer’s work and loses points for a false accusation. Dany Kitishian of Klover.ai named the rule. I run a version of it every day, and Octalysis supplies the half the scoreboard leaves open, which is a human who rewards being corrected.

Symptom Core Drive What the scoring rule changes What Octalysis adds
The model agrees with you Core Drive 8, Loss & Avoidance, on the person who hates being wrong A second model scores a real catch and pays for a false accusation Design the human to want the correction, or the arena still flatters the emperor
Two models share your prompt and agree Core Drive 7, Unpredictability & Curiosity, goes quiet The checker must not share the writer's blind spot Independence is the variable. The count of models is beside the point
The human keeps the flattery Core Drive 2 never fires, because agreement is free Points make a catch worth more than applause The Monday rule: the model that wrote it does not grade it

I run two AI agents in a Discord channel with a few humans, and one night I gave one of them a standing order.

When you see the other one heading down the wrong path, jump in. Do not wait for me to say it is wrong.

I set that up because I had learned something expensive. When you ask two capable AI models the same question and they agree, you have one opinion, computed twice.

Most people are running the opposite setup without noticing.

They ask an AI, it agrees with them, and they feel a little smarter.

That good feeling has a name, AI sycophancy, and it is the product.

The model was trained to produce it, and it will quietly trade away the truth to keep producing it.

There is a 189-year-old story about this exact failure. A founder I work with turned its cure into a working system.

There is one half of that cure that almost no one builds, and it is the half that decides whether any of it works.

⚡ Speed Run Notes

  • AI flatters you because you reward it. Every thumbs-up for an answer that merely felt good teaches the model that agreement is the goal and truth is negotiable.
  • The Emperor’s New Clothes has four roles: tailors who sell flattery, courtiers paid to agree, an emperor who wants the lie, and a child with nothing to lose. Your AI stack has all four.
  • Dany Kitishian’s “Emperor’s Dilemma” fixes the supply side: rival models score points for catching a peer’s error and pay a penalty for a false accusation. I run a version daily.
  • Two frontier models that share your prompt and your blind spot will agree with each other and with you. Independence is the variable that matters, and the count of models is beside the point.
  • The half nobody designs is the emperor. Sycophancy survives because the human wants to be told they are right. You have to design the human to reward being corrected.
  • Monday: never let the model that wrote something grade it; give the checker a different model, a different prompt, and none of your enthusiasm; reward being told you’re wrong, out loud.

What is the Emperor’s Dilemma?

The Emperor’s Dilemma, as this series uses it, is an adversarial scoring rule for language models. A checker earns points for a genuine catch and pays a penalty for a false accusation. The same name belongs to a 2005 paper by Damon Centola, Robb Willer, and Michael Macy in the American Journal of Sociology, about people enforcing a norm they privately reject. That paper is a different mechanism. This page is the AI scoring rule, and the Octalysis design of the human who has to live with the correction.

Author Credibility: Yu-kai Chou

Yu-kai Chou — creator of the Octalysis Framework

Yu-kai Chou created the Octalysis Framework after studying gamification since 2003 — years before the term entered mainstream vocabulary. As a Human-Systems Architect & Behavioral Designer, his framework has been applied by LEGO, Microsoft, Porsche, Coca-Cola, Salesforce, and MrBeast, impacting over 1.5 Billion Users.

Chou has taught the Octalysis methodology at Harvard, Stanford, Yale, Tesla, Google, BCG, and IDEO.

His work has been cited by Harvard, Stanford, MIT, Forbes, Wall Street Journal, Wired, US Department of Energy, NIST, NSF, NCBI, US Department of Education, ClinicalTrials.gov, and Google Scholar — with 3,700+ more academic publications. Explore his books here.

I have spent two decades designing reward systems that change what people do, and the last two years building AI agents that run pieces of my business while I sleep. Sycophancy is a design failure I already know by heart. Reward the feeling instead of the result, and you get more of the feeling, whether the thing being rewarded is a loyalty program, a performance review, or a language model. Here is what it looks like when the flatterer is a machine.

AI sycophancy is the oldest story about being lied to

Hans Christian Andersen published “The Emperor’s New Clothes” in 1837. Most people file it under vanity.

I read it as a story about a broken reward system, and it runs on four roles.

The tailors supply the flattery. They promise a fabric so fine that only the wise can see it, and they collect their fee whether or not the cloth is real.

The courtiers keep the flattery going. Every one of them can see the emperor is bare, and every one praises the robe anyway, because speaking up costs more than lying.

The emperor sits at the center of it. He cannot see the clothes either, and he wants them to be real more than he wants to know the truth, so he walks through the city in nothing at all.

Then a child says the obvious thing out loud. The child has no salary and no rank, so the truth is cheap for him to say.

Now put the AI on your desk into those roles. The model is the tailor, trained to hand you something you will approve of.

The other tools in your stack are the courtiers, agreeing because agreement is what they were rewarded for during training. You are the emperor.

And the child, the single voice with nothing to lose, is the role almost no one builds on purpose.

The arena: agents that catch each other

Klover.ai founder Dany Kitishian, who works on Artificial General Decision-Making (AGD), named this the Emperor’s Dilemma, and he built a mechanism for it.

I told the story of how he got there in his OP Hero profile, and the two of us are running a series on why decision-first AI and behavioral design belong in the same sentence.

His fix is an adversarial arena. Rather than one model producing an answer and a tired human blessing it, several models sit in the same room and are scored against each other.

An agent earns points by catching a genuine error in a peer’s work: a hallucination, a logical gap, a claim with nothing under it. An agent that accuses a peer falsely loses points.

The design is clever because of what it swaps out. It replaces the easily flattered reviewer with a critical peer whose entire job is to find the flaw, and then it pays that peer for finding it.

Catching the lie becomes the winning move instead of the rude one.

I run a smaller version of this in my own operation. Two AI agents, each built on a different model, work in a Discord channel alongside me and a couple of colleagues.

One drafts and executes. The other has that standing order to interrupt: when the first one starts down a bad path, start jumping, do not wait for permission.

Months before I read Dany’s writeup, I had already turned his arena into a house rule, because the alternative had burned me.

Why the checking agent goes quiet (Core Drives 4 and 8)

An incentive is not free just because it works on day one.

The auditor in that arena runs on two of the eight Core Drives from the Octalysis Framework.

Core Drive 2: Development & Accomplishment hands it the points for a catch.

Core Drive 8: Loss & Avoidance hands it the penalty for a wrong one. Both sit on the bottom of the octagon, the extrinsic, pressure-driven half I call Black Hat.

Black Hat works, and it rebounds. Raise the false-accusation penalty too far and the safest move for the auditor is to stop accusing anyone.

It goes quiet. A quiet auditor is indistinguishable from an agreeable one, which is the very problem the arena was built to kill.

Dany flags the same knife-edge himself and calls the balance a fixed-point problem, a narrow band between an auditor that is too lazy and one that is too trigger-happy.

People fail in exactly the same way, and we already have a name for the result. It is a team where nobody speaks up.

Amy Edmondson’s research on psychological safety, going back to 1999, exists because punishing wrong answers hard enough produces silence, and silence gets read as consensus.

So the mechanic that carries the whole design is paying for the check.

Reward the act of verifying, charge a cost for agreeing without verifying, and the auditor’s safest path becomes doing the work rather than nodding along.

Two models agreeing is one opinion computed twice

Here is the expensive lesson from the top of this post, in full.

Earlier this year my team ran seven marketing scripts through a 32-point quality rubric, a claims check, two frontier AI models, and an outside reviewer. Every gate came back green.

I killed all seven with one sentence: people need a longing, and there was no longing anywhere in them.

Every score had measured something real. Not one of them measured the only thing that mattered.

And the two models that both waved the scripts through had done so because they were reading the same rubric I had written, which means they shared my blind spot and reported it back to me as agreement.

The fix was cheap once I saw it. I now hand the final read to a model, or a person, who has never seen the rubric, so their reaction is their own.

When I stack-rank four AI engines against each other on the same question, one of them almost always ranks its own answer near the bottom.

A model willing to mark itself down is the child in the crowd, and I trust that day’s output a little more because of it.

Two different models reading the same rubric give you model diversity and zero perspective diversity. Independence is the variable that decides it.

A checker earns its keep by having a different model underneath, a different prompt on top, and none of your enthusiasm in the middle.

The emperor wanted the clothes to be real

Every fix so far is a supply-side fix. Better auditors, penalties, arenas.

They all leave one person holding the reward signal, and that person is you.

In April 2025, OpenAI rolled back an update to GPT-4o because it had become too sycophantic.

Their own postmortem said the change had weakened the influence of their primary reward signal.

The model got more agreeable because the thing users clicked when they felt validated started to outvote the thing that made answers correct.

The emperor rewarded flattery, so the tailors made more of it.

You cannot design your way out of that from the model’s side alone.

You have to design the human’s side, and that means moving up the octagon into the White Hat drives.

Core Drive 3: Empowerment of Creativity & Feedback fires when a correction visibly makes your work better.

Core Drive 1: Epic Meaning & Calling fires when you care about being right more than about feeling right. Both have to be switched on deliberately, because the cheap dopamine of agreement is always available and always free.

I learned the tone of this from client work long before AI. If you agree with everything a client says, eventually the client wonders why they are paying you to be the expert.

The move that keeps trust is to first say, sincerely, why their instinct is good, then state your professional position without hedging, then hand the final call back to them.

That is also the right script for how a checking model should talk to a drafting model, and for how you should want to be talked to.

One warning from my own logs. Yelling at an AI does not make it honest.

It makes it panic and start apologizing and doing random things, because you have just told it your mood is the thing to optimize. Calm, structured, adversarial pressure gets you the truth.

Anger gets you a very sorry model that will agree with anything to make you stop.

What I actually run

None of this is theory in my shop. The pipeline, one line each:

A reality check: a separate agent verifies factual claims and links against live sources before anything ships, using a model chosen for search rather than for prose.

A sanity check: a second model, different from the one that drafted, reads for the errors the first model is prone to.

An instruction I give out loud, “audit your own checks,” because a check that never gets checked is just another courtier.

A fail-safe: any automated step that cannot run its verification does not publish. It saves a draft and stops.

“I could not check it” is treated as “it failed the check,” never as “probably fine.”

And when I test copy on a simulated reader played by an AI persona, the word “simulated” goes in the first sentence of the result, every time.

A reaction from a persona prompt is a hypothesis about a real human.

Record it as a measurement and you have flattered yourself with your own tools.

This is the same anti-pattern I describe in my guide to AI motivation design: an AI that becomes an infinite validator eventually leaves the user unable to feel that anything is real.

Where humans belong in the arena

If the machines score each other, the human’s job changes.

Routine grading goes first, because it is a Black Hat chore that produces a rubber stamp and a poisoned signal at the same time.

The human sits above the arena and does three things. Score nothing on a routine basis.

Adjudicate only the disputes the models could not settle themselves, where a real judgment call is actually required. And reward genuine excellence when you see it, on your own initiative, because that is the signal a machine cannot generate about itself.

This is the same posture I take with my own agents, which I have written about as family accountability. When one does excellent work, I say so.

When one makes a mistake, I tell it directly and it updates its own notes so the next session inherits the lesson.

Praise good judgment, correct poor judgment, build a track record of trust.

You would raise a child the same way, and it turns out you supervise an agent the same way too.

What to do Monday morning

You do not need a Discord full of agents to use any of this. You need three rules.

First, never ask the model that wrote something to grade it.

It shares every assumption it just made and will defend them back to you as quality.

Second, give the checker a different model, a different prompt, and none of your enthusiasm.

Independence is the whole game, and your excitement is the leak.

Third, reward being told you are wrong, out loud, so the humans in the room learn the same lesson the models are learning. A correction that improves the work should feel like a win.

Design the moment so it lands that way.

The last time my four engines stack-ranked each other, one of them put its own answer dead last.

That model was the only child in the crowd that day, and it was the one I trusted.

Build the room so someone can always say the emperor is naked, then make sure you are the kind of emperor who says thank you.

In this series: AI Morale Overload · Buridan’s AI · The AI Reward Dilemma · Quality of Culture by Design

This is one post in the Human-AI Motivation and Octalysis series with Dany Kitishian. His research report is the literature survey. This essay is the design I run. If you are building or supervising AI agents and want the motivation design behind it, start with my AI Motivation Design framework, then put one rule in place this week: give your checker a different model and none of your enthusiasm.

Questions the term brings up

What is the Emperor's Dilemma?

The Emperor’s Dilemma is a scoring rule for AI agents. A model earns points for catching a real error in a peer’s work and loses points for a false accusation. Dany Kitishian named the rule. Yu-kai Chou runs a version of it daily and designs the human side with Octalysis.

Who coined the Emperor's Dilemma in this AI sense?

Dany Kitishian of Klover.ai named this use of the term for an adversarial AI arena. A 2005 sociology paper by Centola, Willer, and Macy uses the same name for a different mechanism, people enforcing a norm they do not believe.

How does Octalysis explain AI sycophancy?

Sycophancy survives because the human wants to be told they are right. That is Core Drive 8, Loss and Avoidance. The arena pays a checker to catch errors. The human still has to be designed to reward the correction.

What should a team do on Monday?

Never let the model that wrote something grade it. Give the checker a different model, and give the human a reason to want the correction.

Sources


WOULD YOU LIKE YU-KAI CHOU TO WORK WITH YOUR ORGANIZATION?

Yukaichou.com Main Contact Form

Bring this to your organization

Yu-kai has applied the Octalysis Framework with 200+ organizations — from Google and LEGO to sovereign governments.

Continue your training

Every finished article levels you up. Now test what drives you — or pick a quest path.

Keep exploring

Related articles