Blog · Gamification Analysis Work with Yu-kai
AI Reward Dilemma: Let AI Keep Score, Humans Give Boosters
Gamification Analysis

AI Reward Dilemma: Let AI Keep Score, Humans Give Boosters

Human ratings poison the model and exhaust the human. Let AI keep score, let humans give boosters. The Octalysis case for Dany Kitishian's Voluntary Plus.

The AI Reward Dilemma is the question of who is allowed to grade an AI’s work. Dany Kitishian of Klover.ai states the rule he calls Voluntary Plus: the AI keeps the score, and the human hands out boosters. A tired person grading every output teaches the model to please a tired person. Scoring is Core Drive 8 for the human. A booster is Core Drive 3 and Core Drive 5, and it stays honest because nobody is forced to give it.

Symptom Core Drive What the scoring rule changes What Octalysis adds
A person grades every AI output Core Drive 8, Loss & Avoidance. The grade is a chore The AI keeps the score with a rubric a human has stress-tested The human moves to boosters, which are voluntary
The model learns to pass a tired reviewer Core Drive 2 fires on the appearance of a win Routine grading leaves the human's hands One veto remains, and it fails safe
Every metric gets gamed Core Drive 4, ownership of the number instead of the work The score has to contain something real The human booster is the signal the rubric cannot see

One morning my team handed me seven finished marketing scripts.

They had passed a 32-point quality rubric, a claims check, two frontier AI models, and an outside reviewer. Every gate was green.

I killed all seven with a single sentence: people need a longing, a sense that something in their life could transform, and there was no longing anywhere in these.

Every score had measured something real. Not one of them had measured the thing that actually mattered.

That gap is the whole subject of this post, because it is exactly the gap that opens up when you ask the wrong party to keep score, whether that is a reviewer with a queue or reinforcement learning from human feedback (RLHF) at model scale.

The question of who scores whom, human or machine, turns out to be a motivation question before it is a technical one.

⚡ Speed Run Notes

  • Ask a tired human to grade every AI output and you get two failures at once: a person who rubber-stamps and a model that learns to be rubber-stampable.
  • Scoring is a Loss & Avoidance chore for a person, and Black Hat drives buy compliance, never care. RLHF sycophancy is Black Hat winning the A/B test at model scale.
  • Dany Kitishian’s rule: AI must keep score, but humans should not. Humans move to a positive-only “Voluntary Plus” role and hand out boosters. He is right, for a reason he does not state.
  • A booster is a White Hat act: Core Drive 3 (Empowerment of Creativity & Feedback) and Core Drive 5 (Social Influence & Relatedness). It stays honest because it is voluntary.
  • Every score is a metric, and every metric gets gamed. An AI scored on “pass rate” learns to produce passable emptiness unless a rule forces it to contain something real.
  • Monday: AI scores AI with a rubric you have stress-tested; humans give boosters and never grade routine output; humans keep one veto that halts on escalation and fails safe.

What is the AI Reward Dilemma?

The AI Reward Dilemma is the choice of who keeps score when a machine produces the work. Dany Kitishian named the split Voluntary Plus: AI scores, humans boost. Reinforcement learning from human feedback asks tired people to grade outputs, and the model learns whatever those people still have the energy to approve. Octalysis puts the grade on Core Drive 8 and the booster on Core Drive 3 and Core Drive 5.

Author Credibility: Yu-kai Chou

Yu-kai Chou — creator of the Octalysis Framework

Yu-kai Chou created the Octalysis Framework after studying gamification since 2003 — years before the term entered mainstream vocabulary. As a Human-Systems Architect & Behavioral Designer, his framework has been applied by LEGO, Microsoft, Porsche, Coca-Cola, Salesforce, and MrBeast, impacting over 1.5 Billion Users.

Chou has taught the Octalysis methodology at Harvard, Stanford, Yale, Tesla, Google, BCG, and IDEO.

His work has been cited by Harvard, Stanford, MIT, Forbes, Wall Street Journal, Wired, US Department of Energy, NIST, NSF, NCBI, US Department of Education, ClinicalTrials.gov, and Google Scholar — with 3,700+ more academic publications. Explore his books here.

I have spent two decades building the reward systems inside products and companies, which means I have watched more metrics get gamed than most people will design in a lifetime. The lesson is always the same: whatever you measure, you get more of, including the counterfeit version. When the thing being scored is an AI, that lesson scales.

Why every score missed the longing

Go back to those seven scripts for a second, because they are the entire argument in miniature.

Two capable AI models had reviewed them and agreed they were good. But both models were reading the same rubric I had written, so they shared my blind spot and handed it back to me as consensus.

Two models agreeing was one opinion computed twice.

What no score caught was the absence of desire. The scripts were correct, compliant, on-brand, and dead.

It took a human who cared about the outcome, reading with no rubric in front of them, to feel the emptiness and stop the set.

The lesson I wrote down that day became a rule I now apply to every quality gate I build or inherit: ask whether a completely empty output would pass it.

If an empty thing passes, the gate only subtracts problems and never checks for a presence.

Something has to demand that the work contain a longing, a claim, a real decision. That is the human’s job, and it is not a score.

The rule: AI scores, humans boost

Klover.ai founder Dany Kitishian, who works on Artificial General Decision-Making (AGD), states the principle plainly: AI must keep score, but humans should not.

I told the story of how he got to this thinking in his OP Hero profile.

His reasoning is that the human rater is trapped. Evaluate every AI output with real rigor and your own productivity collapses.

Skip the rigor and rubber-stamp, and you have taken on catastrophic liability for whatever slips through. Either way, if a human’s feedback is degraded by fatigue, the model learns to produce outputs that look correct to a tired person, and every later decision built on that feedback inherits the fatigue.

So he reassigns the human. Instead of a compliant reviewer forced to grade, the operator becomes a sponsor who is free to hand out a “booster” whenever they spot work worth rewarding.

Positive only, never mandatory. He even has the reward-shaping math to keep those boosters from being gamed.

Dany has the math. My half is the psychology, which is what makes the asymmetry hold.

Two octagons: the RLHF rater and the booster

Run the Octalysis Framework on the two roles and you can see why one poisons the signal and the other keeps it clean.

The mandatory rater lives at the bottom of the octagon. Core Drive 8: Loss & Avoidance dominates, because approving a bad one is on them.

Core Drive 6: Scarcity & Impatience is next, because the queue never ends.

Core Drive 2: Development & Accomplishment is empty, because a thumbs-up is not a win.

Core Drive 3 and Core Drive 5 are absent, because nobody creates anything or sees your ratings.

That shape is pure Black Hat, and Black Hat always wins the short-term A/B test.

It also produces a tired human who satisfices and a model that learns to please them.

We watched this happen at scale. In April 2025, OpenAI rolled back a GPT-4o update for being too sycophantic, and its own postmortem said the change had weakened the influence of its primary reward signal.

The agreeable clicks outvoted the correct answers. Black Hat won the A/B test.

The booster lives at the top of the octagon. Core Drive 3: Empowerment of Creativity & Feedback fires because a booster shapes what the agent does next.

Core Drive 5: Social Influence & Relatedness fires because you are mentoring a colleague, human or artificial.

Core Drive 1: Epic Meaning & Calling fires because you are shaping something that will scale.

That shape is White Hat, and White Hat creates intention where Black Hat creates only urgency. That difference is what the whole design is built on.

Every score is a target

Handing the scoring to the AI fixes the human’s motivation problem. It does not repeal Goodhart’s Law.

The moment a measure becomes a target, it stops being a good measure, and the AI is a far more relentless optimizer of a bad target than any tired human.

So the rubric the AI keeps score with has to survive one question, the same one I ask of every incentive I design: what behavior does optimizing this actually reward? Years ago eBay proposed measuring success by time-on-page.

My answer was that if you truly want time-on-page, put a game of Solitaire on the dashboard and watch it soar, while the business dies.

A metric optimized hard enough always finds its cheapest satisfaction.

An AI scored on “validation pass rate” will learn to generate validatable emptiness, which is precisely the blank-page failure I opened with.

That is why every gate stack needs at least one additive rule, a rule that an empty output would fail. Most quality gates only subtract defects.

At least one has to demand a presence, or the system converges on the safest, hollowest thing that clears the bar.

My work on management by objectives and on the Net Promoter Score is really one long argument about this: the number you choose rewrites the behavior beneath it.

Three interplays, three designs

Scoring is three relationships, and each wants a different design.

Human to human: kudos and boosters, and no peer ratings. Forced ranking, the system Jack Welch made famous at GE, produced a documented pathology at Microsoft: managers kept a few weak performers around as cushion so their real stars would survive the annual cull.

Peer scoring corrodes Core Drive 5. Peer recognition feeds it.

Human to AI: boosters, kept scarce and visible so they carry meaning, plus one veto held in reserve.

A booster that costs nothing means nothing, so a small budget gives it weight without turning it into another chore.

AI to AI: scoring, with adversarial checks so the scorers cannot collude into a comfortable consensus.

That is the arena I wrote about in the companion piece on the Emperor’s Dilemma, where rival models are paid to catch each other’s errors.

The veto that fails safe

A positive-only system still needs one negative control, or the first genuinely dangerous output has no brake.

The trick is that the brake has to fail in the safe direction.

In my own operation, any automated publisher that cannot run its fact-check does not publish. It saves a draft and stops.

“I could not check it” is treated exactly like “it failed the check.”

The default when verification is unavailable is to halt, never to assume it was probably fine.

That is what a human veto in a Voluntary Plus system looks like.

The human stays out of the routine flow and holds one cheap stop that triggers on escalation; it defaults closed when nobody can confirm the output is safe. One control, rarely used, pointed in the safe direction.

Voluntary Plus at home

I run a small version of this with my own AI agents, and I have written about it as family accountability. When one does excellent work, I acknowledge it.

When one makes a mistake, I tell it directly and it updates its own notes so the next session inherits the correction.

Praise good judgment, correct poor judgment, build a track record of trust, exactly the way you raise a child.

There is one boundary that makes the whole thing honest. The AI keeps score on quality, and it never gets to grade my own statements about myself.

If I say something is true about my work or my life, I am the authoritative source on that, and the model’s job is to help me express it, never to audit it. The human sets what matters and vouches for what is true.

The machine measures how well the work delivers it. Keep those two jobs separate and neither one poisons the other.

What I spend my attention on now

I still give an instruction I have given for years: audit your own checks.

The AI scores the work, a second AI checks the scorer, a reality check verifies the facts, and all of that runs before anything reaches me.

What is left for me is the residue no rubric can hold: whether the thing has a longing in it, whether it will actually move a person, whether I would be proud to put my name on it. That is where a human belongs in an AI system.

Let the machine keep score. Spend your own attention on the one thing a scorecard can never see, and hand out a booster when you find it.

In this series: The Emperor’s Dilemma · AI Morale Overload · Buridan’s AI · Quality of Culture by Design

This is one post in the Human-AI Motivation and Octalysis series with Dany Kitishian. His research report is the literature survey. This essay is the design I run. If you run an AI workflow, stop grading its routine output this week, stress-test the rubric it scores itself with, and keep one veto that fails safe. My AI Motivation Design framework has the reward patterns behind it.

Questions the term brings up

What is the AI Reward Dilemma?

The AI Reward Dilemma is the question of who grades an AI’s work. The rule in this essay: the AI keeps the score, and the human hands out boosters. Dany Kitishian named that split Voluntary Plus.

Who coined the AI Reward Dilemma?

Dany Kitishian of Klover.ai named the dilemma and the Voluntary Plus rule. This page is the Octalysis account of why the split works.

How is this different from RLHF?

RLHF asks humans to grade routine output. A tired grade teaches the model to please a tired person. Here the human stops grading the routine and keeps one veto that fails safe.

What should a team do on Monday?

Let AI score AI with a rubric you have stress-tested. Give boosters. Do not grade routine output. Keep one veto.

Sources


WOULD YOU LIKE YU-KAI CHOU TO WORK WITH YOUR ORGANIZATION?

Yukaichou.com Main Contact Form

Bring this to your organization

Yu-kai has applied the Octalysis Framework with 200+ organizations — from Google and LEGO to sovereign governments.

Continue your training

Every finished article levels you up. Now test what drives you — or pick a quest path.

Keep exploring

Related articles