Big Picture 10 min read

Does Gamification Actually Work? The Evidence

The research includes null results, measured decay curves and documented harm. It also shows what separates gamification that sustains engagement from gamification that fades.

By Asa Goldstein, QuestWorks

TL;DR

The most-cited paper in the field, titled "Does Gamification Work?", concludes that effects depend heavily on context and on the users involved. The best meta-analysis available finds small effects (g = .25 on behavioral outcomes) and calls the factors behind successful gamification still somewhat unresolved. Contingent rewards measurably reduce intrinsic motivation across 128 experiments. A 14-week study of 756 students shows engagement declining from week 4 and partially recovering between weeks six and ten as familiarity takes over from novelty. Klout shows what happens when a gamified score can be optimized independently of the behavior it claims to represent. The systems that hold up share a structure: repeated practice on a fixed schedule that builds competence, participation people choose, and feedback carrying information about performance.

The most-cited paper in the gamification literature is titled "Does Gamification Work?" It has been cited more than 12,000 times and gets quoted constantly as proof that the answer is yes. What the authors concluded was narrower: "gamification provides positive effects, however, the effects are greatly dependent on the context in which the gamification is being implemented, as well as on the users using it" (Hamari, Koivisto and Sarsa, Hawaii International Conference on System Sciences, 2014).

The hedge is the finding. Gamification is a design technique with a very wide outcome distribution. Some deployments produce durable behavior change. Most produce a spike that decays. A few leave people less motivated than they started. Any useful answer to "does gamification work" has to explain which conditions sort those outcomes, and the research does have something to say about that.

What the Evidence Base Actually Shows

The strongest quantitative synthesis available is Sailer and Homner's meta-analysis, which found what the authors themselves call small effects across the studies it pooled: g = .49 on cognitive outcomes, g = .36 on motivational outcomes, and g = .25 on behavioral outcomes (Sailer and Homner, Educational Psychology Review, 2020). Those are positive numbers. They are also modest. An effect size of .25 on behavior means a detectable shift in the average with heavy overlap between the gamified and non-gamified groups. Plenty of individuals in the gamified condition did worse than the median person who got nothing.

Two caveats travel with that number. The first is context: this is an education meta-analysis, and classroom findings transfer to workplaces imperfectly. The second is the authors' own: heterogeneity across studies was substantial, the moderators they tested did not account for it, and they close by calling the factors that contribute to successful gamification "still somewhat unresolved, especially for cognitive learning outcomes". Their verdict on gamification as currently practiced is positive; what the field's best synthesis cannot yet supply is the account of why it works.

That is consistent with how fragmented the underlying theory is. A systematic review of the field found 118 distinct theories invoked across gamification, serious games and game-based learning research to explain the same phenomenon (Krath, Schürmann and von Korflesch, Computers in Human Behavior, 2021). From the interrelations between those theories the authors derive a set of basic principles for how gamification works: illustrating goals and their relevance, nudging users along guided paths, giving immediate feedback, reinforcing good performance, simplifying content into manageable tasks, adapting complexity to a user's abilities, letting users pursue individual goals and choose between progress paths, and enabling social comparison and mutual support. That is the closest thing the field has to consensus: a set of ingredients assembled out of 118 theories built for other purposes, rather than one tested account of why the recipe works.

If you have been sold gamification as settled science, this is the state of the science. (For the definitional boundaries between gamification, serious games and game-based learning, which vendor copy routinely blurs, see serious games vs. gamification vs. game-based learning.)

The Mechanism That Makes Reward Systems Backfire

The sharpest criticism of gamification predates gamification by two decades. Deci, Koestner and Ryan published a meta-analysis of 128 experiments examining what happens to intrinsic motivation when you attach extrinsic rewards to an activity people already find interesting. Engagement-contingent, completion-contingent and performance-contingent rewards all significantly undermined free-choice intrinsic motivation, with effect sizes between -0.28 and -0.40 (Deci, Koestner and Ryan, Psychological Bulletin, 1999).

The detail most summaries drop is the exception. Verbal praise and positive feedback increased both behavioral engagement and interest in the same analysis. The harm is specific to contingent reward structures of the form "do X, receive Y." Recognition delivered as information about how someone performed moves motivation the other way.

That distinction separates a points economy from useful feedback, and most corporate gamification lands on the wrong side of it. Gneezy, Meier and Rey-Biel made the same case from economics: incentives change how a person perceives the task itself, which can produce negative effects on behavior even when the incentive hits its short-term target (Gneezy, Meier and Rey-Biel, Journal of Economic Perspectives, 2011). They traced the pattern across education, public-goods contributions, smoking and exercise, and hedged every verdict: well-specified, well-targeted education incentives show moderate success, "although the jury is still out regarding the long-term success of these incentive programs," while incentives for lifestyle change work in the short and middle run before the habit change disappears again. Their illustration: an incentive for a child to read more might hit that target in the short term and still work against the child's enjoying reading and seeking it out over a lifetime.

Current design literature names the same failure. A 2023 review identifies overreliance on narrow models and shallow design, meaning badges, points and leaderboards and little else, as producing an overjustification effect that actively undermines the intrinsic motivation the gamification was meant to support (Dah John et al., International Journal of Serious Games, 2023). Badly built gamification can leave a behavior in worse shape than no intervention at all.

The Decay Curve Has Been Measured

"The novelty wears off" is the standard objection to gamification, usually asserted without numbers. It has numbers.

A 14-week longitudinal study with 756 students compared gamified and non-gamified assignment environments and tracked engagement week by week (Rodrigues et al., International Journal of Educational Technology in Higher Education, 2022). Engagement followed a U-shaped curve. The initial gains from gamification started declining around week 4. The dip persisted for two to six weeks. Engagement then partially recovered somewhere between weeks six and ten, which the authors attribute to a familiarization effect taking over from novelty.

Two things follow. The first is a measurement problem. A pilot that runs six or eight weeks can land squarely in the decline and stop before the recovery shows. A pilot that runs three weeks captures only the novelty spike. Most corporate pilots are scoped to exactly the windows that produce misleading readings, and most vendor case studies are built from those pilots.

The second concerns what replaced the novelty. Engagement in that study converted into something else. What held it up later in the term was familiarity with a practiced environment, and familiarity accumulates through repetition. Novelty and familiarization are two different engines, and only one of them can run indefinitely.

When the Score Becomes the Target

Klout is the cleanest documented case of a gamified metric measuring the wrong thing at scale. Launched in 2008, it compressed "social influence" into a single 1-100 score and attached a rewards program, Perks, in which companies gave free products to high scorers. By May 2013 users had claimed over 1 million Perks across more than 400 campaigns. Lithium Technologies acquired the company for $200 million in 2014 (Klout, Wikipedia).

Then researcher Sean Golliher reverse-engineered the algorithm and found that the simple logarithm of follower count explained 95% of the variance in the score (Klout, Wikipedia). The machine-learning branding sat on top of a number that was almost entirely follower count, and follower count is trivially gameable. President Obama's Klout score sat below those of various bloggers. The service shut down on 25 May 2018.

Goodhart's law describes the failure precisely: a measure that becomes a target stops being a good measure. Klout incentivized score optimization and got score optimization. It never produced a durable signal of the underlying behavior, because it never made anyone better at the thing being scored.

Workplace Gamification Carries Extra Risk

Most gamification research happens in classrooms. Workplaces add a variable classrooms lack: the person being gamified reports to the person running the game.

One of the few studies to examine the downside in an actual work setting is "Uncovering the dark side of gamification at work: Impacts on engagement and well-being" (Hammedi, Leclercq, Poncin and Alkire, Journal of Business Research, 2021). It found gamified work carrying negative effects on employee engagement and well-being together, with one moderator doing the heavy lifting: how willing employees were to take part in the first place. Willingness is what decides which direction the intervention runs.

The employment-relations literature goes further. Research on algorithmic management describes it as a mediation tool that creates asymmetric information in the working relationship, controlling the supply of labour, targeting different workers with variable incentives and mediating disputes at the platform's discretion (Duggan, Sherman, Carbery and McDonnell, Human Resource Management Journal, 2020). The same review points to platforms using algorithmically set surge pricing as an economic nudge to move workers toward demand, a practice Woodcock and Johnson named "gamification-from-above" (Woodcock and Johnson, The Sociological Review, 2018, 66(3), 542-558). Game mechanics layered over work can operate as a control system while presenting themselves as motivation. Anyone deploying points at work inherits that association whether or not they intend it.

The practical implication concerns consent. Gamification people choose to join, with outcomes disconnected from evaluation, sits in a different category from gamification attached to compensation or review. The mechanics can be identical while the effect on the person differs completely.

What the Durable Cases Have in Common

The durability question has an answer, and habit research supplies it.

Lally, van Jaarsveld, Potts and Wardle tracked 96 people performing a daily behavior over 12 weeks and modeled how automaticity develops (Lally et al., European Journal of Social Psychology, 2010, 40(6), 998-1009). Time to reach 95% of an individual's automaticity plateau ranged from 18 to 254 days, with large individual variation. They also found that missing a single occurrence did not materially disrupt habit formation. Accumulated repetition builds durable behavior, and no single reinforcement event does.

Set that against how most corporate gamification is built. Points and badges reward a discrete action once. There is no practice loop underneath them, no accumulation of competence, nothing that gets easier the tenth time than the first. A review of digital behavior-change interventions concluded that most fail to take habitual behaviour into account, leaning on momentary explicit motivation and leaving automaticity undesigned (Pinder, Vermeulen, Cowan and Beale, ACM Transactions on Computer-Human Interaction, 2018). A 2024 systematic review of digital physical-activity interventions found the same pattern persisting, with the most-applied techniques still self-monitoring, goal setting, and prompts and cues (Zhu et al., Journal of Medical Internet Research, 2024).

Badge fatigue follows directly from that design. A badge system has no mechanism for producing lasting change, because the badge is the reward and the reward is the entire loop.

The positive cases look structurally different. Sailer and Homner found that including game fiction and combining competition with collaboration were both particularly effective for behavioral learning outcomes, which together describe a narrative social practice environment. Fotaris et al. ran a gamified computer programming class against a control class, using leaderboards, quizzes and an interactive coding platform together, and recorded average attendance of 78% against 65% and an average final grade of 61% against 53% (Fotaris, Mastoras, Leinfellner and Rosunally, Electronic Journal of e-Learning, 2016, 14(2)), though the authors caveat the grade result on sample size, and one course over one semester says nothing about durability. The commonality across cases that hold up is that people repeatedly do a thing and get better at it, with the game layer organizing the practice.

Self-determination theory supplies the frame. The current state-of-the-art review holds that autonomy, competence and relatedness are the three basic psychological needs (Vansteenkiste, Ryan and Soenens, Motivation and Emotion, 2020). Mechanics that build competence through repeated practice, leave people autonomy over how they engage, and create relatedness through collaboration satisfy all three. Mechanics that dispense tokens disconnected from skill satisfy none of them, which is why they are the ones the novelty-decay and overjustification findings keep catching.

Design choices follow from this. A leaderboard that ranks individuals on output is a token system with a social multiplier attached. A narrative structure that requires a group to coordinate is a practice environment. Those two produce different curves, which is the subject of narrative vs. competitive gamification.

Six Questions to Ask Before You Buy

  1. Does the activity repeat on a fixed schedule? Habit formation needs accumulated occurrences. A tool people use whenever they remember has no accumulation curve.
  2. Does anyone get better at anything? If the tenth session plays exactly like the first, no competence is being built and engagement has nothing to convert into.
  3. Is the reward contingent or informational? "Complete this, receive that" carries the Deci et al. risk. Feedback describing how someone performed does not.
  4. What happens between week 4 and week 10? Ask any vendor for retention data past the decay window Rodrigues et al. documented. A 30-day case study measures the novelty spike and stops.
  5. Can the headline number move without the behavior moving? If yes, you have built Klout. Ask what the score would look like for someone optimizing it deliberately.
  6. Is participation chosen or assigned? Mandated gamification tied to evaluation inherits every finding in the algorithmic-management literature.

Those six questions separate most tools quickly. The ones built around repeated group practice tend to survive them. The ones built around points tend to fail on at least three. For the case on what strong implementations do achieve, see how gamification improves team engagement.

Where QuestWorks Lands

QuestWorks is gamification. Claiming otherwise would be a dodge.

The design puts its weight on repetition, and tokens play almost no part in it. Teams play a fixed weekly 25-minute session on QuestWorks' own cinematic, voice-controlled platform, with Slack or Microsoft Teams handling install, invites and admin. The whole team plays at the same time, seated dynamically into groups of three to six, so a 20-person team runs several simultaneous quests drawn from its own members. That is a recurring practice environment carrying narrative, competition and collaboration in the same session, which covers the elements Sailer and Homner found particularly effective for behavioral learning outcomes.

Participation is voluntary and opt-in, and the data is never tied to performance reviews, which keeps it clear of the algorithmic-management pattern Duggan et al. describe. It also puts willingness, the moderator Hammedi et al. isolated, on the favorable side. What leaders see is a Team Intelligence Score: a 0-100 composite on a weekly dashboard with an eight-week trend line, named sub-scores including Role Clarity and Decision Velocity, and a "Do this week" list of concrete actions. The eight-week window is deliberate, because it runs through the decline phase and into the recovery window the longitudinal research documents instead of stopping at the novelty peak.

Pricing is $199 per month per team with 10 seats included, or $1,990 per year. The 30-day free trial covers four game days and requires no credit card, which is four repetitions, enough to see whether a team shows up a second, third and fourth time.

None of that exempts QuestWorks from the research. The failure modes it documents became the design constraints.

The Takeaway

Gamification works under conditions most implementations do not meet. Contingent rewards attached to already-motivating work reduce motivation, with 128 experiments behind that claim. Novelty decays on a measurable schedule starting around week 4. Scores that can be optimized independently of behavior get optimized independently of behavior. Workplace deployments tied to evaluation carry risks classroom studies never measure.

What survives is repeated practice that builds competence, participation people choose, and feedback carrying information about performance. That is a narrow target, which is why the average measured result is small and the variance around it is large. It is also specific enough to aim at.

Frequently Asked Questions

Sometimes, and the measured effects are modest. The strongest meta-analysis available (Sailer and Homner, Educational Psychology Review, 2020) found effect sizes of g = .49 on cognitive outcomes, g = .36 on motivational outcomes and g = .25 on behavioral outcomes, with substantial variation between studies and the factors behind successful gamification described by the authors as still somewhat unresolved. The most-cited review in the field concluded that effects depend heavily on implementation context and on the users involved.

A 14-week study of 756 students found engagement following a U-shaped curve: initial gains declined starting around week 4, the dip persisted for two to six weeks, and engagement partially recovered somewhere between weeks six and ten as familiarity took over from novelty (Rodrigues et al., 2022). Systems with no repeated practice underneath them have nothing to recover into, so the decline continues.

Yes. A meta-analysis of 128 experiments found that engagement-contingent, completion-contingent and performance-contingent rewards significantly undermined free-choice intrinsic motivation, with effect sizes between -0.28 and -0.40 (Deci, Koestner and Ryan, Psychological Bulletin, 1999). The same analysis found that verbal praise and positive feedback increased both engagement and interest, so the harm is specific to reward structures of the form do X, receive Y.

Designs built around repeated practice that develops competence. Habit research found that reaching 95% of an individual's automaticity plateau took between 18 and 254 days of accumulated repetition, and that missing one occurrence does not disrupt the process (Lally et al., 2010). Meta-analytic evidence also found that including game fiction and combining competition with collaboration were both particularly effective for behavioral learning outcomes.

It depends on whether participation is chosen and whether outcomes touch evaluation. Research in the Journal of Business Research (2021) found gamified work carrying negative effects on employee engagement and well-being, moderated by how willing employees were to take part, and the employment-relations literature treats algorithmic management as a tool that creates asymmetric information in the working relationship. Voluntary participation with no link to performance reviews avoids most of that exposure.

Teams that quest just work.

A game your team will love to play, and a weekly read on how they perform under pressure. 30-day free trial, four game days, no credit card.

Slack Microsoft Teams Try it free
Team Intelligence™, powered by play. Slack Microsoft Teams Try QuestWorks Free