← Back to list

Cheaters Win Early. Honest Workers Win Later. Here’s the Proof.

A May 2026 paper shows that multi-stage evaluation systems can make long-run honesty the dominant strategy — without punishing gamers at…

Shugo Nozaki · 2026-05-09 13:01 · 21 claps · 5.3 min read
#artificial-intelligence #machine-learning #self-improvement #future-of-work #technology
Open on Medium ↗
Wiki topics: EVAL · Evaluation & Benchmarks ML · Machine Learning AI · AI · General PSY · Psychology EDU · Education & Learning 📐 · Mathematics 🚀 · Self Improvement

Cheaters Win Early. Honest Workers Win Later. Here’s the Proof.

A May 2026 paper shows that multi-stage evaluation systems can make long-run honesty the dominant strategy — without punishing gamers at all.

Everyone has seen this. Someone polishes their LinkedIn to pass an automated screener instead of learning the underlying skill. Someone games their credit score behaviors without building actual financial stability. A student memorizes the exam format instead of the subject matter.

Short term, this is rational. Gaming observable metrics is cheaper than building the real thing. The system can’t tell the difference, so the gamer passes. Punishing this typically requires surveillance, audits, or external incentive structures — all expensive, all imperfect.

A team at the University of Michigan just proved something different. Ziyuan Huang, Lina Alkarmi, and Mingyan Liu published a paper in May 2026 (arXiv:2605.04202) showing that with the right multi-stage design, honest effort dominates gaming in the long run — and the system doesn’t need to penalize cheating to make this happen. The design does the work.

Two kinds of agents

The model formalizes what feel like familiar intuitions.

Improvement — practicing, studying, actually developing the underlying skill. The observable score goes up and so does real ability. Expensive.

Gaming — manipulating observable signals without touching real ability. The score goes up; the skill doesn’t. Cheaper.

The paper accepts, explicitly, that gaming costs less (c−<c+c^- < c^+ c−<c+). In a single-shot evaluation — one job application, one loan decision, one exam — a rational agent often games. That isn’t a conjecture. It’s just looking around.

There’s one more piece: skill depreciation. Ability decays without maintenance (γ<1\gamma < 1 γ<1). Stop practicing and you lose ground. This, too, is just true.

With these two premises, the paper asks: what happens over many rounds?

The multi-stage setup

Most previous research studied either a single evaluation or repeated identical tests. This paper models something closer to how real institutions work: a sequence of levels, each harder than the last, each with higher rewards. Probationary to associate to senior. Freshman to senior. Entry-level credit to premium.

At each level, the classifier doesn’t just accept or reject. It has three possible outputs:

  • Accept (+1): promotes the agent to the next level
  • Reject (−1): demotes the agent one level down
  • Abstain (0): no verdict, agent stays put

The abstain option is doing more work than it first appears. A classifier that withholds judgment when uncertain — rather than forcing a binary call — gives patient, honest agents somewhere to wait. You don’t have to act under pressure with incomplete preparation. You can hold your position until you’re ready to move up.

This is what makes gradual improvement a viable strategy rather than a naive one.

The long-run result

The paper analyzes two pure strategies against each other:

NI (No Improvement): always game, never actually develop. NG (No Gaming): always improve honestly, never game.

Theorem 4.1 shows where each type ends up when the system runs long enough:

  • NI agents cluster around some intermediate level and stall. They can’t pass the threshold above because their actual skill has been decaying the whole time. Gaming moves the observable score; it can’t manufacture the underlying ability the upper levels require.
  • NG agents cluster near the top. Accumulated ability eventually satisfies what the upper classifiers require.

Theorem 4.2 makes the underlying mechanism precise: in the long-run average, the NI agent’s true skill converges to zero. The NG agent’s converges to a positive value.

The gamer — who starts with an advantage, who passes early rounds cheaply, who looks better on paper in round one — ends up at zero. The honest worker ends up at the top.

No external penalty. No bonus for disclosed ability. Just the geometry of the system.

I find this genuinely surprising. The result feels like it should require some detection mechanism, some cost imposed on detected gamers. The proof says it doesn’t.

What makes it work: incremental thresholding

The result doesn’t hold for arbitrary systems. It requires a specific design principle the paper calls incremental thresholding — three conditions on how difficulty escalates:

  1. The first threshold is non-trivial (there’s something to work toward from the start)
  2. Thresholds increase gradually between levels
  3. The inter-level spacing preserves motivation even accounting for skill decay

Theorem 4.3 gives a concrete bound: when the increment Δμ\Delta\mu Δμ between levels is small enough, the NG agent’s long-run utility strictly exceeds the NI agent’s. When increments are too large — when the next level suddenly demands substantially more — gaming becomes rational again.

This is the finding with the most direct practical implication. Raise the bar too fast, and you incentivize shortcuts. Raise it steadily, and you reward actual development. The shape of the difficulty curve isn’t just a pedagogical choice — it determines which strategy is economically dominant.

What this doesn’t prove

The model is clean in ways real systems aren’t. Classifiers have errors. Agents have incomplete information about what’s being measured. Real populations mix many strategies rather than just two pure types.

Whether actual hiring pipelines, credit systems, or admissions processes are close enough to this model for the theorems to transfer is a separate question — and one the math doesn’t answer. The simulation results are encouraging: qualitative findings hold across parameter variations and under reinforcement learning when agents optimize freely. But the gap between theorem and deployed system always requires judgment.

The abstain mechanism also depends on classifiers that know when they don’t know. That’s a real engineering requirement. Many deployed systems don’t have it.

Why it matters anyway

The conventional framing for gaming is moral or cultural: people cheat because they lack integrity, and fixing it requires enforcement. That framing leads to surveillance budgets, fraud detection models, verification overhead.

This paper suggests a different question: if your system rewards gaming, what does the design look like that stops it? Not through detection or punishment — but by making honest development the higher-value strategy in expectation.

Gradual escalation. An abstain option when confidence is low. Those two structural features, the paper shows, are sufficient to flip the long-run incentives.

Whether that’s achievable in any specific real-world system is a design problem. The theorem doesn’t solve it.

But it’s worth knowing: the math, at least, is on the side of honest effort.

References

Primary

Ziyuan Huang, Lina Alkarmi, Mingyan Liu — “Sequential Strategic Classification with Multi-Stage Selective Classifiers,” arXiv:2605.04202 [cs.LG], May 2026. https://arxiv.org/abs/2605.04202

Strategic Classification — Foundations

M. Hardt, N. Megiddo, C. Papadimitriou, M. Wootters — “Strategic Classification,” arXiv:1506.06980, 2015. https://arxiv.org/abs/1506.06980

J. Miller, S. Milli, M. Hardt — “Strategic Classification is Causal Modeling in Disguise,” arXiv:1910.10362, 2020. https://arxiv.org/abs/1910.10362

S. Milli, J. Miller, A. D. Dragan, M. Hardt — “The Social Cost of Strategic Classification,” FAT* 2019. https://arxiv.org/abs/1808.08460

Gaming and Improvement

J. Kleinberg, M. Raghavan — “How Do Classifiers Induce Agents To Invest Effort Strategically?” arXiv:1807.05307, 2019. https://arxiv.org/abs/1807.05307

S. Ahmadi, H. Beyhaghi, A. Blum, K. Naggita — “On Classification of Strategic Agents who can both Game and Improve,” arXiv:2203.00124, 2022. https://arxiv.org/abs/2203.00124

K. Jin, X. Zhang, M. M. Khalili, P. Naghizadeh, M. Liu — “Incentive Mechanisms for Strategic Classification and Regression Problems,” EC 2022. https://arxiv.org/abs/2206.10138

Repeated and Long-Run Settings

J. C. Perdomo, T. Zrnic, C. Mendler-Dünner, M. Hardt — “Performative Prediction,” arXiv:2002.06673, 2021. https://arxiv.org/abs/2002.06673

T. Zrnic, E. Mazumdar, S. S. Sastry, M. I. Jordan — “Who Leads and Who Follows in Strategic Classification?” arXiv:2106.12529, 2022. https://arxiv.org/abs/2106.12529

X. Zhang, M. Khaliligarekani, C. Tekin, M. Liu — “Group Retention when Using Machine Learning in Sequential Decision Making,” NeurIPS 2019. https://arxiv.org/abs/1907.10614

Selective Classification (Abstain Option)

Y. Geifman, R. El-Yaniv — “Selective Classification for Deep Neural Networks,” arXiv:1705.08500, 2017. https://arxiv.org/abs/1705.08500

C. Cortes, G. DeSalvo, M. Mohri — “Learning with Rejection,” ALT 2016. https://link.springer.com/chapter/10.1007/978-3-319-46379-7_10

R. El-Yaniv, Y. Wiener — “On the Foundations of Noise-free Selective Classification,” JMLR 2010. https://jmlr.org/papers/v11/el-yaniv10a.html

Skill Depreciation and Dynamic Models

T. Dohmen, A. Trivedi — “Reinforcement Learning with Depreciating Assets,” arXiv:2302.14176, 2023. https://arxiv.org/abs/2302.14176

J. Alston — “Dynamics in the Creation and Depreciation of Knowledge, and the Returns to Research,” NBER 1998. https://www.nber.org/papers/w6763


메타데이터
post_id
136d68c13b50
slug
cheaters-win-early-honest-workers-win-later-heres-the-proof-136d68c13b50
url
https://medium.com/@shugo/cheaters-win-early-honest-workers-win-later-heres-the-proof-136d68c13b50
canonical_url
https://medium.com/@shugo/cheaters-win-early-honest-workers-win-later-heres-the-proof-136d68c13b50
author_url
https://medium.com/@shugo
status
ok
fetched_at
2026-07-10 18:30:51