← Back to list

Part 2: Software Productivity — What Jeff Sutherland’s Coaches Actually Did

The interventions Sutherland’s coaches used to manufacture hyperproductivity — and the pattern language that reproduces it.

Joakim Marner · 2026-07-15 21:15 · 1 claps · 7.2 min read paywalled
#scrum #agile #metrics #productivity #engineering-mangement
Open on Medium ↗
Wiki topics: 📋 · Product Management ⏱️ · Productivity

Part 2: Software Productivity — What Jeff Sutherland’s Coaches Actually Did

The interventions Sutherland’s coaches used to manufacture hyperproductivity — and the pattern language that reproduces it.

Part 2 of 5 in a series on software productivity — ending with how Agentic AI changes everything.

In Part 1, we looked at what Jeff Sutherland actually measured — the fighter-aircraft dashboard of nine gauges, read team-relative every sprint, that turns “how fast is this team” into an honest, diagnostic question. But a dashboard only tells you a team is slow. It doesn’t make it fast.

Measurement is only half the program. The papers that made Sutherland’s reputation are the ones that describe interventions — the specific, repeatable moves coaches used to push a team in the right direction as measured by the metrics. Read together, they’re less a theory than a playbook. Four cases carry most of the weight, and they all run the same four plays.

What the coaches actually did

Four peer-reviewed cases carry most of the weight. Each pairs a single coaching move with the numbers it produced, measured against the team’s own baseline. They run in escalating order — from bootstrapping one team to scaling the same disciplines across hundreds.

1. Impose the rules before the team can self-organize

Shock Therapy — bootstrapping a team to 400%

Sutherland, Downey & Granvik · MySpace and Jayway · Agile 2009

Many teams first adopting Scrum “under-emphasized or failed to implement critical elements of the Scrum Framework, which sets them up for limited success at best.” To address this, the prescriptive core of the whole Shock Therapy program was the most counterintuitive. Rather than let a new team self-organize prematurely, the coach imposed strict, non-negotiable rules up front: size every story against a shared reference story so points meant the same thing team to team; break stories down until each was small and independently completable inside a sprint; swarm — the whole team finishes one item before starting the next, killing the parallel-work-in-progress habit; and enforce a hard Definition of Done so “done” carried no hidden remaining work. The scaffolding stayed locked until a team met a concrete bar: three consecutive successful sprints and a 240% velocity increase. Only then could it start relaxing the “Default Profile” (never the Scrum framework itself).

The enforced target was a 240% improvement within a few weeks; teams that cleared it typically ran on past 400% of waterfall velocity — into the hyper-productive state — in later sprints. Heavy up-front prescription got them there faster than open-ended coaching. The rules were temporary scaffolding, removed once the team’s own rhythm could hold the pace

2. Deliberately under-commit so the team finishes early

Teams That Finish Early Accelerate Faster — the flywheel

Sutherland, Harrison & Riddle · a pattern language from many teams · HICSS, 2014

The structural companion to Shock Therapy: it explains why the discipline works. The key move is deliberately having teams take on slightly less than full capacity so a sprint reliably finishes early instead of spilling over, then planning the next sprint from the velocity actually achieved (the “Yesterday’s Weather” pattern) rather than an aspirational number. The slack created by finishing early is spent attacking impediments and interruptions — not pulling more scope forward.

Finishing early becomes a flywheel, not slack to be filled. Teams that stop over-committing build a buffer, spend it removing friction, and accelerate sprint over sprint. The paper’s flagship case rode this loop to a 300% velocity gain by Sprint 91 and roughly 1200% by Sprint 211 — described as the first documented, sustainable, hyper-productive company. It codifies the organizational patterns — small size, swarming, realistic planning — that reproduce the effect rather than leaving it to luck.

3. Layer Scrum on a rigorously measured baseline

Scrum on CMMI Level 5 — the most defensible numbers

Sutherland, Jakobsen & Johnson · Systematic Software Engineering, Denmark · Agile 2007

Systematic was a Danish software company serving defense, healthcare, and other industries, certified at CMMI Level 5 — the top rung of the Capability Maturity Model Integration, a five-level scale of how disciplined and measured an organization’s processes are — meaning it already measured everything with unusual care. Coaches layered Scrum on top of that disciplined baseline rather than replacing it, piloting bi-weekly deliveries with tight customer feedback and a story-based early-testing approach in which each story’s test was defined before any code was written. Because the CMMI data infrastructure already existed, they could measure before-and-after on the same teams.

Large projects showed roughly double the productivity (a 201% increase versus baseline) and early testing cut coding defects found in final test by about 40% (38% and 42% on two projects). Because Systematic already measured to CMMI-5 standards, this is the most defensible quantified proof in the set — and it showed Scrum and heavyweight process maturity are complementary, not opposed.

4. Enforce entry and exit criteria at scale

Good to Great — the same levers, at scale

Jakobsen & Sutherland · Systematic, expanded dataset · Agile 2009

The follow-up scaled the ready-ready / done-done discipline across hundreds of teams over multiple years (2006–2009). Making “ready-ready” an enforced entry criterion kept sprints from absorbing under-specified work; making “done-done” an enforced exit criterion eliminated the phantom-progress that inflates velocity. Crucially, the paper ties the measured gains to those specific, nameable practices rather than to Scrum in the abstract.

The gains held at scale and over time, tied to disciplines that could be taught rather than merely admired.

Two business-outcome stories round out the 2014 book: PatientKeeper, the healthcare software firm where Sutherland was CTO, shipping 45 releases a year; and the FBI’s Sentinel case-management system, delivered by a small Scrum team after a large waterfall predecessor had failed expensively. These are coarser than the peer-reviewed cases — release frequency and delivery-against-budget rather than function points — but they answer the question executives actually ask.

The through-line. Every case runs the same play: prescribe a few hard rules, measure against the team’s own baseline, spend early-finish slack on impediments and automation, and only relax the scaffolding once the team’s rhythm can hold the pace. The numbers vary; the levers don’t.

A pattern language for high-performing Scrum teams

Across these interventions and studies, Jeff Sutherland identifies nine foundational patterns that, when adopted and gradually perfected, help bootstrap a team toward high performance:

  • Stable Teams — keep the same 3–9 members together; chemistry and shared habits take time to form, so don’t shuffle people between teams.
  • Swarming — the whole team finishes the highest-priority story before starting the next, minimizing multitasking and context-switching.
  • Yesterday’s Weather — plan the next sprint from what the team actually completed last sprint, not from hope.
  • Interrupt Pattern (Illegitimus Non Interruptus) — don’t block interruptions, budget for them: reserve a buffer sized from history (e.g. ~30% of velocity), route every incoming request through the Product Owner for triage, and abort-and-replan only if the buffer overflows.
  • Happiness Metric — measure morale every sprint; unhappiness precedes falling velocity, so fix the root cause early.
  • Scrumming the Scrum — a continuous retrospective that commits at least one major process improvement every sprint.
  • Teams that Finish Early Accelerate Faster — a team that reliably finishes early spends the slack on innovation, refactoring, and paying down technical debt.
  • Daily Clean Code — fix bugs and keep the code deployable every day so technical debt never accumulates.
  • Emergency Procedure (Stop the Line) — when it’s clear mid-sprint that the Sprint Goal is at risk, the team doesn’t coast into failure. Borrowing a fighter-pilot discipline, it runs a fixed sequence — change how the work is done, get help by offloading backlog, reduce scope, abort and replan the sprint, and inform management of release impact — doing only as much as needed. The point is to make problems visible early (Toyota’s “stop the line”) and take corrective action.

Sutherland also points to the shu-ha-ri (守破離) nature of Scrum: shu, obey the form; ha, break from it; ri, transcend it. Mechanically running the 3–5–3 — the framework’s 3 roles (Product Owner, Scrum Master, Developers), 5 events (the Sprint itself plus Sprint Planning, the Daily Scrum, the Sprint Review, and the Sprint Retrospective), and 3 artifacts (Product Backlog, Sprint Backlog, and Increment) — is only shu. A high-performing team must move beyond mediocre, non-reflective going-through-the-motions Scrum — adopting and perfecting these patterns is how it does so.

Why this matters

Put the two halves together and the origin story loses its mystique. Part 1’s dashboard measures a team against itself; Part 2’s interventions move those instruments on purpose. Neither half is magic, and neither is heroics. Hyperproductivity was the result of a handful of imposed rules, initial coaching on mastery of practices following the rules, a light-but-honest instrument panel, and the discipline to relax the scaffolding only once the team could hold the pace on its own. The numbers varied from case to case — 200%, 400%, 1200% — but the levers never did.

That’s also the limit of the whole program. Every number here measures a team against its own baseline, because no two teams’ story points are comparable (the one excetioon is flow efficiency). It’s the only honest yardstick Sutherland had — and the decade that followed asked a bigger question: could productivity be measured across thousands of organizations, rigorously, without destroying the teams being measured? The answer became DORA, Flow metrics, and the SPACE framework.

Next in the series → Part 3: Beyond Velocity — How DORA, Flow, and SPACE Measure What Actually Matters.

Read more


메타데이터
post_id
d068d4f4b988
slug
software-productivity-what-did-jeff-sutherlands-coaches-actually-do-d068d4f4b988
url
https://medium.com/@joakim.marner/software-productivity-what-did-jeff-sutherlands-coaches-actually-do-d068d4f4b988
canonical_url
https://medium.com/@joakim.marner/software-productivity-what-did-jeff-sutherlands-coaches-actually-do-d068d4f4b988
author_url
https://medium.com/@joakim.marner
status
ok
fetched_at
2026-07-17 17:40:40