← Back to list

How to Measure AI Leverage Without Lying to Yourself

Here’s the scoreboard — including the impressive number we could have led with and didn’t, and why the boring measured ones matter more.

Dian Balta · 2026-06-27 13:03 · 0 claps · 3.6 min read
#software-engineering #developer-productivity #engineering-management #artificial-intelligence #data-science
Open on Medium ↗
Wiki topics: RAG · RAG & Retrieval ML · Machine Learning AI · AI · General BIZ · Business Strategy 🔬 · Science · General ⏱️ · Productivity 📰 · Journalism & News

How to Measure AI Leverage Without Lying to Yourself

Here’s the scoreboard — including the impressive number we could have led with and didn’t, and why the boring measured ones matter more.

The Operating Model · Part 6 of 7

We’ve spent five posts claiming an operating model turns AI from a tax into leverage. This is where we owe you evidence — and the honest move is to say what each number is worth, because the fastest way to lose a technical reader is one clean, flattering figure.

Lead with what stands, not how fast it moved

The load-bearing evidence isn’t velocity; it’s what exists now:

  • roughly 40 services across 6 tenant overlays
  • 210 recorded architectural decisions
  • about 5,400 structural-invariant test functions
  • around 1,500 curated knowledge entries
  • 43 governance gates (a count, not a coverage measure)
  • on the order of 2.2 million tracked lines
  • about 3,100 changes merged to the main branch, each reviewed

These are direct counts of a standing system — evidence the work is durable, which is not yet evidence of what produced it.

“It’s just churn” — answered with the distribution

The reflexive objection to any AI output claim is that it’s noise: a thousand trivial commits dressed up as progress. So here’s the anatomy of a commit, as a distribution rather than a flattering anecdote. The median commit is 80 lines across 2 files. Only 6.1% are two lines or fewer. About 59% are larger than 50 lines. Around 97% are conventionally typed. And there are 0.11 deletions for every insertion.

That last ratio is the one we’d defend. Low deletions-per-insertion means the codebase is accumulating, not thrashing — it isn’t endlessly rewriting the same lines to look busy. You can argue with what the work is worth, but you can’t call this churn with the numbers in front of you.

The number we didn’t lead with

Now the soft ones — each with its caveat in the same breath, because that’s the entire point of this post.

Cadence. At peak, authored commits ran roughly 150 times the eight-month baseline — but that’s authored commits with merges excluded, and we treat it as a secondary signal. It’s the easiest number to inflate and the easiest to misread, so we don’t lean on it.

Leverage. Our own analytics put leverage — output per engineer-hour, measured against our own pre-AI baseline — at about 22–28×. We could have made that the headline. We won’t, because it’s a label-imputed proxy: roughly 99% of the effort base is estimated from size labels, and only about 1% is backed by recorded time. It’s indicative, not measured. We’re telling you this because the platform’s instrumentation surfaced the imputation itself — it wouldn’t let us present an estimate dressed as an observation.

Rework. A small story that taught us to trust this system. An early dashboard reported a 17.4% “rework” rate, which sounds alarming. On inspection, that figure was the share of work items titled “fix” — fix-naming prevalence, not actual defect rework. True rework, where a later change corrects an earlier one, was closer to 3.5%. The honest report is the range, 3.5% to 17.9%, with no claimed trend — because the instrumentation refused to collapse two different things into one clean number.

Cost. A ceiling, self-reported and not independently audited: the whole production cadence runs inside three AI-coding subscriptions (and a LLM run on-prem for some basic tasks) plus a few thousand euros of one-off API spend during the transition, with essentially nothing metered since. The low, model-agnostic cost is itself a point for the thesis — if raw model capacity or spend were doing the work, the bill wouldn’t be this small.

The investment. Those two figures answer the buyer’s question. The build was front-loaded — call it the better part of a year’s platform work, which we can’t cleanly separate from the rest — and the run is the cheap, model-agnostic number above plus one person’s part-time stewardship to keep the memory and tests honest. Expensive to stand up, cheap to operate. We won’t hand you a clean build figure by inventing one.

The caveat we have to put on the caveats

Even the clean, measured numbers are not proof of the thesis. They’re consistent with it. This is one site, observed over time, not a controlled comparison — the durable output is real, but it can’t on its own rule out that we’d have built something similar by other means. We believe the operating model is why it was cheap and fast. We can’t prove that from inside our own project, and we’re not going to pretend the scoreboard does.

A number you can’t caveat is a number you shouldn’t publish. The caveats aren’t the fine print here. They’re the finding.

If your reaction to the 22–28× is “that’s soft,” you’ve read it exactly right — which is why we told you before you had to ask. The last post asks the only question that’s left: does any of this transfer to you?

— -

The Operating Model · Part 6 of 7. ← Previous: “How You Run AI Where Auditors Are Watching” · Next → “Does Any of This Transfer to Your Team?”.

Dian Balta leads the Platform Engineering competence field at fortiss, a research institute in Munich. This series is the public companion to a forthcoming fortiss whitepaper.


메타데이터
post_id
2ea7651b6c15
slug
how-to-measure-ai-leverage-without-lying-to-yourself-2ea7651b6c15
url
https://medium.com/@dian-balta/how-to-measure-ai-leverage-without-lying-to-yourself-2ea7651b6c15
canonical_url
https://medium.com/@dian-balta/how-to-measure-ai-leverage-without-lying-to-yourself-2ea7651b6c15
author_url
https://medium.com/@dian-balta
status
ok
fetched_at
2026-06-29 01:02:39