← Back to list

The Measurement Trap

Why most AI ROI frameworks measure the wrong things

Ian Loe · 2026-05-15 14:02 · 6 claps · 6.6 min read
#ai-roi #ai-governance #roi #agentic-ai
Open on Medium ↗
Wiki topics: AGT · AI Agents GEN · Genomics & Sequencing

The Measurement Trap

Why most AI ROI frameworks measure the wrong things

The boardroom wants a number. Finance wants a payback period. The CEO wants to know if the investment is working. So you build a dashboard. And the dashboard lies to you.

Not because anyone fabricated the data. But because the framework used to collect it was designed to answer the wrong question.

This is the measurement trap. And right now, organisations across every sector are falling into it.

The pressure to prove AI value has never been higher. But the harder organisations push for a clean ROI number, the more likely they are to reach for frameworks that were built for software procurement, not for a fundamentally different class of capability. The result is ROI dashboards that look reassuring and tell you almost nothing useful.

There are three structural flaws in how most organisations approach AI measurement. Each one is individually misleading. Together, they create a compounding illusion of understanding.

Flaw One: Measuring Activity, Not Outcomes

Ask most organisations how they are tracking AI performance and you will hear a familiar list: prompts processed, documents generated, tickets deflected, hours notionally saved. These are activity metrics. They tell you what the system did. They say almost nothing about whether it created value.

Consider a common deployment: an AI assistant integrated into a customer service operation, tasked with deflecting inbound enquiries. After three months, the dashboard shows a 28% deflection rate. The team celebrates. But underneath that number, a different picture is emerging. Customer satisfaction has flatlined, escalation rates on the cases that do reach human agents have risen, and the agents themselves are spending more time correcting AI errors than they saved by not handling the deflected queries.

An activity metric dressed up as an outcome metric is not a measurement. It is a story you are telling yourself.

The deflection rate measured what the AI did. It did not measure what the organisation needed: better customer outcomes, reduced operational cost, or freed capacity for higher-value work. These outcomes require deliberate definition before deployment, not post-hoc rationalisation after the dashboard is built.

This is not a new problem in technology measurement. But AI amplifies it because the outputs are so visible and so easily quantifiable. A language model produces text. Text is countable. Counting text feels like measurement. It is not.

FRAME note: The Metrics dimension within the FRAME governance framework is explicit on this point: measurement must be anchored to the domain outcome the AI system is deployed to serve, not to the throughput of the system itself. Activity without outcome alignment is noise with a confidence interval. https://github.com/ianloe/FRAME

Flaw Two: Measuring Too Early

The second flaw is temporal. Most AI ROI assessments are conducted within the first 90 to 180 days of deployment. This is also the period when almost every AI initiative looks weakest.

In those early months, the organisation is paying full cost: licensing, integration, change management, training, and the hidden cost of disruption to existing workflows. But it is capturing only a fraction of the benefit, because the benefit of AI deployment is not linear. It compounds.

A team that has genuinely integrated an AI capability into its workflow does not merely save time proportional to the number of tasks automated. Over time, it changes how it approaches problems, what it attempts, how it allocates human attention. The productivity dividend from that shift is not visible at month three. It may not be measurable at month six. Organisations that run their ROI assessment at month four and declare the initiative underperforming are not measuring the AI. They are measuring their own impatience.

The compounding curve does not care about your quarterly review cycle.

There is also a capability maturation effect. AI systems, particularly those involving agents or retrieval-augmented generation, tend to improve as they process more organisational data, as prompts are refined, and as workflows are adjusted around their actual behaviour rather than their assumed behaviour. The version of the system you are measuring at month three is not the version that will be running at month twelve.

This has a practical implication: organisations that measure too early often pull back from deployments that would have compounded into significant value. They do not just get the measurement wrong. They make the wrong decision based on it.

AAEF note: The AI Agent Evaluation Framework addresses this directly through structured appraisal cadences. Just as you would not evaluate a new employee’s full contribution after their first month and declare them a poor hire, AI agents in production require evaluation frameworks that account for capability maturation, contextual learning, and the time required for genuine workflow integration. The appraisal cadence is not bureaucracy. It is a guard against the compounding curve penalty. https://github.com/ianloe/AAEF

Flaw Three: Ignoring Adoption Drag

The third flaw is the most undercosted item in almost every AI business case: organisational change.

Most ROI frameworks treat deployment as the finish line. The system is live, the licences are paid, the integration is complete. From this point, the model assumes value accrual begins. But value does not accrue from a system being deployed. It accrues from a system being used, used correctly, and used consistently.

The reality in most enterprise AI deployments is that adoption is partial, uneven, and often reluctant. A tool that 40% of the relevant staff use consistently, 30% use sporadically, and 30% avoid where possible does not deliver 40% of its potential value. It delivers significantly less, because inconsistent adoption creates its own category of problems: process fragmentation, quality variance, the burden of managing both AI-assisted and non-AI-assisted outputs simultaneously, and the cognitive overhead of staff who are uncertain which mode of working applies in which situation.

Partial adoption is not a partial win. It is a full cost with a fractional return.

Adoption drag also interacts badly with the time horizon problem. An organisation that measures early, sees underwhelming numbers, and concludes the tool is underperforming may in fact be looking at an adoption problem, not a capability problem. The tool works. The organisation has not yet changed around it. These are different diagnoses requiring different responses, and a measurement framework that cannot distinguish between them is not fit for purpose.

The cost of change management, training, workflow redesign, and sustained adoption support is rarely modelled in full in the original business case. When it does appear, it tends to be front-loaded as a one-time deployment cost, rather than recognised as an ongoing investment that determines whether the compounding curve is ever reached.

FRAME note: The Evaluation dimension within FRAME addresses this through ongoing governance that includes adoption tracking as a first-class metric alongside performance metrics. Governance without adoption visibility is incomplete governance. You cannot evaluate an AI initiative’s contribution to organisational outcomes if you do not know whether, and how, the organisation is actually using it.

What a Better Measurement Model Looks Like

The three flaws are not unrelated. They compound. An organisation that measures activity rather than outcomes, measures too early, and ignores adoption drag is not just getting the number wrong three times. It is systematically constructing a picture of AI performance that has no reliable relationship to actual value creation.

A better measurement model requires three corresponding corrections.

Define outcomes before deployment, not after

Every AI initiative should begin with an explicit answer to the question: what changes in our operation, for our customers, or for our staff, if this works? That answer defines what gets measured. If it cannot be answered before deployment, the initiative is not ready for deployment. A governance framework that includes outcome definition as a gate condition forces this discipline.

Build measurement cadences that respect the compounding curve

This means distinguishing between leading indicators, which are measurable early and signal whether conditions for value creation are in place, and lagging indicators, which are the actual outcome measures and require time to accumulate. A 90-day review should be assessing adoption rates, workflow integration depth, and system performance quality, not expecting to see the full financial return on a multi-year capability investment.

Model adoption explicitly, and put it in the denominator

The ROI model should reflect realistic adoption trajectories, not theoretical full deployment. If the current adoption rate is 55%, the expected value capture at current adoption should be modelled accordingly, alongside a plan for moving that number. Adoption is not a soft factor. It is a financial variable.

Taken together, these corrections do not make AI ROI harder to demonstrate. They make it more honest, and more defensible, when the conversation happens at board level.

The Dashboard Will Keep Lying Until You Change the Question

The pressure to prove AI value is legitimate. Boards are right to ask. CFOs are right to push. The investment is real, the risks are real, and the opportunity cost of misallocation is real.

But the frameworks most organisations reach for in response to that pressure were built for a different kind of technology investment. They were designed to evaluate software tools with linear value curves, predictable adoption profiles, and stable capability sets. AI in general, and agentic AI in particular, fits none of those assumptions.

The measurement trap is not a data problem. It is not a tooling problem. It is a mental model problem. And until the mental model changes, the dashboard will keep producing numbers that feel reassuring and say almost nothing true.

The good news is that the trap is visible once you know where to look. Measure outcomes, not activity. Respect the compounding curve. Put adoption in the denominator. These are not complicated principles. But they require deliberate governance to enforce, especially in organisations where the pressure to show a number is stronger than the discipline to show the right number.

That governance discipline is exactly what frameworks like FRAME and AAEF are designed to support. Not as theoretical constructs, but as operational tools for organisations that are serious about AI deployment and serious about knowing whether it is actually working.

To learn more about FRAME and AAEF, or to discuss how they apply to your organisation, you can reach me via https://ianloetech.net


메타데이터
post_id
fd50b67e9af4
slug
the-measurement-trap-fd50b67e9af4
url
https://medium.com/@ianloe/the-measurement-trap-fd50b67e9af4
canonical_url
https://medium.com/@ianloe/the-measurement-trap-fd50b67e9af4
author_url
https://medium.com/@ianloe
status
ok
fetched_at
2026-06-09 14:34:10