The North Star Metric Problem in AI Products
Most AI products don’t fail because their models are weak. They fail because they optimise the wrong metric.
The North Star Metric Problem in AI Products

Most AI products don’t fail because their models are weak. They fail because they optimise the wrong metric.
Specifically, they optimise interaction instead of impact.
This mistake shows up everywhere: teams tracking prompts per session, tokens generated, feature usage, or session time as proxies for value. These numbers move easily. They also say very little about whether the product changed what the user actually did.
The North Star Metric problem in AI products is not measurement difficulty. It is value misidentification.
Problem: AI Teams Borrow SaaS Metrics for Non-SaaS Behaviour
Traditional SaaS metrics assume the product executes tasks.
AI products change how users make decisions.
That difference breaks most inherited measurement logic.
Consider three common “North Star” candidates used in early AI products:
queries per user time spent with the assistant outputs generated per session
Each captures activity. None captures usefulness.
A developer can accept zero Copilot suggestions and still generate hundreds of prompts. A recruiter can regenerate candidate summaries repeatedly without trusting any of them. A marketer can iterate copy drafts for twenty minutes and still write the final version manually.
If the metric increases while behaviour does not change, the product is not improving.
It is only getting noisier.
User: The AI User Is Trying to Reduce Uncertainty, Not Complete Workflow Steps
A CRM user logs pipeline updates.
A Figma user exports a design.
An AI user evaluates whether the system’s output is trustworthy enough to act on.
That evaluation step is the product.
Most AI workflows look like execution flows on the surface. Underneath, they are confidence negotiations between human judgment and machine suggestion.
When a recruiter accepts an AI shortlist without re-ranking it, uncertainty drops.
When a developer keeps Copilot’s suggestion without rewriting it, uncertainty drops.
When a support agent sends an AI draft reply unchanged, uncertainty drops.
That is where value appears.
Insight: The Real Output of an AI Product Is Behaviour Change After Generation
Generation is not the product.
Adoption of the generation is the product.
This distinction sounds small, but it changes how the entire metric system should be designed.
Take GitHub Copilot. If the team had optimised for suggestions shown, they would have maximised verbosity. Instead, internal reporting has emphasised acceptance rate and developer productivity improvements, including evidence from Microsoft’s 2023 research showing measurable reductions in task completion time.
Copilot succeeds when developers keep the code.
Not when they see it.
The same logic applies across categories.
Tradeoff: The Metrics That Capture AI Value Are Harder to Instrument
There is a structural reason teams fall back to prompt counts and token usage.
They are cheap to measure.
Confidence is expensive to measure.
Acceptance requires edit-distance tracking. Decision acceleration requires workflow instrumentation. Productivity lift requires controlled comparisons or cohort analysis. These are slower to implement and harder to explain internally.
So teams optimise what dashboards already expose.
The tradeoff is operational convenience versus measurement accuracy.
Most teams choose convenience without noticing the cost.
Decision: A Strong AI North Star Metric Measures Action Taken After Output
Instead of tracking generation volume, strong AI products track what survives contact with the user.
Examples:
In an AI coding assistant, measure suggestion survival into commit history.
In an AI writing assistant, measure draft acceptance without structural rewrite.
In an AI recruiter copilot, measure shortlist-to-interview conversion without manual reranking.
In an AI support assistant, measure response-send rate without edits.
Each captures whether the system influenced a real decision.
That is the behaviour layer where value exists.
Metric: Acceptance Rate Is the Closest Thing AI Products Have to a Universal North Star
Across categories, one metric pattern appears repeatedly useful:
accepted outputs per active user
It works because it encodes three signals simultaneously:
The system generated something The user evaluated it The user trusted it enough to act on it
Prompt volume measures effort.
Acceptance rate measures leverage.
If acceptance increases while prompt volume decreases, the system is improving.
If prompt volume increases while acceptance stays flat, the system is adding friction.
What Most Indian SaaS AI Features Get Wrong About Measurement
Many Indian SaaS products adding AI layers still evaluate success using feature adoption dashboards rather than workflow outcomes.
For example, an AI-generated ticket summary feature in a support product should not be evaluated on how often summaries are opened. It should be evaluated on whether resolution time drops or escalation rates change.
Similarly, an AI CRM email assistant should not be evaluated on drafts generated. It should be evaluated on reply latency reduction or meeting conversion improvement.
If the metric stops at interface interaction, the product team is measuring exposure instead of impact.
A Practical Test for Choosing a North Star Metric in an AI Product
Before selecting a metric, ask one question:
What decision does the user make faster because this model exists?
If the answer is unclear, the product does not yet have a North Star Metric.
It has an activity counter.
The correct North Star Metric for an AI product measures not how often the system produces output, but how often the user changes behaviour because of it.

메타데이터
- post_id
- 441d3b1b28ee
- slug
- the-north-star-metric-problem-in-ai-products-441d3b1b28ee
- url
- https://medium.com/@Prathikkshaa/the-north-star-metric-problem-in-ai-products-441d3b1b28ee
- canonical_url
- https://medium.com/@Prathikkshaa/the-north-star-metric-problem-in-ai-products-441d3b1b28ee
- author_url
- https://medium.com/@Prathikkshaa
- status
- ok
- fetched_at
- 2026-07-16 20:40:14