← Back to list

Where the return on AI investment lives

The return on AI investment is real, but it concentrates in a handful of value pools defined by the shape of the work…not the…

Rakesh Roshan · 2026-05-30 07:30 · 2 claps · 10.8 min read
#artificial-intelligence #return-on-investment #digital-transformation #product-management #finance
Open on Medium ↗
Wiki topics: AI · AI · General INV · Investing & Markets CRY · Crypto & Web3 BIZ · Business Strategy ECO · Economy · General 📋 · Product Management

Where the return on AI investment lives

The return on AI investment is real, but it concentrates in a handful of value pools defined by the shape of the work…not the sophistication of the model.

In May 2025, Sebastian Siemiatkowski did something unusual for a chief executive in the middle of a pre flotation publicity cycle. He told Bloomberg he had been wrong! 18 months earlier, Klarna had announced that its OpenAI powered assistant was handling two thirds of customer service chats and doing the equivalent work of 700 FTEs, on track to add around 40 million dollars to profit. It had become the most cited proof in financial services that AI could replace people at scale.

Klarna then reversed course and began bringing humans back. The all in AI approach, he told Bloomberg, had produced “lower quality” service, and the company would rebuild a human option for customers who wanted one.

The reversal was widely read as a verdict on AI…even though it was nothing of the sort. Klarna’s assistant still handled enormous volumes of routine work, and the company kept investing heavily in AI elsewhere. What failed was not the technology but a judgment about which work to hand over, and how completely.

Most managing directors are asking a binary question this year: will our investment in AI pay?

After three years of pilots and a capital bill that now runs into the hundreds of billions, the question is perhaps fair. The Klarna episode suggests it is also the wrong one!

The better question about the return on AI investment is: Which investments pay, in which parts of the business, and under what conditions does the return show up at all?

The symptom is the pilot, and the cause is the operating model!

The symptom is easy to see as pilots multiply, demonstrations impress, and almost nothing reaches production at a scale that moves the profit and loss. In July 2025, MIT’s NANDA initiative published The GenAI Divide, drawing on 150 executive interviews, a survey of more than 300 employees, and analysis of 300 public deployments. Its central finding was that around 95% of enterprise generative AI pilots had produced no measurable impact on profit and loss, against enterprise spending the report put at $30 to $40bn. Only about 5% were generating real value. Gartner, looking at the same window, expected at least 30% of generative AI projects to be abandoned after proof of concept by the end of 2025, and found that fewer than half of AI projects reach production at all.

Now, let us set that against the supply side. The four largest Hyperscalers told investors they would spend roughly $700 to $725bn on capital expenditure in 2026, close to double the prior year, with around three quarters of it tied to AI infrastructure. At some of them, free cash flow is now compressing sharply to fund the build. The money going into the capability is real, large, and accelerating. The money coming back out, for most enterprises buying that capability, unfortunately is not yet visible.

Where most diagnoses go wrong is in the explanation that the models are not good enough yet, and that patience will close the gap. The evidence, I believe points the other way. MIT found that failure correlated with integration and workflow, not model quality. The same models that stall in 95% of pilots generate millions in the 5% that succeed. Ninety per cent of firms are actively exploring AI, so enthusiasm is not the constraint. The constraint is structural as Companies are buying a general purpose capability and bolting it onto processes, infrastructure, and data that were never built to absorb it.

The risk is not that AI fails to work but the risk is that capital flows to the most visible use cases rather than the most valuable ones. MIT found more than half of generative AI budgets going to sales and marketing, while the clearer returns sat in back office automation. Spending is being allocated by visibility, not by where the return on AI investment lands. That single misalignment explains most of the disappointment in the numbers above.

‘AI investment returns’ should follow the shape of the work!

Return on AI investment is real but unevenly distributed. It concentrates in a small set of value pools defined by the shape of the work and is captured by firms that have made an operating model choice rather than a technology one.

The implication is uncomfortable for anyone hoping for a single answer to the pay or not question. There is no general return on AI and there are specific returns, in specific kinds of work, available to firms that have organised themselves to collect them. The same model can produce opposite results in two companies, and the variable that separates them is rarely the model.

What separates them is whether the work they pointed it at has the right shape, and whether the surrounding operating model (the data, the talent, The infra, the governance, the vendor architecture…) can turn a capable demonstration into a repeatable outcome. The rest of this piece maps where those returns concentrate, sizes them, and sets out what the firms collecting them are doing differently from the firms still funding pilots.

Where return on AI investment concentrates!

Across the relationships we have originated and the practices we have built around this question, the same four value pools keep appearing. They are not defined by industry or technology, but they are defined by two things a leader can assess before committing capital i.e. how much of the work can run without a human in the loop, and how much value is at risk if the work goes wrong. Plotting any AI use case against those two questions places it in one of four pools and each pool has a different return profile, a risk profile, and a different test for whether you are capturing it.

This is a deliberate departure from the way most segmentations are drawn. The common cut is by technology (machine learning here, generative there, agents next) or by function (finance, marketing, operations). Neither predicts return, because the same technology in the same function can land in two different pools depending on the shape of the underlying task. A model that drafts a routine confirmation letter and a model that decides whether to decline a transaction are the same technology in the same department, but they sit at opposite corners of the map and carry opposite risk. The map I propose here sorts by the property that governs the return.

Pool one: the process reinvention!

This is the automation of high volume, rules bounded work where any single error is cheap and the task repeats thousands of times a day. The shift it represents is from human throughput as the ceiling, to machine throughput at marginal cost. For example, JPMorgan Chase’s COiN platform is the canonical proof as it reviews commercial loan agreements that once consumed roughly 360,000 hours of legal and loan officer time a year. You are in this pool when your unit economics improve, the cost per document, per claim, per reconciliation, without you needing to promise headcount cuts to justify the case. The return here is the most certain of the four, and it is the pool most firms underfund because it is invisible to the board.

Pool two: the expert multiplier!

This pool augments expensive, scarce professionals, lawyers, analysts, engineers, underwriters, so each produces more, faster, at the same or better quality. The shift is from the expert as the bottleneck, to the expert as the editor of a machine generated first draft. JPMorgan’s in house LLM Suite now sits with more than 200,000 employees, with over 40,000 engineers using AI coding assistants and around 100 generative AI applications in production. Bank of America reports that roughly 90% of its workforce uses an AI assistant, and that its internal Erica for Employees has cut IT service desk calls by more than 55%. The indicator is cycle time per expert deliverable, and where senior people spend their hours.

Pool three: the always on front door!

This is customer facing engagement that scales personalised service to near zero marginal cost, handling many interactions instantly and continuously. The shift is from service rationed by the cost of human time to service available around the clock at marginal cost approaching zero. Bank of America’s Erica has handled more than 3.2bn client interactions since 2018 and now serves over 20mn clients, with more than 40% of inquiries on its corporate CashPro Chat resolved through Erica. The test is whether your rate of resolution without a human rise while customer satisfaction holds or improves.

Pool four: the autonomy frontier!

This pool covers systems that decide and act with little or no human review, the agentic frontier that platform vendors are racing toward and that attracts the most ambitious capital. The shift is from decisions reviewed before they take effect, to decisions executed autonomously. The pool is real, but today it is thin, and its most useful proof point is a cautionary one. Klarna’s assistant performed brilliantly on routine queries and degraded on complex, emotionally loaded ones, which is precisely where the value at risk was highest. The indicator to watch is the rate at which errors compound before a human can intervene, and the value at risk per autonomous action. In this pool a 95% success rate is not reassuring, because the other 5% can destroy trust or breach a regulation. For regulated firms in particular, the binding constraint is rarely capability. It is explainability, auditability, and the question of who is accountable when an autonomous decision is wrong. Those constraints do not soften as the model improves, which is why this pool will remain the slowest to convert capability into captured return.

Sizing the pools

The pools are not equal in size, and the largest is not where the nearest term return sits. McKinsey’s estimate that generative AI could add $2.6 to $4.4tn annually across the economy remains the most cited sizing, built from an analysis of 850 occupations and 2,100 work activities across 47 countries. It found around 75% of that value concentrated in four functions: customer operations, marketing and sales, software engineering, and research and development. Customer operations split across the front door and process reinvention. Software engineering and parts of marketing sit in the expert multiplier. The autonomy frontier, the pool drawing the most aggressive investment, accounts for the smallest share of currently capturable value and the highest share of risk.

Where this is already playing out!

Three institutions show the map in operation, each an early mover in a different pool, and one of them a study in what happens at the edges of pool four.

JPMorgan Chase: The largest US bank runs an annual technology budget of around $18bn with an AI and data function reporting directly to the chief executive. It chose to build in house rather than rely on consumer tools, deploying COiN for contract review, LLM Suite to more than 200,000 staff, and AI coding assistants to over 40,000 engineers. Generative tools are reported to have lifted gross sales in its asset and wealth management arm by around 20%. Its leadership has described the prize candidly, framing the work as closing “a value gap between what the technology is capable of and the ability to fully capture that within an enterprise.” The bank’s answer has been to spend years connecting models to proprietary data and software, not to chase the newest demonstration.

Bank of America: Its customer facing assistant Erica has passed 3.2bn interactions since 2018 and serves 20mn clients, while inside the firm roughly 90% of employees use an AI assistant and its IT service desk has shed more than 55% of its call volume. The bank spends about $13bn a year on technology, with roughly $4bn of that on new initiatives in 2025. Its stated approach pairs the automation with human oversight, transparency, and accountability, the discipline that keeps a front door deployment in pool three rather than letting it slide into pool four. [Source: Bank of America Annual Report and newsroom, 2025.]

Klarna: Klarna’s OpenAI powered assistant handled 2.3mn conversations in its first month, two thirds of all chats, cut average resolution time from 11 minutes to under two, and was projected to add about $40mn to profit. Then the quality cost of pushing automation into complex, high stakes interactions surfaced, and in May 2025 the chief executive reversed the all in AI posture and moved to a hybrid model with humans on the complex tail. The numbers were never the problem, but the placement was. Klarna had pushed pool three work into pool four. [Source: OpenAI and Klarna press materials, 2024; Bloomberg, May 2025.]

Match autonomy to value at risk!

The surface lesson from these cases is that some firms are simply better at AI, but the structural lesson is more precise, and most leaders miss it. The firms capturing the return matched the degree of autonomy they granted to the distribution of value at risk in the work, not to the average.

Most work has a long, thin tail so routine queries are high in volume but low in value at risk whilst complex, ambiguous, emotionally charged cases are low in volume but carry almost all the value at risk, in trust, in compliance, in lifetime customer value. Klarna automated by volume, which meant it handed a machine the very tail where the value lived whilst JPMorgan and Bank of America did the opposite i.e., they let machines take the bounded majority and kept humans on the high stakes tail. Same models, opposite outcomes, because the firms read the distribution of the work rather than its average.

MIT found that buying from specialised vendors and building partnerships succeeded roughly twice as often as internal builds, around 67% against a third. JPMorgan succeeds while building in house, but it is the exception that proves the rule: it has the scale, the proprietary data flywheel from trillions of dollars in daily transactions, and a governance structure that most firms cannot replicate. The pattern is not “build your own model.” It is “own the data and the integration, and partner for the rest.”

In the 5% that succeeded, someone owned the outcome and was measured on it, not on whether the tool shipped. Where AI is owned by a function with a number to hit, it tends to land in a pool. Where it is owned by an innovation team measured on launches, it tends to stay a pilot!

Decision, decision, decision …and decision!

The map is not an academic exercise, and it changes four decisions a leader should be making right now.

  1. Reallocate from visibility to value and audit where your AI budget goes, by pool, not by project. If more than half of it sits in front office revenue plays, you are probably misallocated against where the return lands this year, exactly as MIT found across the market. The board conversation to have is to stop funding by demonstration and start funding by pool, with the certain pools funded first.

  2. Draw your autonomy line deliberately, by value at risk, not by capability and the fact that a model can pass the demonstration on the complex tail does not mean you should hand it that tail. Decide which work stays under human review because the value at risk is high and say so explicitly. Then measure resolution without a human and customer satisfaction together, never one without the other.

  3. Make the buy versus build call honestly and unless you have a JPMorgan scale combination of capital, proprietary data, infrastructure and governance, partnering will beat building for most pools. The investment that compounds is not a private model but is the proprietary data and the integration work that let any capable model land in your workflow.

  4. Give every AI initiative a profit and loss owner with a number and the single sharpest filter you can apply this quarter is to ask, for each initiative, who is accountable for the outcome and what number they are committed to move.

…And finally!

Klarna’s reversal was not an admission that AI does not pay. It was a correction of where, and how completely, to use it. The $40mn was real and so was the quality cost of pushing a machine into work that still needed a human on the other end. Read that way, Klarna is not a cautionary tale about AI. It is a precise illustration of the map where a firm that captured real value in pool three and then lost some of it by reaching into pool four before the work was ready.

In summary, return on AI investment is real, but it is earned pool by pool, by firms that match the tool to the shape of the work and then build the operating model to collect what the tool produces. Effectively, the capability is now a commodity whilst the discipline to point it at the right work is not and the choice in front of you is which pool you are buying into.


메타데이터
post_id
f0e32ed01bc0
slug
where-the-return-on-ai-investment-lives-f0e32ed01bc0
url
https://medium.com/@RaxRoshan/where-the-return-on-ai-investment-lives-f0e32ed01bc0
canonical_url
https://medium.com/@RaxRoshan/where-the-return-on-ai-investment-lives-f0e32ed01bc0
author_url
https://medium.com/@RaxRoshan
status
ok
fetched_at
2026-07-13 06:23:13