Blog

Six Ways Your AI ROI Calculation Is Lying to You

By Jean-Philippe Bédard · Published: August 29, 2026

In 2026, AI budgets are no longer experimental. Deloitte measured that employee access to AI rose by 50% in 2025 across leading-edge organizations, according to its State of AI in the Enterprise 2026 survey, published on January 21, 2026 at Davos, covering 3,235 leaders across 24 countries. At the same time, McKinsey's State of AI trust in 2026 survey confirms that maturity is progressing, but gaps persist in strategy, governance, and risk management.

That is the paradox of the year. Adoption is at an all-time high. So is the failure rate. MIT's NANDA initiative, in The GenAI Divide: State of AI in Business 2025 (based on 150 executive interviews, a 350-employee survey, and 300 public deployments), found that approximately 95% of generative AI pilots deliver little to no measurable impact on the P&L. BCG's AI Transformation Is a Workforce Transformation report, published on June 3, 2026, puts a sharper point on it: only about 5% of organizations reap substantial financial gains from AI, and roughly 70% of the value comes from the people and process side, not the algorithm.

The difference between the 5% and the 95% almost always comes down to how the business case was built in the first place. Below are six ways the calculation quietly lies, and a question you can ask before signing off on the budget.


1. The "Hours Saved Equal Dollars Saved" Fallacy

The pattern. An employee saves two hours a day. Multiply by an average loaded cost of $50 per hour, project over a year, and the spreadsheet says the company reclaims $25,000 per head. Multiply by the team. Print the slide.

Why it fails. Theoretical productivity gains rarely convert to bottom-line savings unless two things are true: headcount is actually reduced, or the reclaimed hours are explicitly converted into billable output the customer pays for. When those conditions are not met, the time simply reappears as administrative overhead, longer meetings, or slower chat windows. Microsoft's 2025 Work Trend Index, surveying 31,000 workers across 31 countries, found that only around 20% of organizations had reduced headcount after deploying AI. The other 80% absorbed the time.

The real-world cost. Klarna announced in 2024 that an AI assistant had replaced the work of 700 customer service agents, then acknowledged through 2025 that service quality had declined, and began rebuilding a hybrid human-plus-AI operation. The original ROI estimate was not wrong by a percentage. It was wrong by a category. The cost of recovering customer trust was not on the spreadsheet.

Question to ask. If we did not reduce headcount and we did not bill the recovered hours back to a paying customer, where did the time actually go, and what did it cost?


2. Ignoring Total Cost of Ownership

The pattern. The business case is built against the line item in the vendor contract, or the monthly API estimate at current usage.

Why it fails. AI projects have a cost structure that resembles manufacturing, not software licensing. Once a model is in production, the bill grows in three directions at once:

  • Consumption costs. Token usage and inference calls scale with adoption, often non-linearly.
  • Data costs. Cleaning, labeling, augmenting, and engineering the data the model needs.
  • Human-in-the-loop costs. Reviewing outputs, correcting mistakes, retraining on edge cases, and prompt maintenance.

Gartner's press release of February 26, 2025 (Roxane Edjlali, Senior Director Analyst) warned that 60% of AI projects lacking AI-ready data will be abandoned or will fail by the end of 2026. The cost overruns that drive those abandonments rarely show up in the original business case.

Question to ask. What is our fully loaded annual cost per process at steady state, including data, compute, retraining, and human review, and at what adoption level does the unit economics break?


3. Measuring Vanity Metrics Instead of Business Outcomes

The pattern. The dashboard shows active users, query volume, model versions deployed, and adoption curves climbing. The AI committee reports a successful rollout.

Why it fails. High usage does not mean high value. A team can process millions of model calls that yield no measurable increase in revenue, retention, or cost per case. Adoption is an input, not an output.

BCG's June 2026 report makes this concrete: roughly 5% of organizations see substantial financial gains. The same report notes that the other 95% typically report high adoption and low impact, a paradox that resolves only when you stop measuring input metrics and start measuring outcome metrics.

Question to ask. If we strip out adoption numbers and look only at revenue, margin, error rate, and cycle time, does the line go up?


4. Attributing General Growth to the AI Tool

The pattern. Sales rose 14% the quarter after the AI rollout. The AI tool gets the credit in the steering committee deck.

Why it fails. Without controlled comparison, you cannot separate the AI's contribution from seasonal demand, a parallel marketing campaign, a hiring spree, a pricing change, or a competitor's stumble. McKinsey's State of AI chart, AI at work but not at scale, points out that the small share of high performers isolate their AI impact with cohort analysis or A/B testing. Everyone else is guessing.

Gartner's 2026 Hype Cycle for Agentic AI puts it more bluntly: the rapid progress of agentic AI is exceeded by hype and confusion. A growing share of reported "AI wins" sits inside that hype band.

Question to ask. What would this metric have done in the absence of the AI tool, in a comparable cohort or comparable quarter, and can we show that delta in writing?


5. Treating AI as a One-Time Capital Investment

The pattern. The project is approved as a capital expense, depreciated over three years, and forgotten in month four.

Why it fails. AI models degrade as the world moves on. Customer behavior shifts, regulations change, source data drifts, and the prompts that worked in January stop working in July. The MIT NANDA report makes the point operational: the useful life of a generative AI model in production is routinely shorter than the budget cycle that funded it.

Gartner's press release of June 25, 2025 predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, in large part because ongoing governance, retraining, and maintenance were never priced in. The depreciation schedule in the original business case becomes fiction by month nine.

Question to ask. What is our quarterly maintenance budget for monitoring, retraining, prompt updates, and governance review, and what trigger tells us the model has drifted past the point of useful return?


6. Ignoring Risk, Security, and Compliance Liabilities

The pattern. The ROI is calculated on operational speed. Risk is handled in a separate workstream, often by another team, often after the budget is approved.

Why it fails. A single AI incident can erase years of efficiency gains. IBM's Cost of a Data Breach Report 2025 puts the global average cost of a data breach at $4.88 million, up 10% year over year. AI and shadow-AI exposures are flagged as a growing vector inside that number.

Regulatory exposure has become concrete. Under the EU AI Act, fines for the most serious violations reach €35 million or 7% of global annual turnover, whichever is higher. The Digital Omnibus, adopted in 2026, postponed the bulk of high-risk obligations to December 2, 2027, but Article 50 transparency obligations and the full penalty regime took effect on August 2, 2026. Risk-adjusted ROI is no longer a theoretical refinement. It is the law.

Question to ask. If this AI tool produces one materially wrong decision, one privacy breach, or one compliance violation in the next twelve months, what is the expected financial loss, and is it on the same slide as the projected gain?


What the Six Errors Have in Common

Every error above is the same mistake wearing a different costume: the organization commits resources before the business case has been tested against evidence.

The 5% that capture real value tend to do three things the other 95% skip. They score each candidate process against explicit criteria before they engage budget. They isolate the AI's contribution with a baseline they can defend in front of a CFO. And they set a hard profitability threshold, then walk away from any project that fails it.

A feasibility scorecard that forces an answer on decision complexity, input variability, transaction volume, error cost, and latency tolerance catches every error in this article before the first invoice is signed.


What Works: Redesign the Operation First

The natural reaction to a list of ROI failures is to keep doing the same thing with a tighter spreadsheet. The 2026 evidence points the other way.

The same BCG report that puts the winners at 5% of organizations is more precise about where the value actually comes from. About 10% of the value of AI comes from the algorithms themselves. Another 20% comes from the technology needed to deploy them. The remaining 70% comes from rethinking the people and process component. In other words, an organization that drops a model into an unchanged workflow captures at most the 30% that sits in the technology stack. The organization that re-engineers the workflow around what the model is good at captures the rest.

Klarna's reversal in 2025 is the clearest demonstration. The initial 2024 announcement framed AI as a direct replacement for human customer service agents. By 2025, the company acknowledged that service quality had declined and began rebuilding a hybrid operation. The lesson was not "AI failed." The lesson was "replacing a workflow with a model is not the same as redesigning a workflow to use a model." The hybrid that emerged kept the AI for high-volume, low-complexity queries, and routed disputes, complex refunds, and hardship cases back to humans. The cost of the redesign was real. The cost of not doing it, in customer trust and re-acquisition, was larger.

Microsoft's 2025 Work Trend Index makes the same point from the employee side. In organizations that captured value from AI, the freed-up hours were reinvested in higher-judgment tasks: customer conversations, edge-case problem solving, mentoring, and quality control. In organizations that did not, the same hours simply lengthened the average meeting. The AI was identical in both cases. The operational design around it was not.

The pattern holds across the six errors in this article. Every one of them becomes smaller when the workflow is redesigned at the same time as the model is introduced.

  • The "hours saved" fallacy loses force if the operation explicitly routes the saved time into a measurable new output.
  • The TCO surprise shrinks if the data, supervision, and retraining costs are absorbed into the redesigned process from day one.
  • Vanity metrics lose legitimacy if the scorecard tracks outcomes against the redesigned workflow, not adoption against the old one.
  • The growth-attribution error becomes visible if the redesign includes a control group running the previous workflow.
  • The "one-time capital" error fades if the redesign itself is the unit of accounting, with quarterly checkpoints built in.
  • The risk-blind error shrinks if the redesigned workflow includes a human override step at every high-cost decision.

The implication is direct. Before measuring the ROI of an AI tool, measure the ROI of the operational redesign you are willing to commit to at the same time. If that number is zero, the AI's ROI is unlikely to climb above the 30% ceiling that BCG reports.


A Note on Sources and Limits

This article draws on primary sources published in 2024 through 2026. Where a figure originated in an English-language report, the citation has been translated freely into French in the French version, with a note to that effect at the top of the Sources section.

The strongest sources used here are:

  • Deloitte State of AI in the Enterprise 2026 (January 21, 2026, Davos) for adoption velocity.
  • McKinsey State of AI trust in 2026 (April 2026) for the maturity thesis.
  • BCG AI Transformation Is a Workforce Transformation (June 3, 2026) for the 5% and 70% figures.
  • MIT NANDA The GenAI Divide: State of AI in Business 2025 (August 2025) for the 95% pilot figure and methodology.
  • Gartner press releases of February 26, 2025 and June 25, 2025 for the 60% and 40% predictions, and the 2026 Hype Cycle for Agentic AI for the "hype exceeds progress" framing.
  • IBM Cost of a Data Breach Report 2025 for the $4.88 million global average.
  • EU AI Act (Regulation (EU) 2024/1689) for the regulatory ceiling, with the Digital Omnibus delay noted honestly.

One important limit. There is no single study measuring "how much companies overestimate AI ROI on average." The errors above are inferred from converging evidence (Klarna, the 95% pilot figure, the 5% gainers). Anyone claiming a precise percentage for that overestimate is, in our experience, making it up.


Want a Way to Catch These Errors Before They Hit the Budget?

A short process-feasibility scorecard, scored before any commitment, would have flagged every case above. Five axes, one profitability gate, no agent built unless the evidence holds.

If your team is sizing up an AI initiative this quarter, the Agentic Process Value framework (APV) is the methodology we use to do exactly that audit. It is published in five languages at agenticprocessvalue.ai.


Sources

  • Deloitte, State of AI in the Enterprise 2026, published January 21, 2026, Davos. (Official survey, 3,235 leaders across 24 countries.)
  • McKinsey, State of AI trust in 2026: Shifting to the agentic era, AI Trust Maturity Survey, April 2026. (Official survey; title, subtitle, and authors verified through archive.org snapshot of April 22, 2026. Body content not accessible at time of writing.)
  • BCG, AI Transformation Is a Workforce Transformation, fourth annual survey, published June 3, 2026. (Official report.)
  • MIT NANDA initiative, The GenAI Divide: State of AI in Business 2025, methodology: 150 executive interviews, 350-employee survey, 300 public deployments. (Academic research report; coverage in Fortune, August 18, 2025.)
  • Gartner, press release, Lack of AI-Ready Data Puts AI Projects at Risk, February 26, 2025 (Roxane Edjlali). (Official press release.)
  • Gartner, press release, Gartner Predicts Over 40 Percent of Agentic AI Projects Will Be Canceled by End of 2027, June 25, 2025. (Official press release.)
  • Gartner, 2026 Hype Cycle for Agentic AI, summer 2026. (Official research cycle.)
  • Microsoft, 2025 Work Trend Index, 31,000 workers across 31 countries. (Official survey.)
  • IBM, Cost of a Data Breach Report 2025, published July 2025. (Official annual industry report.)
  • Klarna, public statements on AI customer service strategy, 2024 through 2026, with reporting in Forbes (May 2025), Reuters, and Bloomberg. (Official corporate communications plus press coverage.)
  • European Union, Regulation (EU) 2024/1689 (the EU AI Act), Article 99 on penalties, with the Digital Omnibus delay adopted in 2026. (Official regulation.)

← Back to blog