Make this research yours. Add it to a free WorldbyFlow workbench to run follow-ups, ask questions, and re-check it as events move.
Add to your workbench — free
WorldbyFlow•Structured Research
Generated September 20, 2026· technology· 37 sources

Is Enterprise AI Investment Already Delivering Measurable Earnings Returns?

The Arguments
Share
The Proposition
Enterprise AI investment is already delivering measurable earnings returns at scale, rather than mostly producing individual productivity gains and vendor-reported case studies that have not yet shown up in organization-wide financial results.

Overview

Vendors and some enterprises point to agentic AI deployments with triple-digit ROI claims, while McKinsey's 2026 global survey finds only 6% of organizations qualify as AI 'high performers' with a significant, attributable EBIT impact, and independent field research finds most generative AI pilots show no measurable profit-and-loss effect.

Brief

The debate sits on a data set both sides can cite honestly, because the numbers describe different things. McKinsey's State of AI global survey, published August 25, 2026 and based on 1,719 executives across 97 nations, found that 37% of respondents report AI has contributed positively to their organization's EBIT — essentially unchanged from 2025 — while only 6% qualify as 'AI high performers,' defined as attributing at least 5% of EBIT to AI with significant reported value, also flat year over year. Meanwhile 80% of respondents said AI has improved their individual productivity. That 80-versus-37-versus-6 spread is the entire argument in three numbers: capability and individual usefulness are not disputed; enterprise-level earnings attribution is.
Separately, MIT's Project NANDA GenAI Divide report, based on analysis of 300 AI deployments, 52 executive interviews, and a 153-leader survey, found that 95% of generative AI pilots delivered no measurable P&L impact, while only 5% of integrated systems created significant value — a preliminary, non-peer-reviewed finding whose starkest claims rest specifically on the 52 interviews rather than audited financials. Enterprise generative AI spend tripled in a single year, rising from roughly $11.5 billion in 2024 to about $37 billion in 2025 per Menlo Ventures' survey of US enterprise decision-makers, while hyperscaler AI infrastructure capex has climbed into the hundreds of billions for 2026 across Microsoft, Amazon, Alphabet, and Meta. The dollars flowing in are large and rising; the dollars flowing back out as attributable earnings, per McKinsey's own instrument, have not moved.
Against that backdrop sit specific, named case studies that vendors and some enterprises use to argue ROI is real today: Klarna's February 2024 announcement that its OpenAI-powered assistant handled roughly two-thirds of customer service chats in its first month, equivalent to the workload of about 700 agents, and JPMorgan's reported use of AI across hundreds of production use cases. But Klarna's own subsequent history complicates the case for using it as unqualified proof: in May 2025 the company began rehiring human agents after its CEO told Bloomberg that evaluating the shift primarily on cost had produced 'lower quality,' and by 2026 Klarna had settled into a hybrid model routing routine queries to AI and complex or premium interactions to humans. The case shows AI can absorb high-volume, bounded work; it does not show that full-scale AI replacement reliably protects earnings without new costs in quality and re-hiring, a distinction the current debate frequently elides.
McKinsey's 2026 data offers a mechanism for why the split exists: cost reductions from AI concentrate in supply-chain management, service operations, and manufacturing, and revenue gains concentrate in marketing, sales, and product development — narrow, bounded use cases with a measurable output — while broad enterprise chatbot and copilot rollouts tend to produce individual time savings that do not visibly flow to the income statement. The single factor McKinsey ties most strongly to becoming a high performer is workflow redesign: nearly three-quarters of high performers report having fundamentally redesigned workflows around AI, versus about a quarter of everyone else. That finding reframes the entire proposition: the question is not whether AI can generate earnings, since a real subset of organizations demonstrably does, but whether 'investment' alone — capital deployed, tools purchased, pilots launched — is a reliable predictor of earnings return, and the 2026 survey data says it is not.

The Arguments

The Case For(4)
A real, identifiable subset of enterprises is already converting AI investment into measurable earnings impact, and the share attributing any EBIT impact to AI is not trivial.
Reasoning: 37% of organizations reporting a positive EBIT contribution from AI is a substantial minority, not a rounding error, and it demonstrates the mechanism works when applied correctly rather than merely in theory.
Evidence: McKinsey's 2026 survey found that only 37% of respondents said AI has contributed to their organization's earnings before interest and taxes (EBIT), a figure McKinsey says is essentially unchanged from the prior year's survey, and among that group, AI high performers—respondents who attribute an EBIT impact of 5 percent or more to AI use and say their organizations have seen 'significant' value from AI use—account for just 6 percent of survey respondents, unchanged from 2025.
Moderate strength
The gap between high performers and everyone else is explained by a specific, replicable practice — workflow redesign — not by unrepeatable luck or company size, meaning the return is achievable by design rather than accidental.
Reasoning: If the differentiator were random or purely a function of scale, ROI could not be treated as a deliberate strategy; because it correlates with a specific organizational choice, it argues that earnings returns are available on demand to firms that make that choice.
Evidence: McKinsey reports that about three-quarters of AI high performers have fundamentally redesigned their workflows around AI, versus roughly a quarter of everyone else, and they are far more likely to pursue growth and innovation rather than pure cost-cutting, and are about twice as likely to have strong senior-leadership ownership of AI strategy and clear metrics for measuring its impact.
Strong strength
Agentic AI deployments specifically — as distinct from broad copilot and chatbot rollouts — show a materially higher first-year ROI rate, indicating the technology can deliver returns when deployed as an end-to-end workflow replacement rather than an assistant layered onto existing work.
Reasoning: If a specific deployment architecture (agentic, end-to-end) shows sharply higher realized-return rates than the general population, that is evidence the earnings impact exists and is attributable to identifiable technical and organizational choices, not to AI capability broadly failing.
Evidence: Reporting citing McKinsey data states that among agentic AI early adopters, 88% report ROI within the first year, against 74% for gen-AI users overall, and separately that 40% of $1B-plus revenue companies are now scaling AI agents, up sharply from 27% a year earlier — vendor- and survey-adjacent figures that should be read as directional rather than audited.
Contested strength
Named enterprise deployments demonstrate concrete, functional-level financial effects even where enterprise-wide EBIT attribution has not yet been established, showing the return exists at the business-unit level and will likely aggregate over time.
Reasoning: Financial impact often shows up first in specific functions before it consolidates into a reportable enterprise-level metric, so functional-level gains are a leading indicator rather than an absence of return.
Evidence: McKinsey's own survey data show that respondents most often report cost reductions from AI in supply chain management, service operations, and manufacturing — the narrow category of use cases where AI is doing a defined, bounded job with a measurable output — while revenue gains show up in marketing and sales and product development, and separately even respondents who don't report enterprise-level EBIT impact do say their organizations are seeing financial impact from specific business functions' use of AI.
Moderate strength
The Case Against(4)
The dominant, survey-confirmed reality is that AI investment has not moved enterprise earnings at all for the large majority of organizations, and this share has been flat for a full year despite a sharp rise in spending and deployment.
Reasoning: If earnings-impact rates were merely lagging a fast-moving rollout, the share should be climbing; a flat year-over-year rate despite rising capital deployment argues the gap is structural rather than a timing effect.
Evidence: McKinsey's 2026 data shows 31% report some EBIT impact but below the high-performer threshold; the remaining 63% report no measurable enterprise earnings impact at all, even as they expand AI deployment and report strong individual productivity gains, and the 6% high-performer share is unchanged from last year — a figure flat since 2025 despite the sharp rise in enterprise AI budgets over the same period.
Strong strength
Independent field research beyond McKinsey's self-reported survey data finds an even starker failure rate at the pilot level, suggesting the earnings gap is not a McKinsey artifact but converges with data from a separate methodology.
Reasoning: When two differently constructed studies — a large executive perception survey and a smaller but more granular deployment-and-interview study — both point to the same conclusion, the conclusion is harder to dismiss as an artifact of one survey's framing.
Evidence: MIT's Project NANDA GenAI Divide report, based on analysis of 300 AI deployments, 52 executive interviews, and a 153-leader survey, found that 95% of generative AI pilots delivered no measurable P&L impact, while only 5% of integrated systems created significant value — a preliminary, non-peer-reviewed finding whose starkest claim rests specifically on the 52 interviews the report itself flags as directionally accurate rather than official company reporting.
Moderate strength
Individual productivity gains are real and widely reported, but they are demonstrably not the same thing as organizational earnings, and the size of that gap is the core evidence against the proposition.
Reasoning: A worker being measurably faster does not by itself change a company's income statement unless the saved time is captured, reallocated, or converted into headcount or output reduction — and the data shows this conversion mostly is not happening.
Evidence: McKinsey's 2026 survey found that 80% of respondents said AI has improved their individual productivity, and 50% said it helps them make better decisions," yet "only 37% of respondents said AI has contributed to their organization's earnings before interest and taxes, with reporting characterizing this as the finding that organizations' conviction in AI is growing faster than the immediate financial returns they can attribute to it.
Strong strength
Even celebrated, publicly cited case studies of AI replacing enterprise workflows have required costly walk-backs when quality and customer experience deteriorated, undercutting the reliability of ROI claims built on early, self-reported vendor figures.
Reasoning: If the most widely cited proof point for AI-driven earnings gains required a reversal and rehiring within about a year of its announcement, that is direct evidence the initially reported financial benefit was not durable as first presented.
Evidence: Klarna's CEO told Bloomberg in May 2025 that as cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality, and independent analysis of the case notes that the $40 million figure wasn't an audited cost saving; it was described as a projected 'profit improvement to Klarna in 2024,' a forward-looking estimate at the time of announcement," and the "700 full-time agents' figure was a productivity equivalence claim, not a headcount action.
Moderate strength

The Strongest Point on Each Side

Strongest For
A defined, replicable practice — fundamentally redesigning workflows around AI rather than bolting AI onto existing processes — separates a real 6% of organizations that report significant, attributable EBIT impact from the rest, showing the earnings return is achievable by design and not merely theoretical.
Strongest Against
The share of organizations reporting any EBIT impact from AI, and the smaller share reporting significant EBIT impact, has stayed flat for a full year even as AI budgets and deployment scale sharply higher, and independent field research using a different methodology finds a similarly stark gap between AI pilots and measurable profit impact.

What It Turns On (4)

Does 'productivity' at the individual or task level reliably convert into 'earnings' at the enterprise level, or does saved time simply stay at the worker's desk without being captured by the organization?
This is the empirical crux the entire debate turns on: McKinsey's own data shows an 80%-to-37%-to-6% cascade from individual productivity to any EBIT impact to significant EBIT impact, and resolving why that cascade narrows so sharply would settle whether current investment levels are justified or premature.
Is workflow redesign a prerequisite for earnings return, or merely correlated with the kind of company that would have succeeded with AI anyway?
If redesign is causal, the 94% not seeing earnings impact have an identifiable, actionable fix; if redesign is just a marker of already-disciplined, well-resourced organizations, then most companies may be structurally unable to replicate the 6%'s results regardless of how they deploy AI.
Should vendor-reported and single-company case studies (Klarna, JPMorgan, agentic-adopter surveys claiming 171% average ROI) be weighted as evidence of realized returns, or discounted as unaudited, self-selected, and pre-reversal snapshots?
Much of the case for the proposition rests on figures originating from the companies or vendors being evaluated rather than from independent audited financials, and the Klarna case shows an initially celebrated figure was later described by its own source as a forward-looking estimate rather than a realized saving.
Is the flat 6%/37% share evidence of a structural ceiling on enterprise AI earnings impact, or a lagging indicator that has not yet caught up to a genuinely accelerating deployment curve?
The scaling-of-agents share among large companies rose sharply year over year even as the EBIT-impact share stayed flat, so whether earnings attribution is a leading or lagging measure changes whether current investment levels are rational bets on a future payoff or capital being spent ahead of any evidence it will return.

What Each Side Concedes

Proponents must concede that the flat 6%/37% figures are McKinsey's own instrument and have not moved despite a year of rising investment, which is hard evidence against near-universal current-state returns. Skeptics must concede that a genuine, if small, cohort of organizations is producing significant, self-reported EBIT impact through identifiable practices, which means the technology is not incapable of generating returns — the constraint is organizational, not technical.

Where the Evidence Points

The weight of the current evidence favors the skeptical reading of the proposition as stated: McKinsey's flat 6% high-performer share and 37% any-impact share, corroborated directionally by MIT's independent pilot-failure findings, indicate that enterprise AI investment in aggregate has not yet produced earnings returns at scale, even though a genuine minority of well-organized adopters clearly has. The unresolved question is whether that minority represents a preview of where the majority is headed with a lag, or a ceiling defined by organizational capacity that most enterprises will not clear regardless of further spending — the data available as of the August 2026 survey cannot yet distinguish between those two futures.

Common Ground

  • Both sides agree that AI improves individual-level productivity for a large majority of users.
  • Both sides agree that a genuine subset of well-run organizations is achieving measurable, significant EBIT impact from AI.
  • Both sides agree that fundamental workflow redesign, rather than tool deployment alone, is the strongest known differentiator between organizations that see earnings impact and those that do not.

Open Questions

  • Will the 6% high-performer share and 37% any-impact share rise in McKinsey's 2027 survey, or will they remain flat for a second consecutive year despite continued capex growth?
  • Can the workflow-redesign practices of current high performers be replicated by resource-constrained organizations, or do they require capital and leadership commitment only available to a subset of large enterprises?
  • How many additional publicized ROI case studies (in the pattern of Klarna) will require a walk-back once initial vendor- or company-reported savings are tested against audited financials and sustained customer or quality metrics?
medium uncertainty· model's epistemic confidence in this analysis

Facts & Figures (8)

The claims behind this analysis, each with its verification status — including what is contested, unverified, or could not be established. What each grade means
A real, identifiable subset of enterprises is already converting AI investment into measurable earnings impact, and the share attributing any EBIT impact to AI is not trivial.
McKinsey's 2026 survey found that only 37% of respondents said AI has contributed to their organization's earnings before interest and taxes (EBIT), a figure McKinsey says is essentially unchanged from the prior year's survey, and among that group, AI high performers—respondents who attribute an EBIT impact of 5 percent or more to AI use and say their organizations have seen 'significant' value from AI use—account for just 6 percent of survey respondents, unchanged from 2025.
✓ DOCUMENTEDsource ↗case for
The gap between high performers and everyone else is explained by a specific, replicable practice — workflow redesign — not by unrepeatable luck or company size, meaning the return is achievable by design rather than accidental.
McKinsey reports that about three-quarters of AI high performers have fundamentally redesigned their workflows around AI, versus roughly a quarter of everyone else, and they are far more likely to pursue growth and innovation rather than pure cost-cutting, and are about twice as likely to have strong senior-leadership ownership of AI strategy and clear metrics for measuring its impact.
✓ DOCUMENTEDsource ↗case for
Agentic AI deployments specifically — as distinct from broad copilot and chatbot rollouts — show a materially higher first-year ROI rate, indicating the technology can deliver returns when deployed as an end-to-end workflow replacement rather than an assistant layered onto existing work.
Reporting citing McKinsey data states that among agentic AI early adopters, 88% report ROI within the first year, against 74% for gen-AI users overall, and separately that 40% of $1B-plus revenue companies are now scaling AI agents, up sharply from 27% a year earlier — vendor- and survey-adjacent figures that should be read as directional rather than audited.
○ REPORTEDsource ↗case for
Named enterprise deployments demonstrate concrete, functional-level financial effects even where enterprise-wide EBIT attribution has not yet been established, showing the return exists at the business-unit level and will likely aggregate over time.
McKinsey's own survey data show that respondents most often report cost reductions from AI in supply chain management, service operations, and manufacturing — the narrow category of use cases where AI is doing a defined, bounded job with a measurable output — while revenue gains show up in marketing and sales and product development, and separately even respondents who don't report enterprise-level EBIT impact do say their organizations are seeing financial impact from specific business functions' use of AI.
✓ DOCUMENTEDsource ↗case for
The dominant, survey-confirmed reality is that AI investment has not moved enterprise earnings at all for the large majority of organizations, and this share has been flat for a full year despite a sharp rise in spending and deployment.
McKinsey's 2026 data shows 31% report some EBIT impact but below the high-performer threshold; the remaining 63% report no measurable enterprise earnings impact at all, even as they expand AI deployment and report strong individual productivity gains, and the 6% high-performer share is unchanged from last year — a figure flat since 2025 despite the sharp rise in enterprise AI budgets over the same period.
✓ DOCUMENTEDsource ↗case against
Independent field research beyond McKinsey's self-reported survey data finds an even starker failure rate at the pilot level, suggesting the earnings gap is not a McKinsey artifact but converges with data from a separate methodology.
MIT's Project NANDA GenAI Divide report, based on analysis of 300 AI deployments, 52 executive interviews, and a 153-leader survey, found that 95% of generative AI pilots delivered no measurable P&L impact, while only 5% of integrated systems created significant value — a preliminary, non-peer-reviewed finding whose starkest claim rests specifically on the 52 interviews the report itself flags as directionally accurate rather than official company reporting.
◑ CONTESTEDcase against
Individual productivity gains are real and widely reported, but they are demonstrably not the same thing as organizational earnings, and the size of that gap is the core evidence against the proposition.
McKinsey's 2026 survey found that 80% of respondents said AI has improved their individual productivity, and 50% said it helps them make better decisions," yet "only 37% of respondents said AI has contributed to their organization's earnings before interest and taxes, with reporting characterizing this as the finding that organizations' conviction in AI is growing faster than the immediate financial returns they can attribute to it.
✓ DOCUMENTEDsource ↗case against
Even celebrated, publicly cited case studies of AI replacing enterprise workflows have required costly walk-backs when quality and customer experience deteriorated, undercutting the reliability of ROI claims built on early, self-reported vendor figures.
Klarna's CEO told Bloomberg in May 2025 that as cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality, and independent analysis of the case notes that the $40 million figure wasn't an audited cost saving; it was described as a projected 'profit improvement to Klarna in 2024,' a forward-looking estimate at the time of announcement," and the "700 full-time agents' figure was a productivity equivalence claim, not a headcount action.
✓ DOCUMENTEDsource ↗case against

Sources (37)

More technology research
Grounded in 37 web sources · 8 facts on the ledger · 6 verified or grounded · 1 partial or attributed · 1 contested · how the grades work
Analysis generated by WorldbyFlow from publicly available information. WorldbyFlow does not verify claims or endorse conclusions. New here? The two-minute overview.