Researchers at Aalto University ran an experiment that should worry anyone spending money on AI. They gave 698 people LSAT logical reasoning problems to solve — 491 of them working with ChatGPT, the rest as a control. Using AI made people better. Scores rose by about three points. But people’s estimates of their own performance rose further, running roughly four points above what they actually scored.
The Dunning-Kruger effect is supposed to protect the competent. Only the least skilled dramatically overrate themselves. Aalto found that when AI enters the picture, the effect disappears entirely. Everyone overrates themselves. And the people with the highest AI literacy misjudged their own performance the most.
That gap – between what we produce with AI and what we think we produced — might be the most expensive cognitive bias in business right now. Because it’s not the people who can’t use AI who are burning money. It’s the people who are certain they can.
The researchers attributed it to “cognitive offloading” — the tendency to trust AI output without actually evaluating whether it did what you needed.
You’ve felt this. You’ve prompted ChatGPT, skimmed the output, thought that looks about right, and moved on. We all have. It’s the most natural thing in the world. It’s also the mechanism behind the failure rates you’re about to see.
The number everyone's quoting (and getting wrong)
You’ve probably seen the stat: 95% of AI projects fail. It comes from MIT’s Project NANDA report, The GenAI Divide, published July 2025. The number is real. It appears in the study. And it needs serious context that almost nobody provides.
NANDA reviewed 300+ publicly disclosed AI initiatives, interviewed representatives from 52 organisations, and surveyed 153 senior leaders. Their finding: 95% of organisations are getting zero return from generative AI. But “zero return” has a specific definition here – no measurable profit-and-loss impact within six months of pilot. That excludes efficiency gains, cost reductions, customer experience improvements, and anything that doesn’t show up on a P&L statement in half a year.
There’s a bigger problem. Project NANDA builds AI agent infrastructure. Their central finding – that AI fails because it lacks learning, memory, and adaptation – directly describes the product they’re selling. Ajit Jaokar, a course director at Oxford’s Department for Continuing Education, titled his LinkedIn response to the report “a clever marketing gimmick.” The lead author is listed as a Microsoft employee, not MIT faculty. There’s no peer review.
None of this means the finding is wrong. It means we need a better anchor.
BCG surveyed 1,000 CxOs and senior executives in October 2024 – the most rigorous methodology in the space. Their number: 74% of companies struggle to achieve and scale value from AI. Only 4% consistently generate returns worth talking about.
The rest of the research stacks in the same direction. RAND cites estimates putting AI project failure at roughly double the rate of comparable non-AI IT projects. Gartner found fewer than 30% of AI leaders say their CEOs are satisfied with the return on AI investment, and reported in 2026 that only 28% of infrastructure and operations AI projects deliver measurable ROI. S&P Global found 42% of companies abandoned the majority of their AI initiatives before they reached production – up from 17% the year before.
What the AI failure research actually measured
Four studies, four different questions. The headline percentages are not comparable — they count different things against different thresholds.
| Study | Figure | What it counts | What the figure actually means |
|---|---|---|---|
| BCGOctober 2024 1,000 CxOs and senior executives, 59 countries |
74% | Companies All AI |
Have not yet shown tangible value from AI. A maturity measure, not a failure count — only 4% were generating consistent returns. |
| MIT NANDAJuly 2025 300+ initiatives reviewed, 52 interviews, 153 survey responses |
95% | Organisations Generative AI only |
Zero measurable P&L return. Not peer-reviewed, and the 95% has no clearly traceable derivation from any of the three samples. NANDA builds AI agent infrastructure; the lead author is a Microsoft employee. |
| S&P Global / 451 ResearchJanuary 2025 1,006 IT and line-of-business professionals |
42%up from 17% | Companies All AI |
Abandoned the majority of their AI initiatives before reaching production. Not initiatives abandoned, and not abandoned entirely. |
| GartnerApril 2026 782 I&O leaders, surveyed Nov–Dec 2025 |
28%20% fail outright | Use cases Infrastructure & Operations only |
Fully succeed and meet ROI expectations. This is a success rate — the inverse is not a 72% failure rate, because roughly half of the remainder partially succeed. |
On the 80% figure you'll see quoted everywhere. It is usually credited to RAND. RAND did not measure it — its 2024 study interviewed 65 practitioners and academics about the root causes of failure, and quoted the 80% in passing from a Fortune article. It is a journalist's estimate wearing a think tank's name, which is why it isn't in this table.
Why the numbers still matter. Four organisations asked four different questions of four different populations and every one of them found the same direction. That is stronger evidence than any single percentage — and it is the reason no single percentage should be quoted as the AI failure rate.
Different methods. Same direction. No coordination.
The interesting data isn’t in the failure rate. It’s buried in what the successful minority actually did.
Between 2024 and 2025, the same conclusion kept surfacing from different directions. Two research bodies found it in their data: RAND traced failure back to how problems get defined, and MIT NANDA measured the gap between designed and undesigned builds. McKinsey modelled it. BCG built a prescription around it. Two findings, one model, one recommendation – different kinds of evidence, produced without coordination, all pointing at the same differentiator between failed AI projects and successful ones. Not the model. Not the budget. Not the team size or the prompt quality. Whether anyone designed the system before they started building it.
McKinsey put numbers on it. In their June 2025 paper on agentic AI, they model a customer call centre where the same technology gets deployed two ways. Passive deployment – adding AI as a tool alongside existing processes – produces 5–10% improvement. Redesigning the entire workflow around what AI can do produces a 60–90% reduction in resolution time, with up to 80% of common incidents resolved automatically. It’s a model, not a field study, so treat the exact figures as illustrative. But the shape of it matches everything the field data shows: same technology, different design, a different order of outcome.
RAND’s analysis, The Root Causes of Failure for Artificial Intelligence Projects, found the top cause of failure wasn’t technical capability. It was miscommunication about what problem to solve. Not a technology problem. A thinking problem.
BCG’s prescription should make every AI tool buyer uncomfortable. Their 10-20-70 principle says the effort split should be 10% on algorithms, 20% on technology and data, and 70% on people and process. Most organisations run it backwards – pouring budget into tools and platforms while spending almost nothing on figuring out what those tools should actually accomplish.
MIT NANDA’s own data reinforces this from a different angle: external vendor solutions – where someone who had already architected the system sold a ready-made tool – reached full deployment 66% of the time. Internal builds, where teams started from scratch without that design layer, reached deployment 33% of the time. Twice the rate, and the difference was the design work that happened before deployment. NANDA’s own caveat applies: that split comes from its 52-organisation interview sample, not the whole market.
Different methods, different samples, different incentives. The same answer: the gap isn’t in the AI. It’s in the thinking that happens before anyone opens a tool.
"I don't have time to architect anything"
This is where the article is supposed to tell you to slow down, be more strategic, design a comprehensive AI architecture before touching a keyboard.
And you’re already resisting it. Because you’re reading this on a Tuesday with 40 unread emails and a client deliverable due Thursday. You don’t need a framework. You need output. And that urgency – completely legitimate, completely real — is precisely the pressure that produces the failure rate.
But urgency is a permanent condition. There will never be a calm Tuesday where you have four unblocked hours to design your AI architecture. If you wait for that day, you become a statistic.
But “slow down and think first” is also wrong. Or at least incomplete. Because it misidentifies the time cost.
The redesign path doesn’t necessarily cost more time. It spends time in a different place – upfront, defining the problem, mapping the workflow, and specifying exactly what “working” looks like. The passive path starts fast and then pays it all back in debugging, retraining, and manual corrections. Call it a correction tax. You pay for the thinking either way. The only question is whether you pay before the build or after it, with interest.
We went through this exact progression. In 2023, like most people, we were doing one-shot prompts. “Write me a blog post about X.” The output was 80% acceptable – close enough to seem useful, far enough from good to never survive contact with a client. You’d read it and feel that sinking recognition: this sounds like AI wrote it. Every piece needed heavy human editing. We were paying for AI and then paying again to fix what AI produced.
The shift wasn’t learning better prompts. It was stopping to define the system. What does a good outcome actually look like? What does this specific client need? What information does the AI require to produce that outcome consistently?
Once we asked those questions, the architecture followed. We built specialised agents for different content types, each with its own knowledge base and brand guidelines. We created workflows that integrate client strategy before any content gets generated. One coordinating system — we call it “The Brain” – directs specialised tools rather than asking one general tool to do everything.
That’s not a technology achievement. Every one of those components uses tools available to anyone. The achievement was in the design – knowing what each component needed to do before we asked AI to do it.
Where the "AI is failing" narrative is right
It would be dishonest to suggest that architectural thinking solves everything. Some of the failure is real, structural, and technology-driven.
Gartner officially placed generative AI in the Trough of Disillusionment in their 2025 Hype Cycle for Artificial Intelligence. That’s not editorial commentary – it’s a documented, recurring pattern where technology capability temporarily falls short of inflated expectations before eventually maturing into productivity. We’re in the trough right now.
The MIT study’s “learning gap” is a genuine technical limitation – systems that don’t retain feedback, can’t adapt to context, and fail on edge cases. The organisations NANDA interviewed complained repeatedly about tools that repeated errors despite correction and hallucinated on anything outside common scenarios. The report doesn’t attach percentages to those complaints, but they surfaced across organisations, and they match what anyone who has run these tools in production already knows.
IBM’s 2025 CEO study, covering 2,000 chief executives across 33 countries, found only 25% of AI initiatives had delivered their expected ROI. A separate IBM study of 2,500 executives found the return on early generative AI initiatives has settled at around 7% – down from as high as 31% in 2023, and below a typical cost of capital of roughly 10%. Microsoft-funded research from IDC tells a more optimistic story: an average of $3.70 returned per dollar invested, rising to $10.30 among the self-identified leaders. Those figures are self-reported, and they come from a study commissioned by a company with a direct financial interest in AI adoption looking good. The two studies aren’t even measuring the same thing. IBM reports a rate of return; IDC reports a multiple on spend. But even allowing for the different units, one says AI barely covers its cost of capital and the other says it pays for itself several times over. Both are probably true for their respective samples, which tells you how wide the variance is.
There’s also a human cost that the ROI figures don’t capture. A study of 381 South Korean employees, published in Humanities and Social Sciences Communications, a Nature Portfolio journal, found AI adoption is associated with lower psychological safety and higher depression scores. The design is correlational, so it can’t prove cause, but the direction is not comforting. Prosci’s research found more than 73% of surveyed change practitioners say their organisations are near, at, or beyond the saturation point. Failed AI rollouts don’t just waste budget — they erode what researchers call “innovation trust,” making the next transformation effort harder to sell internally.
The failures are real. The question is whether they’re inevitable.
The expensive middle ground
Most AI failure doesn’t look like mediocrity that’s slightly too expensive to justify.
It looks like a chatbot that answers 60% of questions correctly, which is worse than no chatbot at all because now customers don’t trust any of its answers. It looks like a content system that produces grammatically perfect posts that no one reads because they say nothing a reader couldn’t get from typing the same topic into ChatGPT themselves. It looks like an analytics dashboard that tracks everything and reveals nothing, because nobody defined what decisions the data was supposed to inform.
This is the space where most of the 74% lives. Not catastrophic failure. Expensive adequacy. Tools that work well enough to avoid being cancelled but not well enough to justify their cost.
The pattern underneath all of it – every case study, every failure analysis, every piece of independent research – is the same: someone skipped the design phase. They went from “we should use AI” directly to “build me a thing,” without the intermediate step of defining what “working” looks like in specific, measurable, workflow-integrated terms.
The industry taught this. For a decade, the agency model was built on selling first and figuring out delivery later. You could sell SEO, then outsource the actual work to a team in the Philippines. The client rarely knew the difference.
AI doesn’t work that way. You can’t outsource architectural thinking. A prompt doesn’t get better because someone with a nicer title types it. And the gap between “looks like it works” and “actually works” is visible to every client within about three months. That’s when the retention conversation starts.
What "being good at AI" actually means
Back to Aalto University’s overconfidence finding. The researchers discovered that higher AI literacy correlates with more overestimation, not less. Knowing more about AI makes you worse at judging your own output. The reason is cognitive offloading – the better you understand what AI can do, the more likely you are to assume it did do it correctly in this specific instance.
Which means “getting better at AI” in the way most people pursue it – learning more tools, collecting more prompts, watching more tutorials – can actually make the problem worse. You become more fluent in giving instructions without becoming more rigorous in evaluating outcomes.
The minority who succeed aren’t better prompters. They’re better thinkers. They define the outcome before selecting the tool. They map the workflow before writing the first prompt. They build feedback loops that catch failure before it reaches a customer. They document what works so it’s repeatable, not a one-off accident.
It’s a design discipline, not a technical one. And right now, almost nobody is teaching it – because it doesn’t fit in a tweet thread, it can’t be reduced to a prompt template, and it requires admitting that the hard part of AI was never the AI.
The Aalto researchers called it cognitive offloading. The MIT team called it a learning gap. RAND called it miscommunication about the problem. BCG wrote it into a prescription: 10-20-70, with the technology as the smallest slice. McKinsey modelled it: the same technology, deployed with design and without it, producing outcomes an order of magnitude apart.
They’re all describing the same thing. And now that you know the pattern, you can’t unsee it in your own work – that moment where confidence in the tool replaced clarity about the outcome. Where fluency felt like competence.
The people who actually have the skills aren’t the ones who know the most about AI. Aalto’s data says knowing more makes your self-assessment worse. They’re the ones who pause long enough to ask what they’re building, and why, before they start. Everyone else already thinks they’re doing that. That’s the gap.
Your marketing, looked at properly
Thirty minutes on your current setup — what’s working, what’s quietly leaking budget, and what I’d fix first. You’ll leave with a clearer picture whether we work together or not.
Got something specific bugging you? Flag it when you book and I’ll have it looked at before we talk.
