MIT’s Project NANDA found that internally built generative AI tools reach production successfully around a third as often as tools built through vendor partnerships, roughly 33% versus 67%. Combined with the study’s headline finding that 95% of enterprise generative AI pilots show no measurable profit-and-loss impact, the report reframes build versus buy as one of the biggest predictors of whether an AI investment pays off at all.
What did the MIT NANDA report actually measure?
The report, The GenAI Divide: State of AI in Business 2025, is based on around 150 interviews with business leaders, a survey of roughly 350 employees, and an analysis of about 300 public generative AI initiatives across industries (Fortune).
“Success” in the report means a deployment reaching production and demonstrating measurable workflow impact, not merely completing a pilot or demo.
That distinction matters: plenty of internal AI projects reach a working prototype and then stall before anyone downstream actually changes how they work, which is where the report draws its line between a pilot and a real deployment.
What is the actual gap between internal builds and vendor deployments?
Vendor-sourced tools and partnerships reached production successfully around 67% of the time, while internally built tools succeeded at roughly a third of that rate, close to 33%.
Read as a failure rate, that means internal builds fail around 67% of the time versus roughly 33% for vendor deployments, which is about twice as likely to fail, not three times.
The “one-third the success rate” framing and the “twice the failure rate” framing describe the same numbers but sound very different, and the gap between them is exactly the kind of thing that gets mangled as a stat travels from report to headline to LinkedIn post.
Why do vendor-built tools succeed more often than internal builds?
Vendors succeed more often because they have already paid the integration tuition that a first-time internal team is still paying.
A vendor that has shipped the same category of tool into hundreds of businesses has already hit the edge cases, the data-quality failures, and the workflow mismatches that a single internal team meets for the first time, on its own budget, on its own clock.
That is not a claim that internal engineers are less capable than vendor engineers. It is a claim about where the learning has already been amortised.
The report also points to where internal builds tend to break: not at the model, but at the handoff into an existing workflow.
A tool that technically works but does not fit how a team actually operates, its approval steps, its exception handling, its existing systems of record, gets quietly abandoned regardless of how capable the underlying model is.
This is the same failure mode Bedrock AI sees in diagnostic work: the build was sound, but nobody had mapped the workflow it was supposed to join.
| Dimension | Internal build | Vendor partnership |
|---|---|---|
| Reported success rate reaching production | ~33% | ~67% |
| Where the friction typically shows up | Workflow fit and integration, after the model works | Contract terms, customisation limits, data governance |
| Who pays the “integration tuition” | The business, for the first time | The vendor, already amortised across deployments |
| Most common failure mode | Technically working tool nobody adopts | Good tool, wrong scope or poor account fit |
| When it’s still the right call | Narrow, well-diagnosed workflow with a clear owner | Undiagnosed problem, or capability the business lacks in-house |
Does this mean businesses should always buy AI tools instead of building?
No. The data says vendor partnerships succeed more often on average, not that building is always the wrong call.
Internal builds still make sense for a narrowly scoped, well-diagnosed workflow with a named owner and a realistic maintenance plan, especially where the workflow is specific enough that no vendor tool fits it cleanly. What the data argues against is the default instinct to build first and diagnose the workflow later.
Most of the internal builds behind that 33% figure were not failed because the team wrote bad code; they failed because the team started writing code before anyone had mapped where the workflow would actually break.
It is also worth reading the finding with some scepticism about its edges.
The 95% and 67/33 figures come from self-reported interviews and a sample skewed toward large organisations, so “success” reflects what interviewees judged as impact rather than an independently audited outcome.
Treat the numbers as a strong directional signal, not a precise universal ratio: the underlying pattern, that unscoped internal builds fail more often than diagnosed, vendor-supported ones, shows up consistently enough across other industry reporting to be worth acting on.
What should a business do differently because of this data?
Treat the build-or-buy decision as a diagnostic output, not a starting assumption.
That means running the workflow diagnosis before recommending either path.
It means only recommending an internal build when three things are true. The workflow is well enough understood to specify exactly what “done” looks like.
There is a named owner who will maintain the tool after launch. And no vendor tool fits the specific constraint driving the project (data residency, an unusual system of record, a workflow too niche for a general product).
Where those three conditions are not met, a vendor partnership is very likely to reach production faster and stay there longer, and the diagnosis should say so before the business spends the budget finding out the hard way.
FAQ
What is Project NANDA? Project NANDA is a research initiative from MIT’s Media Lab studying enterprise AI adoption. Its 2025 report, The GenAI Divide: State of AI in Business, is based on interviews, an employee survey, and analysis of around 300 public generative AI deployments.
Is the 95% failure figure about pilots specifically, or all AI use? It refers to generative AI pilots and initiatives that failed to show measurable profit-and-loss impact within the study’s window. It is not a claim that 95% of all AI use, including established machine learning and automation, delivers no value.
Does the internal-build failure rate apply to small and mid-sized businesses too, or just large enterprises? The report’s sample leans toward larger organisations, so the exact ratio may not transfer precisely to smaller businesses. The underlying mechanism, unscoped builds outrunning workflow diagnosis, applies at any size; smaller businesses often feel it faster because there is no spare budget to absorb a failed build.
Should a business avoid building AI internally altogether? No. The data argues for diagnosing the workflow and naming an owner before deciding to build, not for avoiding internal builds entirely. Narrow, well-scoped internal tools with a clear owner remain a sound choice.
How recent is this data, and does it still hold? The report was published in mid-2025 and remains the most cited dataset on build-versus-buy outcomes in enterprise generative AI as of mid-2026. No comparably sized follow-up study has superseded its core finding.
Bedrock AI finds where AI will actually save your business time and money, before you spend anything on building. Book a free 15-minute call