Where to Start With AI: Score Your Workflows First

6 min read

Business capability mapping for AI diagnosis means scoring every business function against how much of its work is high-volume, structured and repeatable, then ranking those scores against the cost of the delay or error it currently produces. That score, not which department asked first or shouted loudest, is what should decide which workflow gets the first AI pilot.

What is business capability mapping, and why does it matter for AI diagnosis?

Business capability mapping is a decades-old enterprise architecture technique that lists everything a business does (its capabilities, such as “process a claim” or “onboard a supplier”) independent of who currently does it or which system supports it.

Orbus Software describes the map as a structured, visual representation organised into a hierarchy, built specifically so the same capability can be examined for maturity, cost, risk and strategic importance in one view (Orbus Software).

Ardoq makes the same point differently: the map lets business and IT teams “jointly develop AI roadmaps,” because it is the one artefact both sides already recognise (Ardoq).

The technique matters for AI diagnosis because it forces the question most AI projects skip: which capability, out of everything the business does, actually has an AI-shaped problem.

Without a map, that question gets answered by whoever has the most political capital to request a pilot. With a map, it gets answered by comparing every capability against the same criteria, which is the difference between an AI roadmap and an AI wish list.

Why does scoring capabilities by “loudest request” produce the wrong AI roadmap?

The loudest request is a volume signal about internal politics, not a signal about whether AI actually fits the workflow.

A department head who has recently been burned by a slow process, read a competitor’s AI press release, or sat through a vendor demo will ask for a pilot regardless of whether their workflow has the volume, structure or repeatability that makes AI viable there.

Sourcing candidate use cases this way means the businesses that get an AI pilot first are the ones with the most assertive stakeholders, not the ones with the biggest AI-shaped bottleneck.

This is the same failure mode enterprise architecture spent decades solving for IT investment generally, which is exactly why the classic capability heat map is the right tool to adapt.

A heat map exists to replace gut-feel prioritisation with a defensible, capability-anchored rationale, typically scoring every capability on the same one-to-five scale for strategic importance, maturity and performance so that leadership can see strength, progress and urgent gaps at a glance rather than relying on whoever pitched hardest in the last budget meeting (Acorn).

AI diagnosis needs the same discipline, with different scoring dimensions.

How do you score a capability for an AI-shaped bottleneck?

An AI-shaped bottleneck score replaces the classic heat map’s strategic-importance and maturity dimensions with four questions that predict whether AI, specifically, will fix the workflow rather than just automate it badly.

Score each capability from 1 (low) to 5 (high) on each dimension, then multiply volume by cost of delay to weight the result toward capabilities where the problem is both frequent and expensive.

DimensionWhat it measuresLow score (1-2) looks likeHigh score (4-5) looks like
VolumeHow often the underlying decision or task repeatsA handful of cases a month, each materially differentHundreds or thousands of near-identical cases a week
StructureHow consistent the inputs and decision logic areJudgement calls with no consistent inputs (contract negotiation, hiring)Structured inputs feeding a repeatable decision (claims triage, invoice coding)
Data readinessWhether the data the decision needs already exists in usable formData lives in people’s heads or unstructured email threadsData already sits in a system of record with consistent fields
Cost of delay or errorWhat the business loses while the bottleneck persistsDelay is an inconvenience with no measurable costDelay or error directly costs revenue, compliance exposure, or customer churn

A capability that scores high on volume and structure but low on data readiness is not a “no”: it is a diagnosis that the first project is a data cleanup, not a model.

A capability that scores high on cost of delay but low on volume and structure (a rare, high-stakes judgement call) is usually a poor AI fit regardless of how loudly it gets requested, because there is no repeatable pattern for a model to learn from.

This mirrors what process-intelligence practitioners describe when they prioritise automation targets by frequency multiplied by time cost multiplied by risk, rather than by department priority alone (Rebel Force).

How is this different from a classic EA capability heat map?

The classic capability heat map and the AI-bottleneck score share a structure (list every capability, score it consistently, compare across the whole business) but they answer different questions.

The heat map asks where the business should invest generally; the AI-bottleneck score asks specifically whether AI is the right tool for that investment, which most capability maps were never built to answer because they predate AI agents as an option at all.

SAP LeanIX’s business capability assessment approach scores capabilities on dimensions like process maturity, application fit and business criticality to decide where any technology investment should go next (LeanIX).

That is the right first pass for a portfolio review, but it will happily rank a low-volume, judgement-heavy capability above a high-volume structured one if the judgement-heavy one is more strategically critical.

An AI-bottleneck score run on top of that same map catches the cases where strategic importance and AI-fit diverge, which is precisely where businesses waste a pilot budget: funding the capability that matters most to the board, not the one a model can actually fix.

What does an AI-shaped bottleneck score look like in practice?

Take a mid-sized insurer with three candidate capabilities on the table: claims triage, underwriting appeals, and customer onboarding.

Claims triage scores 5 on volume (thousands of claims a month), 4 on structure (most claims follow a handful of patterns), 4 on data readiness (claims already sit in a claims management system), and 4 on cost of delay (slow triage directly delays payout and drives complaints).

Underwriting appeals score 2 on volume (a few dozen a month), 2 on structure (each appeal turns on case-specific judgement), and 3 on data readiness, despite scoring 5 on cost of delay because a mishandled appeal carries real regulatory exposure.

Customer onboarding scores 4 on volume and structure but only 2 on data readiness, because the required documents arrive as unstructured email attachments with no consistent format.

The scoring makes the diagnosis obvious without a single stakeholder interview: claims triage is the pilot, underwriting appeals is not an AI project at all (it needs a better escalation process, not a model), and customer onboarding is an AI project with a data-engineering prerequisite that has to happen first.

None of that ordering matches which team asked loudest; the underwriting team, protecting the highest-stakes capability, would have made the most compelling case in a stakeholder meeting and would have been wrong.

FAQ

Does business capability mapping for AI diagnosis replace a discovery workshop? No. The map tells you which capability to investigate first; the discovery workshop is where you confirm the score against what actually happens on the ground, because self-reported process maturity and observed system behaviour rarely match exactly.

Who should build the capability map? Whoever owns it, the map only works for AI diagnosis if the same person or team applies the AI-bottleneck score consistently across every capability. A map built by different people scoring different capabilities against different criteria produces a ranking that looks objective but is not comparable.

How often should the AI-bottleneck score be re-run? At minimum every time a new capability is proposed for a pilot, and ideally quarterly, because data readiness and volume both shift as other projects change what data exists and how much of a process is already automated.

What if two capabilities score identically? Break the tie on cost of delay or error, since that dimension most directly reflects what the business loses by not acting, and it is the dimension most resistant to political inflation compared with claimed strategic importance.

Can a capability with a low AI-bottleneck score still be worth automating? Yes, but with a different tool. A low-volume, low-structure capability might still benefit from simple automation or better process design; the score specifically flags whether AI, as opposed to automation generally, is the right answer.


Bedrock AI finds where AI will actually save your business time and money, before you spend anything on building. Book a free 15-minute call

Keep reading

More for owners

All articles
AI adoption Why AI Tools Get Abandoned: The Workflow-Fit Failure Most AI tools that get abandoned work technically. They fail on workflow fit. Here are the diagnosis questions that catch it before build starts. / 7 min Build vs buy Why Internal AI Builds Fail: The MIT NANDA Findings MIT's NANDA report found internally built AI tools reach production at roughly a third the rate of vendor-built ones. Here's what the data shows and why. / 5 min AI strategy The AI Vendor Proof of Concept: A Three-Week Framework How to test a vendor's AI tool on your own work in three weeks before committing budget, and what each week must prove. / 6 min
Book a free call