An AI opportunity assessment framework is a scoring method for ranking candidate use cases before any budget is committed. Each candidate is scored on five dimensions: business value, data readiness, workflow fit, risk and regulatory exposure, and adoption effort. The highest scorer gets funded as one pilot with a decision gate agreed in advance.
This post is about choosing what to fund. Proving the return afterwards is a separate job, covered in our guide to measuring AI ROI. Read that one once money has moved. Read this one while the shortlist is still a list.
My position, after a lot of these conversations: fund the use case that reaches a verifiable result fastest, not the one that looks best on slide four. Someone demos the impressive thing, the room gets excited, and nine months later there's a prototype no data can feed.
dimensions every candidate use case is scored against
the weighted score below which a use case does not go first
scorers per session, including someone who has queried the data
Why does the first AI use case decide whether there's a second?
The first AI project sets a company's internal reference point for whether AI works. A use case that produces a measurable result within one quarter usually earns a second budget cycle. One that produces a demo and no before-and-after comparison usually ends the program, whatever the technology could do.
Budgets don't get renewed on technical merit. They get renewed because whoever approved the last one can point at something specific and say it worked. You'll see the claim that most AI pilots never reach production in every vendor deck, and I won't defend a number nobody will show the working for. The mechanism behind it holds up though. Ambitious first use cases need more integrations, touch more teams, and outlast a budget cycle. They run out of political runway before technical road.
What should you score every AI use case on?
Score each candidate on five weighted dimensions: business value, data readiness, workflow fit, risk and regulatory exposure, and adoption effort. Each is scored 1 to 5 against fixed anchors, multiplied by its weight, then summed into one weighted score out of 5. Candidates rank against each other, not an abstract standard.
| Dimension | What you score | 1-5 anchors | Weight |
|---|---|---|---|
| Business value | Size of the one-quarter gain, and whether it's provable | 1 = nobody can name the number that moves. 3 = plausible saving, no agreed figure. 5 = named metric, current value, agreed target | 25% |
| Data readiness | Whether the data exists, is reachable at decision time, and is complete | 1 = in people's heads or scanned PDFs. 3 = in a system, needs manual export. 5 = API-queryable, refreshed daily, fields populated | 25% |
| Workflow fit | How repetitive, rule-bound and stable the task is | 1 = judgment-heavy, rules rewritten monthly. 3 = stable rules, heavy exceptions. 5 = high volume, stable rules, under one in five exceptions | 20% |
| Risk and regulatory exposure | What a wrong output costs, and which rules apply | 1 = regulated data, no clear legal basis. 3 = internal, sensitive data, review in place. 5 = internal, non-sensitive, human approves before anything leaves | 15% |
| Adoption effort | How much must change for the people using it | 1 = new tool, new process, three teams, no sponsor. 3 = one team, some retraining. 5 = output lands in a tool they already open | 15% |
Two rules make the weights work. A floor rule: any dimension scoring 1 goes on hold regardless of total, because one broken dimension kills projects, not a mediocre average. And a threshold, below 3.5 weighted it doesn't go first.
Business value is capped at 25% deliberately. Everyone over-weights it in their head already, so handing it 40% formalizes the bias you're removing. Data readiness earns an equal share because it's what turns a confident yes into six weeks of unbudgeted discovery.
Is this workflow ready to be automated at all?
A workflow is ready for automation when it runs often enough to matter, follows rules that have held steady, throws a manageable share of exceptions, exposes its data programmatically, and has a clear answer on whether a human must approve it. Failing one isn't fatal. Failing three means you're funding a research project.
Volume. Count how often the task runs per week, not per year. At eleven times a month, recovered time won't cover the build. Our rule of thumb: a few hundred repetitions a quarter.
Rule stability. Ask when the rules last changed. If the policy was rewritten twice in six months, you're automating a moving target and maintenance will outgrow the saving.
Exception rate. Ask what share of cases leave the standard path. Past roughly a quarter, you've stopped building an automation and started building an exception-handling system, which costs more and pays back worse.
Data access. Existing and reachable aren't the same thing. A warehouse refreshed overnight can't support a decision made at 10am. Check the API, permissions and refresh interval before anyone scores this a 5.
Human approval. If a person must sign off, fine, but it changes what you're buying: drafting time rather than decision time. Say so in the value model. Required approval also lifts the risk score, which is why supervised use cases make better first picks.
Why does "remove one hour from one role" beat "replace the whole process"?
Narrow use cases score better because they reach a verifiable result faster. A single-role, single-task automation has one owner, one baseline number, and one clean before-and-after. A whole-process replacement spans several owners and systems, offers no clear comparison point, and fails completely when one assumption turns out wrong.
"Replace the claims process" sounds like leadership. "Cut first-pass document review from 50 minutes to 12 for the four people who do it" sounds small, and it's the one I'd fund. It carries a number before you start. Someone's Tuesday visibly changes.
A narrow win also tells you what your data can support, which sharpens the next round. The counter-argument is fair: some valuable problems are genuinely large. Fund those second.
Which use cases should be disqualified before you score them?
Three conditions disqualify a use case before scoring begins. There's no named owner whose work actually changes. There's no baseline measurement of current performance. Or the data required can't legally be used for this purpose. Treat these as pass/fail gates rather than low scores, because no weighting compensates for them.
No owner. If you can't name the individual whose week gets different, you have a theme, not a use case. Themes don't get adopted. "Operations wants this" is not an owner.
No baseline. Without a current number you can't prove improvement, and the pilot ends in an argument. Spend two weeks measuring before ten weeks building. Setting that baseline properly sits in the AI ROI guide, not here.
Data you can't legally use. Consent scope, contractual limits on third-party data, residency rules, personal data collected for another purpose. Get it in writing from whoever owns compliance, at scoring time. Finding out in week seven is the most avoidable way to burn a budget.
Who runs the scoring session, and how do you stop the loudest person deciding?
An effective session has four to six participants: the process owner who does the work, the executive accountable for the metric, an engineer who has queried the relevant data, and a compliance representative when regulated data is involved. Scores are submitted independently before discussion. A facilitator runs the session and scores nothing.
Mechanics matter more than the guest list. Everyone scores all five dimensions silently in a sheet nobody reads until it closes. Reveal at once. Then discuss only cells where scores differ by 2 or more, usually two or three out of thirty.
One rule stops seniority from deciding: the data readiness score belongs to the engineer. If the person who ran the query says 2, it's 2. Executives can argue about business value, which is their call. Make everyone write a six-word reason beside each score.
How do you turn a scorecard into a funded pilot with a decision gate?
Fund the top-scoring use case only. Before the build starts, write down four things: the metric being moved, its baseline value, the threshold that counts as success, and the date the decision gets made. Release budget in two tranches, the second conditional on that gate. Then continue, expand, or stop against the number.
A gate with a date separates a pilot from a project that quietly never ends. Set the date up front and hold it even when the result is unflattering. A pilot stopped on schedule buys credibility for the next round.
Only after the use case is chosen does the build question arrive: template, managed API, or custom, which we compare in the build-vs-buy breakdown for AI agents. Budget sizing is separate, covered in our AI app development cost guide. Run these in the wrong order and you'll pick an architecture before knowing what it's for.
Want outside eyes before the money moves? Our AI consulting practice runs this session with client teams, and our delivery model puts a working build in front of users in about two weeks. Book a 30-minute call with your candidate list.
What should you do with the use cases that scored badly?
Low scorers go into one of four buckets: re-scope smaller and score again, fix the blocking dimension and revisit on a set date, park with a documented trigger condition, or drop. Keep them in a visible register with scores and review dates, and re-score quarterly as rules and data change.
Re-scoping is the valuable bucket. Plenty of use cases score 2.9 as written and 4.1 once cut to the single step carrying the volume. Before discarding anything, ask what its smallest version would score.
The fix bucket usually hides a data problem with a known remedy: an export that should be an API, an optional field that should be mandatory. Parking needs a trigger rather than a vague "later", so write the condition down.
What does the scoring process look like, step by step?
- Collect every candidate use case, without filtering for feasibility yet.
- Apply the three disqualifiers: no owner, no baseline, no legal basis.
- Write each survivor as one sentence: role, task, unit of time or money.
- Run the readiness test: volume, rule stability, exception rate, data access, human approval.
- Assemble four to six scorers, including someone who has queried the data.
- Score all five dimensions independently and silently, with a six-word reason each.
- Reveal simultaneously and discuss only cells differing by 2 or more.
- Apply the floor rule: any dimension at 1 goes on hold.
- Calculate weighted scores and rank. Drop anything under 3.5 from first place.
- Sanity-check the leader: could this show a verifiable result inside one quarter?
- Define the gate before the build starts: metric, baseline, threshold, date.
- File everything else in the register with a bucket and a review date.
Frequently Asked Questions
How long should an AI opportunity assessment take?
For six to ten candidates, budget about two weeks. Most of that is gathering baselines and checking data access, not scoring. The session itself runs in 90 minutes.
What if two use cases score within a few points of each other?
Break the tie on time-to-verifiable-result rather than total value. Still level? Take the one with the more engaged owner, because adoption is the dimension people underestimate.
Should generative AI use cases be scored differently from predictive ones?
The dimensions hold, but the anchors shift. Generative use cases usually score higher on data readiness, since they need no labeled history, and lower on risk, because output quality varies per request.
What if leadership has already chosen the use case?
Score it anyway, alongside two or three alternatives, using the same anchors. A scorecard moves the conversation from opinion to evidence. If it scores 2 on data readiness, surface that now.
Do we need our data cleaned before we can score?
No. Scoring is how you find out what cleaning is required. A 2 on data readiness isn't a verdict, it's a cost line that belongs in the decision.
Is a low-risk internal use case always the right first pick?
Usually, though not automatically. Internal use cases are easier to measure and safer to get wrong. A customer-facing one can still lead when there's a real approval step and a baseline you already track.
Have a project in mind? Let's scope it together.
You get a named team, written estimates, full code and IP ownership, and 48-hour response times. CMMI Level 5 certified. 700+ projects delivered across the UK, US, UAE, and Australia.