Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 13 min read

The AI Capability Gap Analysis Method We Use With New Clients

The AI Capability Gap Analysis Method We Use With New Clients

Most new client engagement starts the same way: the client believes they have an AI roadmap problem, and we discover that they have a capability inventory problem first. The roadmap they describe is a list of features they want to build; the inventory they do not have is the list of capabilities they already run, the sourcing decisions behind each, the eval discipline applied to each, and the cost trajectory of each. Without the inventory, the roadmap is a wish list pretending to be a plan, and the engagement begins with two parties solving different problems. The fix is the four-phase capability gap analysis we run in the first 21 days of most new client engagement: capability inventory, sourcing assessment, eval discipline check, prioritized roadmap. Each phase produces a specific artifact; the four artifacts together are a defensible AI strategy in a way that no individual one of them is alone. This piece names the four phases, the artifacts each produces, and the failure modes the method is designed to surface in the first three weeks rather than the third quarter.

This puts the AI build-vs-buy-vs-hire decision matrix for 2026 into practice. The matrix’s seventh principle is that most sourcing decision is re-litigated quarterly; the gap analysis is the artifact that bootstraps the cadence for an organization that has not run it before.

Why a method, not an audit

A gap analysis is not an audit. An audit measures current state against a known standard and reports compliance. A gap analysis measures current state against the state the organization needs to reach and produces a prioritized path between them.

Two consequences follow. The standard is the client’s, not ours; the destination is shaped by the client’s strategic priorities, constraints, and risk tolerance, and the roadmap is owned by the client. The output is a path, not a verdict; “you are at X today, need to be at Y in twelve months, and here are the four sequenced moves that get you there.”

The method is structured because unstructured gap analyses produce wish lists. The structure; four phases, each with a fixed artifact, time-box, and inputs; makes the analysis repeatable across engagements and comparable across quarters.

Phase 1: capability inventory

Phase 1 produces the capability list. It runs in week 1 of the engagement.

The work is interview-based. We meet with engineering, product, business unit, and operations leaders, ask “what AI runs in your area today?” and triangulate. Answers diverge in interesting ways: the feature engineering treats as primary is one product has rarely used; the workflow operations relies on heavily was unmentioned by engineering.

The phase produces three artifacts.

The capability list. Most AI capability the organization runs in production; typically 8-15 entries for mid-market. The altitude is “customer support response drafting” or “documentation search,” larger than features, smaller than functional areas, with one namable owner. We cross-reference each with the capability portfolio framing.

The shadow-AI map. AI usage outside the formal list; individuals using ChatGPT for ad-hoc work, marketing using copy tools without procurement, support agents using extensions on their own laptops. Not bad per se but invisible to governance.

The orphaned-capability list. Capabilities in production with no clear owner; typically from acquisitions, departures, or lost sponsorship. Orphans are highest-priority remediation in phase 4.

A typical first-time inventory finds 20-40% more capabilities than the client’s mental model included; the gap is mostly orphans and shadow AI.

Phase 2: sourcing assessment

Phase 2 assesses how each capability is sourced and whether the sourcing matches the matrix’s principles. It runs in week 2.

For each capability on the inventory, we determine current sourcing (build, buy, hire, compose), document the primary vendor or in-house team, and stress-test the sourcing decision against the matrix.

The stress test asks four questions per capability.

Is the verb explicit? Has someone deliberately chosen build, buy, hire, or compose, or did the verb emerge by default from whoever moved first? Default-verb capabilities are the most often miss-sourced.

Does the verb match the principle? The matrix’s principles 3 (foundation models are buy), 4 (agent orchestration is build), and 5 (eval infrastructure is build or hire, rarely buy) constrain which verb is right for which kind of capability. We flag capabilities whose current verb violates a principle.

Is the sourcing renewable? A capability sourced from a vendor whose contract is renewing in the next 90 days is in a different state than one with 18 months left. Renewable capabilities are more decision-ripe in the roadmap.

Is the sourcing reversible? Capabilities with low exit cost are more readily re-sourced; capabilities with high exit cost are constrained even when the principle would prefer a different verb.

The sourcing assessment phase produces the sourcing scorecard; a per-capability rating of sourcing alignment with the matrix, with notes on which principle is violated where. The scorecard is the input to phase 4’s roadmap prioritization.

Phase 3: eval discipline check

Phase 3 measures the eval discipline applied to each capability. It runs in week 3 and is consistently the most uncomfortable phase for the client.

For each capability, we ask:

  • Does an eval set exist?
  • If so, where does it live (in-house repo, vendor product, scattered across notebooks)?
  • How recently was the eval set updated?
  • What is the most recent pass rate?
  • Is the pass rate trended over time?
  • Is the eval set known to cover the long tail or only common cases?
  • Has anyone validated that the eval scoring agrees with human judgment?

The answers fall into four discipline tiers.

Tier 4; disciplined. Eval set in repo, updated quarterly, trended, scoring validated against human judgment, long tail tested. Rare on first engagement.

Tier 3; partial. Eval set exists, lives in a defensible place, pass rate computed but not trended; scoring plausible but not validated. About 30% of capabilities.

Tier 2; vibes. Informal eval; a few test cases in a notebook, “looks good to QA” sign-off. About 35% of capabilities.

Tier 1; none. No eval set. Quality judged by complaints and gut feel. About 25% of capabilities, including ones the client believes are working well.

The check surfaces the highest-impact remediation opportunities. The matrix’s fifth principle; eval infrastructure must be built or hired, rarely bought; is most often violated at tiers 1 and 2. The role that owns this discipline at the executive level is in why AI agencies need a chief evaluation officer before a chief AI officer.

Output: per-capability tier plus rolled-up portfolio average. Organizations below 2.5 are routinely surprised by regressions and miscalibrated about where AI investment pays off.

Phase 4: prioritized roadmap

Phase 4 produces the roadmap. It runs as a synthesis of phases 1-3 in the final 3-5 days of the engagement.

The roadmap has three layers.

Layer 1; the immediate (next 90 days). Orphan remediation, shadow-AI normalization, eval bootstrapping for tier 1 capabilities. These are the moves that pay off fastest and that constitute the table stakes for any defensible AI strategy. Typical immediate-layer items: name owners for orphan capabilities, bring shadow-AI usage under procurement governance, write minimum-viable evals for the three highest-traffic capabilities at tier 1.

Layer 2; the structural (next 6 months). Sourcing realignments, eval discipline upgrades, contract renegotiations. These are the moves that change the long-run trajectory of the AI portfolio. Typical structural items: switch a violator capability from buy to compose, upgrade two tier-2 capabilities to tier 3 by writing structured eval sets, renegotiate a vendor contract that scored high on lock-in.

Layer 3; the strategic (next 12-18 months). Capability additions, capability retirements, organizational design changes. These are the moves that change what the company does, not just how. Typical strategic items: add a new capability that the matrix would justify but the team has not yet started, retire a capability that has been outpaced by foundation model improvements, hire a chief evaluation officer if portfolio eval discipline is below tier 2.5.

The roadmap is sequenced. Layer 1 must precede layer 2 must precede layer 3, because the artifacts produced in earlier layers are inputs to later layers. We do not let clients jump to layer 3 without doing layer 1 first; the strategic moves cost more and fail more when the organization has not done the structural and immediate work.

The roadmap includes named owners, time-boxes, budget envelopes, and measurable success criteria for each item; the same structure as the kept/changed/retired output of a steady-state portfolio review. The gap analysis is the bootstrap; the portfolio review is the steady state.

Failure modes the method surfaces early

The four-phase method surfaces specific failure modes in the first 21 days rather than at the decline of a quarter or a year.

Roadmap-without-inventory. The client arrives with a roadmap of features to add and discovers no inventory of features they already run. The roadmap is rewritten on a foundation that exists.

Eval debt masquerading as capability strength. A capability the client believes is high-performing has no eval set. The “high-performing” judgment was absence of complaints, not presence of evidence.

Vendor lock-in normalized as strategic dependence. A capability with severe lock-in scores has been mentally re-labeled as strategic dependence. Phase 2 surfaces the difference.

Shadow-AI as governance gap. Significant AI work happens outside the formal list. Phase 1 inventories it; phase 4 normalizes it.

Owner-less capabilities. The discipline of giving most capability one named owner is the cheapest governance improvement available.

What the method does not do

The method does not replace strategic AI thinking; it grounds it. The strategic work; what should we be building, what differentiates us, where is the moat; happens after the gap analysis, on the foundation the analysis produces. A client who arrives expecting strategy and gets inventory may feel underserved in week 1 and decisively well-served by week 3.

The method does not validate vendor selection. New vendor selection happens in a separate evaluation forum; the gap analysis points at where new selection is needed but does not itself run the selection.

The method does not produce specific implementation plans. Layer 1, 2, 3 items are sized and sequenced but not project-planned. The implementation planning happens in the engagement weeks after the gap analysis is delivered.

The method does not work without engineering and product participation. A gap analysis built only on procurement and finance interviews misses the capability inventory and the eval discipline checks. The method requires cross-functional participation; refusing to provide it is itself a finding.

Frequently asked questions

How long does the full method take?

Twenty-one days end-to-end with engaged client participation. Phase 1 is one week, phase 2 is one week, phase 3 is one week, phase 4 is the final 3-5 days. Compressing below 21 days lowers fidelity; extending beyond 28 days lowers urgency.

Who at the client needs to participate?

A minimum of: VP/Director Engineering, VP/Director Product, CFO or finance director, head of procurement, head of AI (if exists), and 2-3 business-unit leaders. Without this participation the method does not work.

How does this compare to an AI maturity assessment?

A maturity assessment scores the organization against a fixed framework. The gap analysis includes a similar scoring inside phase 3 but goes further: it produces the prioritized path to a higher state, not just the score.

What if the client has fewer than five AI capabilities in production?

The method still runs but compresses. Phase 1 takes 2 days instead of a week; phase 4 emphasizes capability addition (layer 3) more than realignment (layer 2). The four phases still happen.

How does this connect to the lock-in audit and exit-cost model?

The lock-in audit and exit-cost model are inputs to phase 2’s sourcing assessment for buy-row capabilities. We run them inside phase 2 if the client does not already have them.

What does success look like at day 21?

Four artifacts delivered: capability list with shadow-AI and orphan annotations, sourcing scorecard, eval discipline scorecard, three-layer prioritized roadmap. The client owns the roadmap and can begin executing layer 1 immediately.

How often is the method re-run?

Once per engagement at start; thereafter the quarterly portfolio review replaces it. Re-running the full method makes sense after major organizational events; acquisition, restructuring, leadership change.

What is the most common surprise in the analysis?

The shadow-AI volume in phase 1. Clients consistently underestimate by 30-50% how much AI is being used outside formal procurement governance.

Can the client run the method themselves?

Yes, after they have seen it run once. The methodology is documented; the discipline is in the time-box and the cross-functional participation. Clients who run subsequent gap analyses themselves typically take longer the first time and faster most time after.

How does this connect to the matrix?

The matrix is the strategic frame; the gap analysis is the diagnostic that locates the client on the matrix and produces the move list to a better location. The two are designed to work together; using either alone misses the value the other provides.

Key takeaways

The capability gap analysis runs in four phases over 21 days: capability inventory (week 1), sourcing assessment (week 2), eval discipline check (week 3), prioritized roadmap (final 3-5 days). Each phase produces a specific artifact; the four artifacts together are a defensible AI strategy in a way no individual artifact is alone.

The method consistently surfaces five characteristic failure modes early: roadmap-without-inventory, eval debt masquerading as capability strength, vendor lock-in normalized as strategic dependence, shadow-AI as governance gap, and owner-less capabilities. Surfacing these in the first 21 days is the difference between an engagement that produces a defensible strategy and one that produces a wish list.

The roadmap is sequenced into three layers; immediate (next 90 days), structural (next 6 months), strategic (next 12-18 months). The sequencing is firm: clients who skip from layer 1 directly to layer 3 routinely produce strategic moves that fail because the organizational foundation is not in place.

The gap analysis is a bootstrap, not a steady state. Once delivered, the quarterly portfolio review takes over as the cadence that keeps the analysis current. The discipline of running the gap once and the review most quarter thereafter is what differentiates organizations that compound their AI sourcing decisions from organizations that drift into avoidable lock-in, eval debt, and shadow governance. The method is the entry point; the cadence is what makes it pay off.

Last Updated: Jun 25, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles