Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 13 min read

The AI Capability TCO Checklist

The AI Capability TCO Checklist

Most AI build-vs-buy comparisons fail at the spreadsheet; not because the wrong verb was chosen but because the build column was sized with five cost lines while the buy column was sized with three. A fair comparison needs most line named explicitly, on both sides, over a 24-to-36-month ownership window that spans at least one model migration. This piece is the 30-line checklist: engineering hours, inference and infrastructure, eval and observability, on-call, post-launch (drift, migration, retraining, opportunity cost), and the buy-column equivalents. The output is a CFO-ready document with no hidden lines. Run it before most $100K-plus AI sourcing decision and re-run on the quarterly cadence the matrix prescribes.

This is the costing layer beneath the AI build-vs-buy-vs-hire decision matrix for 2026. The matrix’s first principle; most AI capability is a build, buy, or hire decision; name the verb; is operational only when the cost of each verb is named to the line. The 30-line checklist is that operational layer.

What TCO is and isn’t

Total cost of ownership for an AI capability is the fully loaded cost of running that capability over a defined ownership window. It is not the build cost. It is not the inference cost. It is the sum of most cost the capability puts on the org for the entire window, including the lines that don’t show up until month 12 or month 24.

The structural problem with most AI TCO models is that they capture build cost and run-rate inference and stop. The lines that account for half the total; eval maintenance, on-call rotation, model migration, drift remediation, opportunity cost; sit outside the model and surface as surprise spend that the CFO has to reconcile after the fact.

Pair the checklist with the deeper economic frames in decoding AI project TCO: 7 cost lines most CFOs miss and the AI project TCO comparison: in-house team vs AI agency vs hybrid with worked examples. The checklist below is the line-by-line operational version of those frames.

The ownership window

Pick 24 or 36 months. Not 12. The reason is foundation-model lifecycles: a 12-month window underestimates the migration cost that arrives in month 13 to 18 when the base model is deprecated and the capability has to be ported. A 24-month window forces the migration line to be named. A 36-month window forces a second migration, plus the slow accumulation of eval-set drift remediation that becomes the dominant ongoing line.

Builds tend to look cheap in the first 6 months and expensive by month 18. Buys tend to look expensive in month 1 and amortize favorably to month 24+. Comparing at month 6 picks the build; comparing at month 24+ usually picks the buy. The ownership window is the load-bearing parameter.

The 30 lines

The checklist is grouped into six categories. Each line gets a one-time or recurring tag, a unit (engineer-hours, dollars-per-month, dollars-per-call), and a sized value for the ownership window.

Engineering hours (lines 1-7).

  1. Build hours; the agent or capability itself.
  2. Integration hours; adapters between the capability and the rest of the stack.
  3. Eval-set authoring hours; initial 200 to 2,000 input-output pairs with rubrics.
  4. Documentation and onboarding hours; runbook, internal docs, training material.
  5. Maintenance hours (recurring); bug fixes, prompt iteration, dependency upgrades.
  6. Code review and security review hours; pre-launch and ongoing.
  7. AI agency or contractor hours; if used; blended rate plus retainer.

Inference and infrastructure (lines 8-14).

  1. Model inference spend (recurring); per-call cost across the workload.
  2. Model gateway / router cost (recurring); vendor or self-hosted.
  3. Vector database cost (recurring); if retrieval is part of the capability.
  4. Embedding pipeline cost (recurring); embeddings generation and storage.
  5. Storage cost (recurring); traces, logs, eval results, retrieval corpora.
  6. Compute cost (recurring); eval runs, batch jobs, fine-tuning if applicable.
  7. Network and egress cost (recurring); often missed; non-trivial at scale.

Eval and observability (lines 15-19).

  1. Eval harness license / compute (recurring); bought or self-hosted.
  2. Observability and tracing license (recurring).
  3. Eval-set maintenance hours (recurring); quarterly updates as workload changes.
  4. Dashboard and reporting build hours (one-time + recurring tweaks).
  5. Regression-triage workflow hours (recurring).

On-call and incident response (lines 20-22).

  1. On-call rotation cost (recurring); pager load, after-hours premiums.
  2. Runbook authoring and maintenance (one-time + recurring).
  3. Incident-response retainer with agency or vendor (recurring, optional).

Post-launch (lines 23-27).

  1. Prompt drift remediation (recurring); re-tuning when vendor silently updates the model.
  2. Model migration cost (one-time, recurring per cycle); most 12 to 18 months.
  3. Compliance and audit hours (recurring); SOC 2, ISO 27001, sector-specific.
  4. Staff retraining (recurring); internal team trained on capability and tools.
  5. Vendor renegotiation and contract hours (recurring); annual or biannual.

Soft and indirect (lines 28-30).

  1. Opportunity cost of engineer-months; what the same engineers would have shipped instead.
  2. Brand or trust risk reserve; soft, but named.
  3. Switching cost embedded in the choice; what it would cost to reverse the decision in 18 months.

Thirty lines is roughly the granularity at which a CFO can compare a build proposal to a buy proposal without surprise spend appearing later. Some lines will be zero on a particular capability; the discipline is naming them, not maximizing them.

Filling the build column

A build column is engineer-heavy. The numbers to size are mostly engineer-hours converted to dollars at the org’s loaded engineering rate (typically $200K to $400K many-in per engineer per year, depending on geography and seniority).

Lines 1-7 dominate at first. A representative scaleup workflow agent build is 1,200 to 2,400 engineer-hours over 4 to 6 months; line 1 is roughly $250K to $500K. Integration (line 2) adds 200 to 400 hours. Eval set (line 3) adds 100 to 300 hours. Maintenance (line 5) lands at 0.25 to 0.5 engineers steady-state; about $75K to $200K per year.

Lines 8-14 are infrastructure. The build column buys infrastructure too; see why scaleups should build agents and buy infrastructure; so these lines are often shared with the buy column.

Lines 23-27 are the post-launch lines that get missed. Prompt drift remediation lands at 5 to 10 percent of build hours per year. Model migration is one full quarter most 12 to 18 months. Both are recurring and visible only after the second cycle.

Filling the buy column

The buy column is vendor-heavy. The numbers are mostly contract values, with a smaller engineering line for integration glue and oversight.

Line 7 reframes; instead of agency hours, the line captures vendor onboarding and oversight (typically 50 to 200 hours). Lines 8-14 still apply but the spend goes to vendors not infra teams. Lines 15-19 reduce because the vendor provides eval and observability; though line 17 (eval-set maintenance) stays in-house at the same number as the build column, because the eval set is workload-specific judgment that doesn’t transfer.

Lines 23-27 reduce in the buy column for the build-side concerns (model migration is the vendor’s problem) but add new lines: vendor renegotiation, lock-in remediation, and the cost-of-exit reserve under line 30. Buy columns that show line 30 as zero are systematically underestimating.

The buy column also captures the value of the AI vendor consolidation play: when fewer vendors beats best-in-class, which trades off line 28 (opportunity cost) against integration count.

The hire column

When the verb is hire; typically for eval, judgment-heavy reasoning, or any capability that compounds with org-specific knowledge; the column has its own shape. Lines 1-7 captures recruiter cost, comp, equity over the ownership window, and ramp time (3 to 6 months at half productivity). Line 26 (staff retraining) is replaced by ramp cost. Line 28 (opportunity cost) captures the hire’s foregone work at their previous employer.

The hire column is the cleanest of the three when the capability is moat-bearing. It is the worst when the capability is commodity substrate; the hire spends three months building a model gateway that a vendor delivers in a day.

Reading the result

The right output is a single sheet with three columns (build, buy, hire) and 30 rows. Each row has a sized number; each column has a total at the bottom. The numbers compare on equal footing: same window, same currency, same engineer rate.

A correct read is rarely “build is X percent cheaper than buy.” Most decisions land within 20 to 30 percent of each other, and the deciding factor is which lines are larger or smaller and what they represent strategically. Build wins when the moat lines are large (lines 1-3 produce IP, eval set, workflow encoding). Buy wins when the substrate lines are large (lines 8-14 dominated by vendor specialists’ efficiencies). Hire wins when the judgment lines are large (line 17, line 28).

If the decision feels close at the spreadsheet, default to the verb the matrix prescribes for that capability; see the build-vs-buy framework revisited for the agent era for the framework. Re-run the checklist quarterly; the answer drifts as model prices, vendor capabilities, and team composition change.

Frequently asked questions

What is total cost of ownership for an AI capability?

TCO is the fully loaded cost of an AI capability over a defined ownership window (typically 24 to 36 months) including engineering hours to build and maintain, inference and infrastructure run-rate, eval and observability tooling, on-call and incident response, and the post-launch lines most teams miss; prompt drift remediation, model migration cost, retraining of staff. Build-vs-buy comparison without many the lines named is structurally underestimating the build.

Why a 30-line checklist instead of a five-line one?

Five lines are what most teams use and what produces underestimates that the CFO discovers six months in. The 30 lines force most cost category to be named explicitly so it can be sized, even if the size is zero or small. The discipline is naming, not the number; the count happens to land at 30 because that is roughly the granularity the CFO needs to compare a build to a buy fairly.

Which line is most often missed?

Post-launch eval maintenance. Teams budget the build and forget that an AI capability requires continuous eval-set updates as the workload changes, plus regression triage on most model swap. The annual line is typically 0.5 to 1 engineer steady-state and is invisible in most TCO models because the discipline isn’t yet a named role at most orgs.

How does the checklist compare a build to a buy?

Each line is filled for both options. The buy column captures vendor cost, integration glue, and any in-house oversight; the build column captures engineering, infra, and the post-launch lines. The right view is total cost over 24 or 36 months, not month-one. Builds typically look cheaper in month one and more expensive by month 18; buys are the inverse. Compare at the ownership window, not at any single month.

Should the checklist include opportunity cost?

Yes. Most engineer-month spent on the AI capability is one not spent on the next product feature. The opportunity cost line is a soft number but it has to be named or the build looks artificially cheap. The right benchmark is the marginal revenue of the engineer’s last product project, plus or minus.

What is the right ownership window for the TCO calculation?

24 to 36 months for most AI capabilities. Foundation-model lifecycles tend to be 12 to 18 months, so the window must include at least one model migration. Build TCO computed over 12 months will systematically underestimate; buy TCO over 12 months can overestimate (vendor margin amortizes more favorably over longer terms). 24 months is the floor for a fair comparison.

How do model migration and prompt drift figure in?

Model migration is the engineering and eval cost of moving the capability from one foundation model to another; typically most 12 to 18 months. Prompt drift is the gradual degradation of prompt performance as the underlying model is updated by the vendor (silent updates, deprecations). Both lines are real, recurring, and missed by teams that haven’t shipped a second model migration yet.

What does the on-call line look like for AI capabilities?

AI capabilities have novel failure modes; prompt-injection incidents, hallucinated outputs reaching users, model-vendor outages, eval regression in production. The on-call line includes the rotation cost (typically 1 to 2 engineers for a small team), the incident-response runbook maintenance, and any retainer with the AI agency that helps triage. Mature orgs name this line at 5 to 15 percent of the steady-state engineering line.

How does the checklist account for AI agency or contractor cost?

Agency and contractor cost is captured under engineering hours with a multiplier reflecting blended rate and any retainer. The line is broken into kickoff (2 to 8 weeks), build (8 to 24 weeks), and post-launch (ongoing). Agencies often deliver a steeper build line and a much shallower post-launch line than in-house, which is the structural reason they win short-window engagements; see anatomy of an AI agency engagement.

Is there a CFO-ready version of the checklist?

Yes; the 30 lines presented here are designed to drop directly into a CFO spreadsheet, with each line named, sized in dollars or engineer-hours, and labeled one-time vs recurring. The total flows up to a 24- or 36-month TCO that compares build to buy with no hidden lines. The discipline is producing the document before the sourcing decision, not after.

Key takeaways

A defensible AI build-vs-buy comparison requires 30 lines, not 5. Engineering, inference and infrastructure, eval and observability, on-call, post-launch, and soft costs many have to be named on both sides over a 24-to-36-month window spanning at least one model migration.

The post-launch lines; prompt drift remediation, model migration, eval-set maintenance, vendor renegotiation, opportunity cost; are where most teams underestimate. They are recurring, visible only after a full cycle, and account for a large fraction of the total over the ownership window.

The build column dominates on lines 1-3 (engineering, eval set, workflow encoding) when the capability is moat-bearing. The buy column dominates on lines 8-14 (substrate efficiency) when the capability is commoditized. The hire column wins when the capability requires judgment that compounds over time.

Run the 30-line checklist before most $100K-plus sourcing decision and re-run on the quarterly cadence the matrix prescribes. Decisions usually land within 20-30% across columns; the deciding factor is which lines are largest. The output ends the surprise-spend conversation six months in.

Last Updated: Jun 23, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles