Buy looks cheaper than build at the sticker. The vendor charges $50K to $200K annually for a capability the build alternative would cost $400K to $1.2M to stand up. The buy decision looks easy. It is not. Buy’s true cost; the actual annual run-rate the org pays once the capability is in production; is typically 3 to 5x the sticker, and the gap is structural rather than incidental. Integration work, workload-specific eval suites, data plumbing, on-call rotation, observability extension, and the engineering capacity to operate the vendor relationship are the load-bearing costs that the sticker does not include and the procurement framework rarely surfaces. The gap is not a sign that buy is the wrong choice; buy is often still cheaper than build at the true TCO. The gap is a sign that build-versus-buy comparisons made on sticker pricing produce systematically wrong arithmetic, and the orgs that win at AI sourcing are the ones whose comparison includes the full integration tax. This piece names the load-bearing costs, gives the multipliers for each, and provides the worksheet for computing true TCO before buy decisions are signed.
This sharpens one edge of the AI build-vs-buy-vs-hire decision matrix for 2026. The matrix’s principles say what to source how; this piece operationalizes the cost arithmetic that buy decisions need to pass to be defensible at renewal.
Why the sticker is misleading
The vendor’s sticker is the price of API access or seat licenses. It does not include the work the org has to do to make those useful in production. The gap is structural: vendors price against capability access; orgs pay for capability use, which requires more than access.
The gap manifests in two patterns. Pattern 1: small initial buy, no integration tax surfaced. Org signs a $50K vendor contract; six months later actual run-rate is $200K; the $50K plus engineering, eval suite work, data integration, on-call. CFO sees only the original $50K because the rest is buried in engineering payroll. Pattern 2: vendor wins the bake-off by ignoring its own integration tax. Build budget includes integration costs; buy budget is just the vendor quote. Buy wins on apparent cost; six months later the integration tax surfaces.
Both produce the same outcome: buy looks cheaper than it is, build looks more expensive than it is, the org under-builds capabilities it should have built and over-buys capabilities with high integration tax. The fix: make the integration tax legible at decision time.
Cost component 1: integration engineering
The largest of the six and most predictable. Most vendor capability requires integration: authentication wiring, rate limiting per user, prompt template management, retrieval logic, error handling, fallback paths, frontend integration.
Typical annualized costs: RAG capability $150K to $300K (3-6 engineering weeks initial plus 0.5-1 FTE ongoing); agent capability $300K to $600K (6-12 engineering weeks plus 1-2 FTE); foundation model API $50K to $150K (1-3 engineering weeks plus 0.25-0.5 FTE). Integration cost does not scale with vendor pricing; a $30K vendor contract has approximately the same integration tax as a $300K contract because the work is determined by product complexity, not vendor pricing. Implication: low-sticker buys often have the worst sticker-to-TCO ratios.
Cost component 2: workload-specific eval suites
The most under-budgeted. Vendor general scores apply to general benchmarks; production quality on the org’s workload requires workload-specific eval suites testing the vendor against the actual distribution. Per the matrix’s fifth principle, eval suites are build or hire; rarely buy. The vendor cannot write eval suites for the org’s workload, and the vendor’s incentive to find quality issues in their own product is misaligned with the org’s incentive to find them.
A production-grade eval suite typically costs $100K to $250K annualized (2-4 engineering weeks initial plus 0.25-0.5 FTE ongoing). The cost is approximately the same whether the underlying capability is buy or build; the eval suite is a moat layer per the AI plumbing-vs-moat piece. Orgs that skip the workload-specific suite ship vendors that look good in the bake-off and produce visible quality failures within 3 to 6 months.
Cost component 3: data plumbing
The most variable. Some capabilities need almost no plumbing; others require substantial data flows; getting the org’s data into the vendor’s index, keeping it synced, handling deletes and updates, managing schema evolution, access control. RAG plumbing is typically $100K to $200K annualized; agent capabilities consuming internal tooling run $250K to $600K because most tool integration is a plumbing problem; embedding services run $150K to $250K.
Data plumbing is also where data residency, compliance, and governance overhead show up. For regulated industries; financial services, healthcare, regulated SaaS; this category can dominate true TCO, often 1.5x to 2x what an unregulated org would pay for the same capability.
Cost component 4: on-call and incident response
The most overlooked. Once a capability is in production, someone responds when it breaks. The vendor handles the vendor’s service; the org handles everything downstream; quality regressions on the workload, integration failures, data sync issues, customer-visible incidents.
On-call cost is typically 0.25 to 0.75 FTE-equivalent, $75K to $225K annualized. AI capabilities have novel failure modes (eval drift, hallucination spikes, vendor API behavior changes) that legacy on-call playbooks do not cover. The cost grows roughly linearly with capability count and can be partially amortized through a hub structure (per the hub-and-spoke org piece) but not eliminated.
Cost component 5: observability extension
Vendor capabilities ship with vendor observability; vendor traces, vendor dashboards. The org needs more: end-to-end observability across the product, cost attribution by team, eval-score-in-production monitoring, prompt/response logging in the data warehouse. Observability extension is typically $50K to $150K per capability annualized (2-4 engineering weeks plus 0.1-0.3 FTE).
A hub team running centralized observability (per the hub-and-spoke pattern) absorbs the leverage layer; the hub itself costs $400K to $1M annually but amortizes across most AI capability. For orgs with 10+ capabilities the hub is cheaper; for orgs with fewer capabilities the per-capability path is cheaper.
Cost component 6: vendor relationship operations
The smallest but persistent. Most vendor relationship has operational overhead: contract management, renewal negotiation, vendor health monitoring, escalation handling, vendor-specific procurement work. For a single vendor this is small; 0.05 to 0.15 FTE; but it scales linearly with vendor count.
An org running 6 AI vendor relationships pays approximately $50K to $150K annually in vendor operations cost. An org running 20 vendor relationships pays $150K to $500K. The cost is one of the inputs that drives vendor consolidation as orgs mature: cutting from 20 vendors to 8 saves vendor operations cost as well as license cost.
The TCO worksheet
For each candidate vendor capability, compute:
Sticker. Vendor’s annual quote.
Integration engineering. From the per-capability table above. Annualized.
Eval suite cost. $100K to $250K, with workload-specific adjustment.
Data plumbing cost. From the per-capability table, with regulatory adjustment if applicable.
On-call cost. $75K to $225K per capability for orgs without a hub; less for orgs with a hub absorbing the rotation.
Observability extension. $50K to $150K per capability without a hub; less with one.
Vendor relationship operations. $25K to $50K per vendor.
True TCO. Sum of the above.
Sticker-to-TCO multiplier. TCO divided by sticker.
For most AI buy decisions in 2026 the multiplier comes out to between 3x and 5x. Capabilities with low sticker and high integration complexity (RAG, agent capabilities) skew toward 5x. Capabilities with high sticker and low integration complexity (foundation model APIs at scale) skew toward 1.5x to 2x because the sticker dominates.
The true TCO is the number that should be compared to build cost in the build-versus-buy decision. Comparing build cost to sticker is the systematic error that drives most under-build outcomes.
What this changes about build-vs-buy
The integration cost gap does not flip the decision in most cases; buy is still cheaper than build for low-volume, low-differentiation capabilities. But it changes the threshold at which build becomes correct.
Concretely, a capability where the build cost is $700K annually compared against a vendor sticker of $100K looks like an obvious buy on sticker. Compared against the vendor’s true TCO of $400K to $500K, the math tightens substantially: build is now within 1.4x to 1.75x of buy. At that ratio, considerations beyond pure cost (lock-in, differentiation, capacity for moat work) often tip the decision toward build.
The cost crossover for the buy-then-build progression (per the buy-then-build progression piece) should also be computed against true TCO rather than sticker. Many orgs miss the crossover because they track the sticker only; the actual crossover happened 12 to 18 months earlier when the integration tax pushed total cost above the in-house alternative.
Frequently asked questions
Why does the integration tax not show up in vendor case studies?
Vendor case studies report what the customer paid the vendor, not what the customer paid total. The integration tax lives in the customer’s engineering payroll, which the vendor does not have visibility into and would not publish if they did. The case studies are not deceptive; they are scoped to vendor cost. The buyer’s job is to add the integration tax outside the vendor’s reported numbers.
How do we estimate integration cost before signing?
Three diagnostics. (1) How many existing customers have this capability in production at our volume? Talk to two or three of them; ask specifically about their integration timeline and the engineering capacity they had to allocate. (2) Build a small proof-of-concept against the vendor’s API; the time to get the POC working is a rough proxy for the integration scale. (3) Map the data flows the capability requires; each data flow is an integration cost line.
What if the vendor includes professional services?
Vendor professional services can offset integration cost but rarely eliminate it. The professional services scope typically covers initial setup; ongoing maintenance, eval suite construction, on-call, and observability extension are still the org’s job. Discount the integration tax by the value of professional services received, but do not zero it out.
How does this compare to legacy SaaS integration costs?
Legacy SaaS integration costs are typically 1.5x to 2.5x sticker, lower than AI’s 3x to 5x. The difference is driven primarily by the eval suite component (legacy SaaS has no analog) and the workload-specific data plumbing (which is more elaborate for AI than for legacy SaaS). The AI sticker-to-TCO multiplier is structurally higher because AI capabilities have more places for the org to do work that the vendor cannot do for them.
What about open-source AI tooling versus commercial?
Open-source has no sticker but has many the same integration tax components; sometimes higher because there is no vendor support. Total TCO can be similar; the components shift from license to engineering capacity. Open-source is not free; it is a different cost mix.
How does the hub-and-spoke org structure affect TCO?
A central hub absorbs the leverage components (observability extension, eval harness, on-call rotation foundation, vendor management overhead) across many AI capabilities. The per-capability TCO falls; the hub itself is a fixed cost. For orgs with 5+ AI capabilities the hub usually reduces total TCO; for orgs with 1 to 2 capabilities the hub is overhead. The hub is appropriate when the leverage to capture is bigger than the hub’s fixed cost.
Why does the eval suite cost not scale with capability complexity?
Because the workload-specific eval suite work is determined by the workload’s coverage requirements, not by the underlying capability’s complexity. A simple capability operating in a complex workload has a complex eval suite; a complex capability operating in a simple workload may have a simple eval suite. The cost is a workload property more than a capability property.
How does this interact with the AI capability ladder?
The capability ladder provides default verbs based on 2026 conditions. The integration cost gap argument is upstream of the ladder: the ladder’s default verbs are correct given an honest TCO comparison, and orgs whose comparison undercounts the integration tax will land on the wrong verb even when applying the ladder correctly. The full ladder is in the AI capability ladder piece.
Key takeaways
- The vendor sticker for AI capabilities is typically 1/3 to 1/5 of the true TCO; the gap is structural and predictable, not incidental.
- Six load-bearing cost components show up post-sticker: integration engineering, workload-specific eval suites, data plumbing, on-call rotation, observability extension, and vendor relationship operations.
- Per-capability cost components total $500K to $1.5M annually for moderate-complexity capabilities; the multiplier on sticker is 3x to 5x for typical AI buys.
- The integration cost gap does not flip most build-vs-buy decisions but tightens the math substantially; capabilities that look like obvious buys on sticker often become close calls on true TCO.
- The cost-crossover for buy-then-build progressions should be computed against true TCO, not sticker; orgs that track sticker only miss the crossover by 12 to 18 months.
The orgs that procure AI well in 2026 do not pay the sticker as their decision input; they pay the true TCO, including the integration tax that the sticker does not advertise. The discipline is structural; write the TCO worksheet at most buy decision, review it at most renewal, and use true TCO as the comparison number against build alternatives. The orgs that skip the discipline end up with a buy stack that costs 3 to 5x what their procurement framework predicted, and the gap surfaces 12 to 18 months later as a budget surprise that nobody can quite explain.
Arthur Wandzel