Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 13 min read

The AI Build-vs-Buy Decision Graveyard: 8 Wrong Calls and How to Spot Them

The AI Build-vs-Buy Decision Graveyard: 8 Wrong Calls and How to Spot Them

Eight build-vs-buy wrong calls show up across post-mortems often enough to deserve names. Each has a recognizable shape; the stated rationale, the leading indicator visible months before failure, the wreckage, and the correct call. Naming the wrong call before signing avoids the graveyard. These eight account for most AI sourcing regret between 2024 and 2026. Treat the catalog as a checklist before any AI sourcing decision over $100K.

This catalog draws on the AI build-vs-buy-vs-hire decision matrix for 2026. It operationalizes the matrix’s seventh principle; most sourcing decision is re-litigated on a quarterly cadence; by surfacing the eight calls most likely to need re-litigation.

Reading the catalog

Each entry has four fields: failure shape (what happened), leading indicator (the earliest signal in the proposal or kickoff doc), wreckage (sunk cost, missed window, morale damage), correct call (the verb the team should have used).

The leading indicators are the load-bearing column. Most entry could have been caught months before failure if the indicator was named. They are the screening questions for the procurement desk.

1. Built model gateway from scratch

Failure shape. Team builds a multi-provider gateway with routing, fallbacks, caching, structured outputs. 12-18 months later it exists at 60-70% of LiteLLM/Portkey/Bedrock parity; engineers are bored; the agent team is a release behind.

Leading indicator. Proposal lists “flexibility” or “control” or “avoid lock-in” without naming a specific capability the bought alternative cannot deliver. Or the proposal compares against an outdated version of the buy. Or fewer than 4 engineers are committed steady-state.

Wreckage. 12-18 engineer-months. Two to four engineer-quarters of agent work delayed. The gateway becomes a maintenance liability nobody wants.

Correct call. Buy from a specialist. The flexibility argument is almost usually solved by an adapter pattern (a 200-line vendor adapter the team owns). Save engineering capacity for the agent layer where moat density is real.

2. Bought “AI platform” no one used

Failure shape. Procurement signs a six- or seven-figure annual contract for an end-to-end “AI platform” meant to host many internal agents and copilots. Engineers ship without it because the workflow shape doesn’t match the platform’s idioms. Sub-10% utilization through renewal.

Leading indicator. Selected by procurement or executive sponsor without engineering users running a real pilot measured in shipped features. Or sold as “supports any AI use case,” which translates as “supports no specific case well.”

Wreckage. $250K-$2M in unused contract value. An org perception that AI tools “don’t work” that biases the next decision toward in-house build.

Correct call. Run a 4-6 week pilot with engineering users on a real workload before signing. See the AI procurement maturity model and apply level-3 procurement to anything over $100K.

3. Hired one AI engineer instead of an agency

Failure shape. Exec is told the org “needs AI” and hires one AI/ML engineer. The hire spends three months negotiating compute, two picking infra and eval tooling, and quits at six months citing “no support, no roadmap.” Zero shipped capability.

Leading indicator. The JD asks the hire to also pick the model gateway, eval framework, observability vendor, and deployment target. Or the role lacks a product partner, eval lead, or platform peer. Or the exec describes the role as “build out AI for us.”

Wreckage. 6-9 months. One regretted hire. Brand damage in the AI hiring market when the exit interview circulates.

Correct call. Engage an AI agency for the first capability; see a field guide to evaluating an AI agency in under 90 minutes; to ship in 8-12 weeks while the hire pipeline matures. Then hire two engineers minimum with a defined product partner.

4. Fine-tuned a foundation model without an eval set

Failure shape. Team decides the foundation model “doesn’t know our domain” and fine-tunes on a curated dataset. The tune improves on a vanity metric and regresses on real production workload; undetected for three months because there is no eval set. Base is deprecated six months later; the tune is unmaintainable.

Leading indicator. Team pursues fine-tuning before having a versioned eval set with regression tracking. Or the eval is “we’ll generate it after the tune ships.” Or the rationale is “prompts are too long” rather than a documented quality gap on a specific workload.

Wreckage. 4-8 engineer-months. Inference cost increase from custom hosting. An orphaned model when the base is deprecated.

Correct call. Build the eval set first. Most workloads improve more from prompt engineering, retrieval, and tool design than from fine-tuning. Fine-tune only after the eval has surfaced a gap prompt-and-retrieval cannot close, against a base the vendor commits to maintain.

5. Bought end-to-end agent platform

Failure shape. Scaleup buys an “end-to-end agent platform” promising orchestration through deployment. Product team ships an AI feature that is a thin wrapper around the vendor’s agent. A competitor on the same vendor ships parity in a quarter. The buyer is competitively undifferentiated.

Leading indicator. Procurement spec asks for “end-to-end” or “turnkey” without specifying which workflow logic stays in-house. Or the demo is on the vendor’s preferred workload, not the buyer’s. Or the contract has no kill clause tied to performance.

Wreckage. Loss of moat. Per-call vendor margin into the vendor’s gross margin. Switching cost that grows quarter over quarter as workflow logic accretes inside the vendor’s product.

Correct call. Build the workflow agent and buy the substrate beneath; see why scaleups should build agents and buy infrastructure. The agent is the moat; commoditizing it via vendor is paying to undifferentiate the product.

6. Bought eval-as-a-service for a workload no vendor had seen

Failure shape. Enterprise buys eval-as-a-service running generic-benchmark eval. Dashboard turns green in two weeks. Three months in, a production regression the eval missed reaches customers. The org rebuilds in-house anyway, having paid for the masked version.

Leading indicator. Vendor doesn’t ask to see actual production traffic and failure modes during sales. Or cannot articulate why their eval would catch a domain-specific regression. Or the dashboard is green on day one.

Wreckage. 6-12 months of false confidence. A production regression that costs customer trust. Sunk SaaS cost.

Correct call. Build the evaluator. Buy compute or harness if useful; see the case for buying your AI evaluation stack and building your AI evaluator. Eval-as-a-service for org-specific workloads is the entry that recurs most often in the graveyard.

7. Built observability and tracing in-house

Failure shape. Team builds in-house tracing because “we want control of the data” or “vendors are expensive.” 12 months later the schema lags most published vendor schema by 6-12 months, the agent team uses a worse view, and one engineer maintains the build permanently.

Leading indicator. Team writes the tracing schema before the agent that needs tracing. Or rationale is “data sovereignty” without a specific compliance requirement the bought vendor cannot meet (most run in-region or on-VPC). Or the build is sized at one engineer.

Wreckage. 8-12 engineer-months. A tracing system that ages out and is swapped for the bought alternative anyway. Agents shipped without observability for the first six months.

Correct call. Buy observability from LangSmith, Langfuse, Helicone, Arize, or similar. The data-sovereignty argument is solved by on-VPC deployment most vendors offer. Save the engineer-months for the agent layer.

8. Picked a vendor on demo polish without a kill clause

Failure shape. Procurement picks the vendor with the most polished demo. Six months in, production behavior diverges from the demo and the buyer cannot exit without paying out a 24-month term. Stuck with a worse capability than the second-place vendor would have delivered.

Leading indicator. Contract has no kill clause tied to a named exit-trigger metric (eval threshold, latency SLA, cost-per-call ceiling). Or the buyer has no eval set to measure the vendor against. Or the procurement timeline did not include a real-workload pilot.

Wreckage. 12-24 months locked in. Sunk contract value. A second procurement cycle at higher cost because the buyer is now under deadline pressure.

Correct call. Run a real-workload pilot. Negotiate a kill clause tied to the buyer’s evaluator metric. Cap term to 12 months on first contract. Read the AI vendor reference-call playbook and check references with structured questions before signing.

Using the catalog at decision time

Run any pending decision against the eight indicators before signing or starting:

  1. Build rationale is “flexibility” or “control” without a named capability gap?
  2. Platform selected without an engineering-led pilot?
  3. One engineer being asked to do everything?
  4. Fine-tuning pursued before the eval set exists?
  5. Procurement asking for “end-to-end” without specifying in-house logic?
  6. Eval vendor skipping production-traffic discussion?
  7. Tracing schema written before the agent it should trace?
  8. Contract missing a kill clause tied to a measured metric?

Two or more yes answers and the decision is in graveyard territory. Stop, name the verb at each layer, redo. See why AI build-vs-buy decisions made in 2024 should be re-litigated this quarter for the re-litigation cadence.

Frequently asked questions

What is a build-vs-buy graveyard?

A graveyard is the catalogued set of past build-vs-buy decisions that produced bad outcomes; missed shipping windows, abandoned platforms, sunk-cost write-downs, regretted hires. The graveyard exists because the same wrong calls recur across orgs, and naming each one with its leading indicator is how the next team avoids it.

Why does a model gateway built from scratch land in the graveyard?

Model gateways are commoditized substrate that vendors with multi-year head starts and full-time platform teams build better than a 3-engineer in-house rotation. The build absorbs 12 to 18 months of capacity, lags the buy alternative on features (caching, fallbacks, structured outputs), and produces no moat because most customer of the bought alternative gets the same capability. The leading indicator is the team listing “flexibility” as the rationale without naming a specific feature the bought version cannot deliver.

Why do bought “AI platforms” end up unused?

Because the platform was selected by procurement or by an executive sponsor, not by the engineers who would adopt it. The platform solves a generic problem with a generic interface; the engineering team has specific workflows that require specific shapes. The platform sits unused while engineers ship without it. The leading indicator is a procurement-led selection where the eventual users do not run a real pilot before the contract.

Why is hiring one AI engineer instead of an agency a graveyard call?

One AI engineer cannot ship a production-grade AI capability alone in any reasonable window. They need eval, observability, integration, and product framing; a multi-discipline workflow that an agency’s bench supplies in week one. The hire spends six months building the substrate and quits when the org has no second engineer to pair with. The leading indicator is the hire being asked to also pick the model gateway, the eval framework, and the observability vendor.

Why does fine-tuning a foundation model usually fail?

Fine-tuning is a build path; most teams pursue it without the eval discipline that distinguishes a genuinely useful tune from a regression. The model improves on a vanity metric, regresses on real workloads, and becomes unmaintainable when the base model is deprecated 6 months later. The leading indicator is the team pursuing fine-tuning before having a versioned eval set that can detect regressions.

Why does buying an end-to-end agent platform commoditize the buyer?

End-to-end agent platforms encode the workflow inside the vendor’s product. The buyer’s product becomes a thin wrapper that any competitor with the same vendor can replicate. Switching costs accumulate without moat being built. The leading indicator is the procurement spec asking for “end-to-end” or “turnkey” agent capability without specifying which workflow logic stays in-house.

Why is buying eval-as-a-service almost usually a graveyard call?

Generic eval vendors test generic problems; they cannot test the specific workload that the org runs. Green dashboards mask production regressions until the first real failure. The org rebuilds the eval in-house anyway. The leading indicator is the eval vendor not asking to see actual production traffic and failure modes during the sales process.

Why does building observability and tracing in-house misallocate engineering?

Observability and tracing for AI workloads is a moving target; prompt and response capture, structured-output handling, eval-result attachment, latency and cost telemetry. Vendors who specialize iterate weekly. The in-house build lags by 6 to 12 months in capability, ages out, and consumes the engineers who should be shipping agents. The leading indicator is the team writing the tracing schema before writing the agent that needs to be traced.

Why does picking a vendor on demo polish without a kill clause produce regret?

Demo polish is uncorrelated with production-grade behavior. Without a kill clause; a contractual exit on missed eval thresholds or named SLA violations; the buyer cannot leave without paying out the term. The leading indicator is the procurement contract lacking a named exit-trigger metric and a published timeline for invoking it.

How can a team use the graveyard to avoid the next wrong call?

Run each pending decision against the eight leading indicators before signing or starting the build. If two or more fire, the decision is in graveyard territory and needs a redo. The redo usually flips the verb; if the team was going to build commodity substrate, buy it; if the team was going to buy a moat-encoding agent, build it. The discipline is naming the verb explicitly and re-litigating it on the quarterly cadence the matrix prescribes.

Key takeaways

The graveyard is not random. Eight wrong calls account for most 2024-2026 AI sourcing regret: building a model gateway from scratch, buying an unused AI platform, hiring one engineer instead of engaging an agency, fine-tuning without an eval set, buying an end-to-end agent platform, buying generic eval-as-a-service, building observability in-house, picking a vendor on demo polish without a kill clause.

Each has a leading indicator visible in the proposal; flexibility-as-rationale, procurement-led selection, lone-engineer JD, fine-tune-before-eval, end-to-end RFP language, eval-vendor avoiding production traffic, schema-before-agent, kill-clause absent. Two or more firing is the threshold for stopping and redoing.

The redo is almost usually a verb flip. Build the moat (workflow agents, eval set, prompt registry); buy the substrate (model gateway, observability, vector DB, eval compute); hire the judgment (eval-fluent senior engineer, AI agency that ships while the team scales). Run the catalog before most $100K+ contract and most two-engineer-quarter build.

Last Updated: Jun 23, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles