A foundation-model leapfrog is a frontier release whose head-to-head benchmark on your eval set crosses 5 percent quality lift, 30 percent cost reduction at equivalent quality, or a new capability tier unavailable in the prior generation. When the trigger fires, most fine-tune, distilled model, model-routing config, and cost-sensitive buy workload is re-litigated within two to four weeks. The agent orchestration, eval infrastructure, prompt registry, and integration plumbing stay put; those are moat work that should not churn with model releases. The reset is on the model layer, not the system. Orgs that skip the drill carry two to four quarters of unit-economics premium on most workload that should have moved. This piece is the leapfrog drill: what counts as a leapfrog, what gets re-litigated, what stays put, who owns it, and how to keep the eval set ready so the reset runs on a two-to-four-week shape rather than a quarter.
It sits within the AI build-vs-buy-vs-hire decision matrix for 2026. The matrix’s seventh principle holds that most decisions are re-litigated on a quarterly cadence; the leapfrog drill is the faster overlay that fires when the frontier moves between quarters.
Why a leapfrog resets the matrix
The build-vs-buy matrix encodes a snapshot of a moving target. Most entry was made against a specific frontier capability and a specific cost curve. Most frontier shifts are small enough that the matrix’s quarterly cadence absorbs them; some are not.
A leapfrog is a shift large enough that prior decisions inverted. A workload correctly fine-tuned against last generation’s frontier may lose to this generation’s frontier even before eval drift is counted. A workload correctly bought from provider A may now be cheaper from provider B by a margin exceeding switching cost. A workload held off the agent roadmap because the model couldn’t handle it may now be unblocked.
The reset exists because the prior decisions were rational at the time and are no longer rational against the new frontier. Carrying them forward unchanged is not stability; it is unrecognized regret. The drill is the structural counterweight that re-syncs the matrix to the current frontier within a bounded window.
What counts as a leapfrog
Not most frontier release is a leapfrog. The trigger criteria are specific.
- 5 percent quality lift on the org’s eval set. Head-to-head benchmark, current production model versus new frontier release, on the eval set the org uses. A 5 percent lift in the aggregate quality score is the threshold; smaller lifts are absorbed by the quarterly cadence.
- 30 percent cost reduction at equivalent quality. The new release matches or exceeds the current production model on the eval set and prices 30 percent or more below the current model’s cost-per-useful-task. The threshold is on cost-per-useful-task, not cost-per-token; token-price drops that don’t translate into useful-task cost don’t trigger.
- New capability tier. Extended context (e.g., 1M+ tokens stable), native tool use that was previously brittle, video or audio reasoning that was previously absent, or a step-change in agentic horizon (multi-hour task completion). The capability must be unblocking; work the org wanted to do and couldn’t.
Any one of those three triggers fires the drill. Smaller releases are noted in the next quarterly review and otherwise rolled into normal operations. The discipline matters: a drill that fires on most release becomes the org’s full-time job and stops being a drill.
What gets re-litigated
The reset is scoped to the model layer. Within that scope, most artifact decided against the prior frontier is re-litigated.
Most fine-tune. A fine-tune wins by some specific margin against a specific buy alternative; the leapfrog moves the alternative. Re-litigation runs the eval against the new frontier and compares cost-per-useful-task. Fine-tunes that no longer win are sunset.
Most distilled model. Same logic. A distillation case justified by 1/10 to 1/50 inference-cost ratio at near-frontier quality can have the ratio collapse (frontier price drops 30 percent) or the quality gap widen (distilled model now meaningfully behind on the eval).
Most model-routing config. The router decides which model handles which call. The leapfrog changes the cost-per-useful-task curves and quality-percentile distributions, and the routing config must regenerate against the new curves. Detail in the AI project model-routing economics piece.
Most cost-sensitive buy workload. Buy decisions that survived prior reviews at prior frontier prices may flip when new prices arrive.
The drill is a sweep, not a redesign; each artifact has a memo template, an eval to run, and a decision to make.
What stays put
The reset is on the model layer, not the system. Several layers do not churn with leapfrogs.
Agent orchestration logic. The code that decides which model to call, which tools to use, how to handle errors, how to compose multi-step workflows. Built against the org’s workload, not any specific model. A leapfrog lets orchestration call a different model on the same code path; the code path is moat work and stays.
Eval infrastructure. Harness, threshold-locking, regression-triage workflow, eval-set design. Workload-shaped, not model-shaped. A leapfrog runs through the eval; the eval doesn’t change because of the leapfrog.
Prompt registry. Prompts may be optimized for the new model in the regular cadence, but the registry, versioning, serving layer, and tooling stay put.
Integration plumbing. Connectors to internal data, auth flows, observability traces, cost-monitoring dashboards. These layers see model traffic; they don’t depend on which model.
Orgs that touch the moat layer most leapfrog ship slower; orgs that hold it steady ship faster.
The two-to-four-week shape
The drill runs on a two-to-four-week shape from leapfrog release to written go/no-go on each affected workload. Faster is panic; slower misses the window.
- Week 1: run the eval against the new release, produce head-to-head numbers (cost-per-useful-task, quality-percentile, latency, capability deltas). Identify affected artifacts.
- Week 2: write the memo per artifact; prior decision, new evidence, recommended decision, migration cost.
- Week 3 (if needed): brief decision-makers, finalize the go/no-go set.
- Week 4 (if needed): produce migration plans, name owners, schedule the work.
The shape is set by eval speed, memo count (typically 5 to 30), and decision-maker alignment. Orgs running cleanly hit the two-week end; orgs running cold hit the four-week end. Orgs that miss four weeks consistently are running a quarterly project, not a drill.
Sun-setting losing fine-tunes
A losing fine-tune carries ongoing eval-drift monitoring that no longer earns its keep. Migrate the workload to the buy alternative on a 30-to-90-day timeline:
- Re-route traffic to the buy alternative behind a feature flag, fine-tune still available as fallback.
- Run on the buy alternative for two to four weeks against production traffic with continuous eval.
- Confirm production quality matches the eval’s prediction.
- Remove the fine-tune, decommission serving infrastructure, archive artifacts for audit.
The detail on when fine-tuning is correct in the first place is in build, buy, or fine-tune: a decision frame for foundation-model choices.
Keeping the eval set leapfrog-ready
The drill’s binding constraint is the eval set. If running it against a new model takes a multi-week stand-up, the drill misses the window. The eval must be production infrastructure: versioned, CI-style harness producing a complete report within a day, owned by the eval team, with coverage sufficient to detect the 5-percent-lift threshold. If the eval doesn’t meet those requirements, invest there before the next leapfrog.
Who owns the drill
A single owner with authority to redirect engineering capacity; typically a head of AI engineering, a principal engineer with platform scope, or a CTO at smaller orgs. Committees are too slow; a five-person weekly committee cannot produce a written go/no-go on 15 workloads inside four weeks. One owner, an artifact, a published cadence; everyone else feeds inputs. The owner reports to the executive team after each drill; what fired, what changed; for transparency, not permission.
Frequently asked questions
What counts as a foundation-model leapfrog event?
A frontier release whose head-to-head benchmark on your eval set crosses 5 percent quality lift, or 30 percent cost reduction at equivalent quality, or a new capability tier (extended context, native tool use, video reasoning) that was unavailable in the prior generation. Any of those triggers a reset; smaller releases roll into the normal quarterly cycle.
What gets re-litigated and what stays put?
Most fine-tune, most distilled model, most model-routing config, and most workload classified as cost-sensitive on the buy side gets re-litigated. The agent orchestration layer, eval infrastructure, prompt registry, and integration plumbing stay put; those are moat work that shouldn’t churn with model releases. The reset is on the model layer, not the system.
How fast should the reset run?
Two to four weeks from leapfrog release to a written go/no-go on each affected workload. Faster is engineering-led panic; slower is the org missing the cost or quality window. The two-to-four-week shape is set by how long it takes to run the eval, write the memo, and brief the decision-makers.
Doesn’t this churn the org most quarter?
Only the model layer churns. The orchestration, eval, and integration layers are stable. The drill exists precisely so the model-layer churn is bounded; you re-run the same eval, you re-read the same memo template, you write a new decision against the same gates. The repeated rhythm makes the work cheaper, not more expensive.
What happens to a fine-tune that loses to the new frontier?
It is sunset on a published timeline; typically 30 to 90 days from the reset decision. The fine-tune carries ongoing eval-drift monitoring cost that no longer earns its keep, so the right move is to migrate the workload to the buy alternative and decommission the fine-tune. Holding a losing fine-tune in production is a fixed cost against an improving alternative.
Who owns the reset drill?
A senior engineering decision-maker with the matrix as a standing artifact and the authority to redirect engineering capacity. The drill is not run by a model-selection committee; committees are too slow for a two-to-four-week shape. One owner, an artifact, and a published cadence is the structure that holds.
What if no leapfrog happens for several quarters?
The quarterly re-litigation still runs on its normal cadence; the leapfrog drill is a faster overlay that fires when the trigger is met. Quiet quarters are fine. The risk is treating the absence of a recent leapfrog as evidence the next one won’t come; the frontier release schedule is bursty, not smooth.
How do we keep the eval set ready for leapfrog reset?
Treat the eval set as production infrastructure. Versioned, owned by the eval team, run on most model release as a CI-style benchmark, with cost-per-useful-task and quality-percentile reported in a dashboard. If the eval set requires a multi-week stand-up to test a new model, the drill is too slow and the eval is the bottleneck; invest there first.
Does the reset apply to closed-source frontier releases only or to open-weights too?
Both, with different rhythms. Closed-source frontier releases reset the buy default; open-weights releases reset the fine-tune-base and the open-source distillation options. The trigger criteria are the same; 5 percent quality lift, 30 percent cost reduction, or a new capability tier; applied to whichever layer the release affects.
What’s the cost of skipping a leapfrog reset?
Two to four quarters of unit-economics premium on most workload that should have moved, plus a competitive position eroded by anyone who did move. The aggregate cost is rarely visible in a single line item; it shows up as the org’s AI feature velocity falling behind a peer that quietly resets most leapfrog.
Key takeaways
A leapfrog is a frontier release crossing one of three thresholds: 5 percent quality lift, 30 percent cost reduction at equivalent quality, or a new capability tier. Smaller releases roll into the quarterly cadence; leapfrogs trigger a faster reset.
The reset is scoped to the model layer. Most fine-tune, distilled model, routing config, and cost-sensitive buy workload is re-litigated. The agent orchestration, eval infrastructure, prompt registry, and integration plumbing stay put; those are moat work that should not churn.
The drill runs on a two-to-four-week shape from release to written go/no-go. The shape is held by production-grade eval infrastructure, a memo template per artifact, a single owner with authority, and a published cadence. Committees and cold evals miss the window.
Losing fine-tunes are sunset on a 30-to-90-day timeline with feature-flagged migration to the buy alternative. The eval set is the binding constraint and must be invested in as production infrastructure. Orgs that skip the drill carry two to four quarters of unit-economics premium on most workload that should have moved.
Arthur Wandzel