Milestone billing is the only AI MVP billing model that pays a vendor for progress the buyer can verify. Time-and-materials pays for activity. Lump-sum pays for a promise. Eval-anchored milestones pay for unlock events the founder controls. The case for milestones is structural, not aesthetic: an AI MVP produces four distinct categories of information across a 6–12 week engagement, each one changes what a defensible price looks like, and each belongs at a payment gate where the founder can either release the next tranche or stop the work. This piece defends the four-milestone canonical structure, names what unlocks payment at each, names the founder’s veto, and argues against the two alternatives.
This is the POV companion to AI MVP pricing explained: fixed-price vs hourly vs milestone and the fixed-price AI MVP contract: 7 clauses worth negotiating. It sits inside the AI MVP economics playbook under the idea-to-product manifesto. For the failure-mode counterpart, see the AI agency milestone trap.
Why feature-based milestones don’t work for AI
Pre-AI software borrowed the milestone pattern from construction and tied each gate to a feature (“UI complete,” “API integrated,” “QA pass”). It worked because feature completion was self-evident — an endpoint either returned 200 OK against the test suite or it did not.
That self-evidence collapses for a 2026 AI MVP. The endpoint returns a string. The string is grammatically correct. The string is also, sometimes, wrong. Whether the feature is done is not a binary check; it is a graded sample of representative inputs scored against a written rubric, with a numeric pass threshold.
Feature-anchored milestones therefore concede the most important decision back to the vendor. The vendor declares the feature complete; the founder has no objective basis for pushback because no number was agreed in writing. The quality question gets deferred to production.
McKinsey’s State of AI work has named scope discipline and acceptance criteria as the two largest gaps between AI projects that scale and the 78% that stall. BCG’s Where’s the Value in AI? corroborates: 74% of corporate AI investments fail to scale. Acceptance ambiguity is the dominant contributing cause — and feature-anchored milestones manufacture that ambiguity by design. Eval-anchored milestones do not. The acceptance event is a number against a rubric the founder co-authored. The milestone unlocks when the number clears.
The four canonical milestones
A 2026-defensible AI MVP under milestone billing has exactly four payment gates. Three collapses two distinct buyer decisions into one. Five adds administrative drag without adding new buyer information.
| # | Name | Default split | Default duration | Unlock artifact |
|---|---|---|---|---|
| 1 | Planning | 15% | Week 1 | Signed PRD intent and ICP document |
| 2 | PRD plus evals | 25% | Weeks 2–3 | Written PRD, eval set with rubric, threshold, model pick |
| 3 | Build complete | 45% | Weeks 4–10 | Working build that clears the eval threshold |
| 4 | Handoff | 15% | Weeks 11–12 | Runbook, code transfer, on-call window opens |
The percentages are defaults for a 6–12 week single-feature MVP in the $80K–$250K range. The four-gate structure stays across engagement size; the splits adjust.
The argument for exactly four is informational. Each milestone exposes a category of buyer information that did not exist at the previous gate. Planning surfaces the founder’s clarity on intent. PRD-plus-evals surfaces the team’s eval discipline and the model dependency. Build-complete surfaces actual model behavior against the rubric. Handoff surfaces documentation discipline and on-call posture. Combining any two collapses decisions; splitting any one releases on the same information twice.
Milestone 1 — Planning
Default 15% of fee. Week 1. The tranche unlocks on a signed PRD intent document — two to four pages naming the ICP, the one feature, the one persona, the one primary task, the deployment surface, and the partner team roster. No eval set, no architecture, no budget breakdown yet. The document answers what we are buying.
Founder’s veto. The intent document. If the founder cannot read it back to a non-technical stakeholder in two sentences (“we are building X for ICP Y, measured against Z”), it does not unlock payment.
Why this exists. Three weeks into a build, the most common discovery is that the founder and the partner are pricing different products. Surfacing that delta in week 1, when 15% is at stake, is cheaper than discovering it in week 6 when 70% has cleared. Folding planning into PRD-plus-evals bundles “what we are buying” with “how we measure it” into one decision the partner picks.
Milestone 2 — PRD plus evals
Default 25% of fee. Weeks 2–3. The artifact is the full PRD plus the eval set plus the named model dependency. The eval set names roughly 60–150 representative inputs with a written rubric, anchored sub-criteria, and a single numeric pass threshold (e.g., “≥85% of cases scored ‘pass’ on a four-criterion rubric”). The model dependency names primary plus documented fallback (e.g., Claude Opus 4.8 primary, GPT-5 fallback).
Founder’s veto. The eval set and the threshold. The rubric is the contract’s quality referee. The founder writes it (with partner support) and signs it before the build starts.
Why this exists. It surfaces the team’s eval discipline. A team that ships a 30-example eval with a vague rubric is signaling they expect to negotiate quality at the build-complete gate. A team that ships a 120-example eval with a four-criterion rubric is signaling they have the discipline to ship at threshold. Folding evals into build-complete converts the engagement to T&M by default — the threshold drifts to whatever the partner can hit in the budget remaining.
Milestone 3 — Build complete
Default 45% of fee. Weeks 4–10. The single largest tranche, and the only one gated on a number rather than a document. The artifact is a working build that scores at-or-above the threshold on the milestone-2 eval set, with the grader run reproducibly and the score recorded in writing.
Founder’s veto. The eval score against the threshold. If the score is below threshold, the tranche does not release. The partner re-runs prompt engineering, retrieval, guardrails, or model fallback at no incremental cost until the score clears. The remediation budget lives inside the 45% the founder has not yet released.
Why this exists. Every previous milestone is paid for documents; this one is paid for measured quality. The partner who scoped tightly clears the threshold inside the build budget; the partner who scoped loosely either eats the overrun or asks for a change order. Releasing the 45% on “build complete, eval to follow” concedes verification back to the vendor — the number arrives after the money has cleared, when nobody is incentivized to find a regression.
Milestone 4 — Handoff
Default 15% of fee. Weeks 11–12. The artifact is the runbook plus the code transfer plus the opening of the on-call window. The runbook documents the model dependency, prompt library, eval set with grader and rerun script, observability instrumentation, failure-mode taxonomy, on-call escalation tree, and change-management process for model migrations.
Founder’s veto. The runbook. Specifically: whether a competent successor team could re-run the eval, ship a model migration, or triage a sev-1 incident from the document alone.
Why this exists. A 15% incentive at the end is the cheapest mechanism for ensuring the team writes a transferable runbook. Without it, runbooks are written in the last 48 hours by whoever was least busy. Folding handoff into build-complete releases 100% before the documentation exists — the founder discovers the gap on day 31 when the on-call window closes.
Why milestones beat time-and-materials
T&M moves budget control to the buyer. Every invoice is a pay/pause/terminate decision. For a non-engineer founder, this is the wrong decision to vest. The founder cannot adjudicate whether 60 hours of prompt engineering was reasonable or 120. Invoice review is calibrated for a buyer who reads code; most AI MVP buyers do not. Milestones move the decision to acceptance review — whether the artifact clears the rubric the founder co-authored.
T&M also makes the partner indifferent to overruns. ACM Queue’s 2024 review reports T&M engagements under unstable dependencies systematically overrun by 40–80%. AI MVPs run on a frontier-model dependency that ships behavior changes on cycles of weeks — exactly that case.
Why milestones beat lump-sum fixed-price
Lump-sum has the opposite failure: it collapses four buyer decisions into one signing event. At SOW signing, the founder commits to a price for a product they have not defined (no PRD), have not scoped (no eval set), have not model-anchored (no architecture), and have not planned for handoff. Each commitment hides a tail risk the founder is not positioned to price. A partner pricing competitively under lump-sum will protect margin by interpreting scope tightly — the buyer’s veto is one moment, the vendor’s interpretive control runs across every subsequent week.
Milestones spread the decisions. Each tranche is priced against information that did not exist at the previous one. The founder is buying progress with verification, not a promise with hope. This is the same argument Harvard Business Review’s contracting-under-uncertainty literature has made for decades about milestone contracts in pharma, defense, and aerospace: when output quality cannot be verified at signing, the rational contract pays for partial output at the points where verification becomes possible.
The founder’s veto, written into the SOW
The veto is the load-bearing primitive. Without it, milestones become ceremonial — the partner declares “build complete,” the founder feels socially obliged to release the tranche, and the eval threshold becomes a recommendation rather than a contract term.
The SOW language that holds the line has four properties:
- Unilateral. The founder approves; the partner does not co-approve.
- Time-bounded. A defined review window per milestone (typically 5 business days for milestones 1–2, 10 for milestones 3–4).
- Rubric-anchored. The veto names the document or score being evaluated, e.g.: “The Founder may withhold acceptance of Milestone 3 if the eval set fails to clear the threshold defined in Appendix B.”
- Remediation-defined. A defined remediation path inside the original fee for the first cure attempt (typically 2 calendar weeks), with subsequent attempts at a published rate.
A vendor draft will typically replace “the Founder may withhold acceptance” with “the parties shall mutually agree” or “the partner shall demonstrate.” Both are veto removal. The vendor’s objection is usually framed as fairness: “You can hold us hostage at every milestone.” The response: “You can be held to a rubric you co-authored. If you can clear it, the milestone unlocks. The veto is the contract’s quality referee, not a hostage.”
When milestone billing breaks
Three failure modes recur in 2026 contracts.
Ambiguous unlock criteria. A milestone that unlocks on “build complete to the satisfaction of both parties” is a renegotiation, not a milestone. If the unlock cannot be reduced to a sentence with a number or named artifact, rewrite it before signing.
Milestones too large. A two-milestone 50/50 contract is lump-sum with a haircut. Push for four.
Milestones used as cover for T&M. A common pattern: four milestones at fixed percentages, plus “hours over budget will be billed at the standard rate.” This is T&M with a marketing veneer. Strike the overage clause.
FAQ
What is the typical split percentage between the four milestones?
A 2026-defensible default is 15/25/45/15 — planning, PRD-plus-evals, build-complete, handoff. Build-complete carries the largest weight because it is the only milestone gated on measured quality rather than a document. Engagements under $80K can compress planning to 10%; engagements over $250K can lift build-complete to 50%.
What happens if a milestone fails the eval threshold?
The tranche does not release. The partner re-runs the work — prompt engineering, retrieval tuning, guardrails, model fallback — inside the original fee for the first remediation attempt (typically two calendar weeks). Subsequent attempts run at the published change-order rate.
Can milestone billing be combined with time-and-materials?
Yes, in a hybrid: milestones cover the eval-bound core build, and a separate T&M envelope covers the integration surface where eval thresholds do not apply (bespoke Salesforce work, custom auth flows).
How long should each milestone take?
A 6–12 week MVP typically runs: planning week 1, PRD-plus-evals weeks 2–3, build-complete weeks 4–10, handoff weeks 11–12. Compressing planning or PRD is a common partner request and a common founder mistake.
How do I write the founder’s veto into the SOW?
Four properties: unilateral founder acceptance authority, a defined review window per milestone, acceptance anchored to a named rubric (not to “mutual agreement”), and a defined remediation path inside the original fee for the first cure attempt.
What if the partner refuses milestone billing?
Two cases. Vendors whose internal accounting is T&M — usually solvable; most back offices can split invoicing into four tranches. Vendors who lose margin under milestones — a signal they priced the project assuming the buyer would not verify acceptance. That is exactly the case milestones defend against.
Should milestones be tied to dates or to deliverables?
To deliverables. Date-tied milestones convert acceptance to a calendar event and reintroduce the worst feature of fixed-price. Eval-threshold-tied milestones keep the bar fixed on quality and let the date float.
Are milestones overkill for a small AI MVP under $50K?
Below roughly $40K, a two-milestone 50/50 structure can work because downside is bounded. Above $40K, the information advantages of four milestones exceed the administrative cost. Above $150K, four milestones are the default.
How does milestone billing handle mid-engagement model migrations?
A milestone contract can absorb one in-engagement migration inside the build-complete tranche if the SOW names the primary model and a documented fallback. A founder-requested migration to a not-yet-named model is a change order.
What if my partner wants more than four milestones?
Five-plus usually means decomposing build-complete into sub-milestones (UI, backend, eval-clear), which reintroduces feature-anchored problems. Build-complete is one milestone, gated on one number.
Closing
The 15/25/45/15 four-milestone structure — planning, PRD-plus-evals, build-complete on eval-threshold pass, handoff — is the default a non-engineer founder should ask for. It is the consequence of building software whose unit of “done” is an eval threshold rather than a working endpoint. Take it into the next vendor conversation. The partners who can hold the bar are the partners who can ship at quality.
For the operational worksheet that produces the eval set and milestone splits before negotiation, download the AI MVP scoping worksheet. For the wider picture, see the AI MVP economics playbook and the idea-to-product manifesto, and the companion piece on the end of the fixed-price AI project for the broader market argument.
Arthur Wandzel