Seven specific patterns in an AI MVP proposal predict a 30%+ budget overrun before a line of code is written. They are not the soft signals every procurement guide warns about — slow email replies, junior staffing, missing case studies. Budgets blow up on vague eval language, missing acceptance criteria, hourly billing with no weekly milestones, a thin IP clause, no named technical lead, no on-call window past launch, and no graceful-exit clause. Each is a sentence the proposal either contains or does not. This article names the seven, explains why each predicts overrun, gives you the question that surfaces it, and tells you what to demand on a Friday night before a Monday signature.
This is a companion to idea validation to a defensible PRD and the master frame on how non-engineers ship AI products in 2026. Where those tell you what good looks like, this one names what wrong looks like at the proposal-review stage, with hours to go before a six-figure commitment. Pair it with the upstream partner-prior-work evaluation guide and the engagement checklist before you sign.
Table of Contents
- Why these seven and not the usual ones
- Signal 1: Vague eval language
- Signal 2: No fixed eval acceptance criteria
- Signal 3: Hourly billing without weekly milestones
- Signal 4: No IP and weights clause
- Signal 5: No named technical lead
- Signal 6: No on-call window past launch
- Signal 7: No graceful-exit clause
- The Friday-night worksheet
- Frequently Asked Questions
- Book a 30-minute idea review
Why these seven and not the usual ones
Most red-flag lists for AI partners were written before evals became the contract. “Weekly demos” was a credible quality story and “you own the code” was a sufficient IP clause. Both statements are now half-true. Demos do not catch eval regressions; code ownership does not cover fine-tuned weights, prompt versions, or evaluation sets.
Each signal below was chosen by one test: its presence in a proposal correlates with a measurable downside. Vague eval language and missing acceptance criteria predict scope drift and the slow-creep overrun that takes a $100k MVP to $140k. Hourly billing without milestones and a thin IP clause predict billing surprises and post-launch dependency. No named lead and no on-call window predict the most common operational failure — the team that built the system leaves before the founder can operate it. The graceful-exit clause caps your downside if the first six were misjudged.
McKinsey’s State of AI and Gartner’s pilot-to-production surveys consistently report that roughly 85% of AI pilots stall before production. The seven signals are the contractual fingerprints of the projects that stall.
A caveat: this is a screening tool, not legal advice. If a signal trips, treat it as a request for clarification first, a renegotiation second, and a walk-away third. Reputable partners rewrite the clause within 48 hours; the ones that will not have told you something.
Signal 1: Vague eval language
Why it predicts overrun. AI features are delivered to a quality bar on a distribution of inputs. If the SOW does not name the quality bar, the engagement has no definition of done — only a definition of shipped. The vendor ships on the timeline, the founder discovers at acceptance that the bar they assumed was implicit was not the bar the vendor built to, and remediation runs on the founder’s budget.
The language to listen for is soft. “We will iterate to quality.” “We will define success together as the project progresses.” “The model will be tuned during the build.” Each is grammatically a commitment and operationally a deferral. They sound collaborative; they bill as scope.
Question that surfaces it. “What is the numeric eval threshold at acceptance, on what fixed evaluation set, with what failure-mode breakdown?” An eval-first partner answers in one sentence with three numbers — for example, “92% pass on a 150-case evaluation set, decomposed into hallucination, refusal, and extraction-accuracy failure modes, with a 1% week-over-week regression budget.” Without that practice they answer in three sentences with one number.
Do now or walk. Demand a one-page Eval Acceptance Appendix: fixed evaluation set size, per-failure-mode breakdown, pass threshold, model versions, regression budget. If the partner cannot produce this within 72 hours, the engagement is not eval-first.
Signal 2: No fixed eval acceptance criteria
Why it predicts overrun. This is Signal 1 encoded into the document. Without acceptance criteria, every milestone payment becomes a negotiation. The vendor delivers, the founder asks “does this work?”, and the conversation collapses into qualitative back-and-forth because the SOW gives neither side a measurable answer. Each round is a week of slipped timeline and another week of burn.
The pattern is recognizable. Deliverables section lists features (“inbox-triage agent”, “intent classifier”, “draft-reply generator”); acceptance section says “deliverables to be reviewed and accepted by client.” No eval-table appendix, no numeric threshold, no fixed test set referenced by name. A defensible SOW puts each AI-bearing deliverable in an acceptance-appendix row naming the evaluation set version, the metric, the threshold, and the model version. For what production-ready actually means at the contract level, see the breakdown of “production-ready” in AI agency proposals.
Question that surfaces it. “Where in this SOW is the test that decides whether milestone three is paid?” The right answer is “Appendix B, table 2, row 4.” Any other answer means the test does not exist as a contractual artifact.
Do now or walk. Refuse to countersign until each AI-bearing deliverable has a row in an acceptance appendix. A competent partner has the appendix in a template. If they push back with “we keep the SOW flexible to support iteration,” translate: they keep it flexible to support invoicing.
Signal 3: Hourly billing without weekly milestones
Why it predicts overrun. Hourly billing is not the red flag alone. The red flag is hourly without a weekly milestone structure that converts hours into committed deliverables. Pure time-and-materials engagements blow budgets because the work that produces evals, observability, and recovery code is invisible to the founder reviewing a weekly invoice. The vendor logs 220 hours, the founder sees a number, and no shared artifact maps hours to outcomes.
A defensible hourly engagement names a milestone per week: “Week 1 — PRD signed and evaluation set v0 written; Week 2 — capability-feasibility prototype demoed against the evaluation set; Week 3 — architecture document signed and cost model published.” The invoice references the milestone the hours produced. If the milestone slips, the founder sees it that Friday, not in month two when the budget is half-burned.
The right structure for an $80k–$180k MVP is milestone-based fixed price with hourly overage above a cap — vendor margin floor, founder budget ceiling, weekly accountability.
Question that surfaces it. “If I look at the invoice on a Friday at end of week 3, what is the milestone the hours billed that week were supposed to produce, and how do I tell whether it was hit?” A weekly-milestone partner answers in 30 seconds with a specific milestone from the SOW. A pure-hourly partner references velocity instead of outcomes.
Do now or walk. Insert a weekly-milestone schedule into the SOW: 6–12 rows, one per week, one named deliverable, one acceptance check. Tie at least 70% of contract value to milestone acceptance. Reserve hourly billing for an overage band above a defined cap.
Signal 4: No IP and weights clause
Why it predicts overrun. The standard SaaS IP clause — “all source code is the property of the Client” — was written for deterministic software. It does not cover the four artifacts that decide whether an AI MVP is genuinely yours: fine-tuned model weights, the evaluation set, the prompt-version history, and any vendor-hosted model artifacts. A partner that omits these clauses can deliver the codebase and still leave you unable to operate the product. That is post-launch lock-in dressed as code ownership.
The fine-tuned weights and the evaluation set are the bulk of what makes the system perform; the codebase is the wrapper. If the vendor retains either, the founder cannot meaningfully migrate, renegotiate, or operate independently. Six months later, the founder pays ongoing fees for access to artifacts they paid the vendor to produce.
Question that surfaces it. “On termination, what specific artifacts transfer to me, in what format?” A defensible answer names four to six: source code; fine-tuned weights as a registry artifact; the evaluation set as a versioned dataset; the prompt-version history; the architecture documents; the runbook. “We’ll provide the code and any documentation” has confirmed the missing clause.
Do now or walk. Add a five-item IP-and-artifacts schedule: (1) source code with history; (2) fine-tuned weights in a client-owned registry; (3) evaluation set as a versioned dataset; (4) prompt-version history exported in a structured format; (5) vendor-hosted model artifacts with a 30-day handover. If the partner declines item (2) or (3), the engagement is a managed service with switching costs, not idea-to-product.
Signal 5: No named technical lead
Why it predicts overrun. Most proposals name a “delivery team” — engineer count, a PM, sometimes a tech-lead role in the abstract. The red flag is when no specific senior engineer is named, by name, as the architect of record. Without a named lead, accountability diffuses. The engineer who designs the architecture is rarely the one who writes the implementation, and the one who writes is rarely the one on the demo call.
The pattern is in the team section. Roles (“Senior AI Engineer”, “Full-Stack Engineer”, “PM”) and headcount, no names. Names appear at kickoff as bios of engineers the founder did not interview. The architecture document gets signed by “the team.” Six weeks in, the founder cannot tell who is responsible for eval failures because no one in particular is.
Question that surfaces it. “Who is the single named senior engineer who signs the architecture document, attends each weekly demo, and is on the page for the first production incident?” The right answer is a name, a LinkedIn link, and a calendar invite for the first six demos. The partner-prior-work evaluation guide goes deeper on screening that specific engineer’s prior shipped work.
Do now or walk. Insert a named-lead clause: one named engineer is the architect of record, signs the architecture document, attends each weekly demo, and is on call for the first 30 days post-launch. If they leave mid-engagement, the partner notifies the client in writing within 48 hours and proposes a successor the client can interview. Of all the clauses in this article, this one delivers the largest reduction in operational risk per word added to the SOW.
Signal 6: No on-call window past launch
Why it predicts overrun. The default idea-to-product shape is “we build, you operate.” That model collapses on day eight of production traffic. The non-engineer founder hits an eval regression, a cost spike, a model-deprecation notice, or a prompt-injection probe, and has no one to call. The remediation either happens through paid change orders at the partner’s premium rate, or the founder cobbles together a fix that breaks something else, or the system silently degrades and customers churn.
The signal is silence past the launch date. Timeline ends at “MVP live.” Transition is vague — “knowledge transfer” or “documentation handover.” No rotation, no response-time SLA, no kill-switch protocol. The implicit message: operations are the founder’s problem from day one.
A defensible structure has a 30-60-90 on-call window. Days 1–30 are partner-led with the partner’s named lead on call. Days 31–60 are joint: the founder’s nominated operator shadows the partner. Days 61–90 are founder-led with the partner on backup. Day 91, the engagement transitions to a documented support contract or a clean exit.
Question that surfaces it. “If the system pages me at 11 p.m. on day 14 post-launch, who picks up the phone, what is their response SLA, and what is the per-incident cost?” Defensible answers are specific: a named engineer, a 30-minute SLA, a defined hourly rate, a written incident process. Anything vaguer is a placeholder.
Do now or walk. Add a three-row post-launch on-call schedule with named engineer, response SLA, channels, and rate per tier. If the partner declines to staff days 1–30 with a named lead, they are shipping a system the founder cannot operate.
Signal 7: No graceful-exit clause
Why it predicts overrun. The first six signals are predictive. The seventh is the cap. The graceful-exit clause determines your downside if signals one through six were misjudged or if the engagement degrades for reasons no contract anticipated — a co-founder dispute, a pivot, a partner-side staffing change. Without a written exit clause, the only off-ramp is a renegotiation with the party you are trying to leave. The renegotiation is more expensive than the exit clause you wish you had signed.
The pattern is what is absent. The contract has start dates, milestone payments, and a termination clause naming a 30-day notice period. It does not specify what is delivered on exit, in what state, at what cap on remediation cost, with what handover support. The 30-day notice is only useful if both parties want a clean exit; if the engagement degrades, those 30 days are spent disputing scope rather than transitioning artifacts. The engagement checklist before you sign covers the full set of clauses that belong here.
A defensible exit clause looks like a miniature SOW for the wind-down. Code delivered to a named repository. Evaluation set exported. Weights handed over to a client-owned registry. Prompt-version history exported. Architecture document and runbook delivered. A capped remediation budget — typically $5k–$15k — for follow-on questions in the first 30 days after exit. A symmetric non-disparagement clause. A reference-handling clause.
Question that surfaces it. “If we terminate at week 6, what is the exact list of artifacts I receive, in what state, by what date, at what cost?” A partner that has done this before describes the exit schedule in two minutes. A partner that has not will improvise.
Do now or walk. Require an exit schedule as a numbered appendix. If the partner has a template, this takes 24 hours to negotiate. If they do not, you are their pilot for this clause; price accordingly or escalate. The cost of the exit clause is approximately zero; the cost of not having one runs to mid-five figures.
The Friday-night worksheet
Open the SOW. For each signal, find the sentence in the proposal that addresses it.
| Signal | Where in the SOW | Status | Action |
|---|---|---|---|
| 1. Vague eval language | Methodology / scope | Pass / fail | If fail: demand eval appendix |
| 2. Fixed eval acceptance | Acceptance appendix | Present / absent | If absent: refuse signature until added |
| 3. Weekly milestones | Schedule / billing | Present / absent | If absent: insert weekly schedule |
| 4. IP and weights clause | IP section + artifacts appendix | Complete / partial / absent | If partial: add 5-item artifacts schedule |
| 5. Named technical lead | Team section | Named / unnamed | If unnamed: require name, LinkedIn, demo attendance |
| 6. On-call window past launch | Timeline + support | 30-60-90 / vague / absent | If vague: add three-row on-call schedule |
| 7. Graceful-exit clause | Termination + exit appendix | Detailed / generic / absent | If generic: require exit appendix |
If two or fewer signals trip, the proposal is workable; raise them and expect rewrites within 72 hours. Three to four: the proposal needs structural rework; budget an extra week. Five or more: the partner is not idea-to-product capable in the sense this article means; walk, or treat it as a learning engagement with a small budget and a hard cap.
None of these seven signals are obscure. Each corresponds to a clause a competent partner already has in a template. Omissions are information about how the partner operates. The Friday-night worksheet is, in part, a competence screen disguised as a checklist.
The right framing with the partner is professional, not adversarial: “I read every AI-MVP proposal this way; these are clauses I expect; I would like to add them.” A reputable partner agrees, often with relief that the founder is asking the right questions. A partner that resists has revealed the omission was deliberate.
Frequently Asked Questions
What is the single most predictive red flag in an AI MVP proposal?
The absence of fixed eval acceptance criteria in the statement of work. This is the contractual encoding of “we’ll iterate to quality” and the most reliable predictor of scope drift, milestone disputes, and slow-creep overrun. Every other signal correlates with overrun, but missing acceptance criteria is the one that converts every milestone review into a negotiation rather than a measurement. If a proposal is missing only one clause, this is the clause to fight for.
How much should an AI MVP cost in 2026?
A typical idea-to-product engagement in 2026 lands in the $80k–$180k range for a 6–12 week build, split across a PRD phase ($20k–$30k), the MVP build ($60k–$110k), and a hardening period ($25k–$40k). Prices below $50k almost always omit one of the seven items in this article. Prices above $250k for a single-feature MVP usually reflect either enterprise overhead or undefined scope. For deeper mechanics, see the AI MVP economics playbook.
Is hourly billing always a red flag?
No. Pure time-and-materials without weekly milestones is the red flag. Hourly billing combined with a milestone-by-week schedule, an outcome-tied invoice format, and a capped overage band is workable for engagements where scope cannot be fully fixed at signing. The screen is whether the founder can read a Friday invoice and tell whether the week’s hours produced the week’s milestone. If yes, the structure is defensible. If no, the structure invites drift.
Should the partner own any AI artifacts after delivery?
The partner should retain nothing required to operate the system. They may reasonably retain reusable internal tooling, redacted reference snippets for case studies with written consent, and proprietary frameworks they used as scaffolding. They should not retain the fine-tuned weights, the evaluation set, the prompt history, or any vendor-hosted model artifact tied to your data. The five-item artifacts schedule in Signal 4 is the contractual line.
What if the partner says “we don’t share team member names until contract signing”?
Treat this as a near-decisive red flag. The signal you are buying with the named-lead clause is accountability, which cannot be deferred behind confidentiality. Reasonable confidentiality is “we share names under NDA after a shortlist conversation.” Unreasonable is “you’ll meet the team at kickoff.” The latter is a staffing-reservation pattern in which the partner sells one engineer to win the deal and assigns another to do the work.
How long does it take to renegotiate a SOW with these clauses inserted?
A competent partner can rewrite a SOW with the seven clauses inserted in 48–72 hours. Each clause is already in their internal templates; they simply did not include it in the sales version. A partner that takes more than a week, or pushes back on every clause, is signaling that the omissions were structural rather than incidental. Plan for a full week of negotiation if more than three signals tripped.
What does the post-launch on-call clause cost?
In a fixed-price MVP, a 30-day named-lead on-call is typically priced as 5–10% of the build budget — $5k–$15k for a $100k–$150k MVP. The 30-60-90 transition with reduced partner availability in months two and three is usually included at no incremental cost or priced as a small monthly retainer. A partner that quotes more than 15% of the build budget for a 30-day on-call is either pricing for undisclosed risk or pricing to discourage the clause.
What if I see only soft red flags but the seven signals pass?
You are likely fine. Soft signals — slow replies, junior staffing — are friction signals and matter at the margin, but the seven contractual signals are the ones that move budget outcomes. A partner occasionally slow on email but writing a defensible SOW with all seven clauses present is a better risk than one replying in 20 minutes and shipping a SOW missing three of them. Optimize for the document, not the email cadence.
Where can I get the SOW reviewed if I do not have a technical co-founder?
Three workable options: a fractional CTO with one or two days of paid time; an engineer-friend with the SOW shared under NDA; or a structured 30-minute review with the SFAI Labs team. The review is checklist-based, takes 30 minutes, and produces a yes/no recommendation on each of the seven signals. It is not a substitute for legal review of the contract, which a competent business attorney should perform in parallel.
Book a 30-minute idea review
If you have an idea-to-product proposal on your desk this week and any of the seven signals tripped, the next useful step is a 30-minute review with a senior engineer who has run this exact screen on dozens of SOWs. Bring the proposal; we walk the seven-signal checklist; you leave with a yes/no on each signal and a list of clauses to insert before signing. The session is no-charge and confidential. Book a 30-minute idea review with SFAI Labs.
For the upstream work — turning the hunch into a PRD worth proposing against — start with the idea validation playbook. For the broader frame, the idea-to-product manifesto is the place to start.
Arthur Wandzel