Solo founders sign AI MVP contracts alone, and that is the problem this list is written for. No technical co-founder reads the SOW with you. No fractional CTO is on retainer. The discovery call was charming, the team looked sharp, the demo was crisp — and now a six-figure proposal sits on your desk on a Friday afternoon, with a Monday signature expectation. This article is the print-and-keep red-flag list to run against that proposal in the next hour. Ten flags. Ordered by severity. Each one tells you the diagnosis, the walk-away verdict, and roughly what it costs you in dollars and months of runway when you sign anyway.
This is a decision-stage companion to the founder × AI partner operating manual and the master frame on how non-engineers ship AI products in 2026. Where the operating manual tells you how to run the engagement week by week, this list tells you whether the engagement is worth starting at all. Pair it with the inverse — 11 must-haves in a founder-friendly AI partner — and the longer-form 7 SOW signals that predict overrun. For the contract-level definition of what a partner means when they promise “production-ready”, see decoding “production-ready” in AI agency proposals.
Table of Contents
- How to use this list
- Tier 1 — Walk now (flags 1–3)
- Flag 1: No eval discipline in the methodology
- Flag 2: No fixed eval acceptance threshold
- Flag 3: Hourly billing without a weekly milestone schedule
- Tier 2 — Renegotiate or walk within 72 hours (flags 4–7)
- Flag 4: No IP and weights clause
- Flag 5: No named technical lead
- Flag 6: No on-call window past launch
- Flag 7: No graceful-exit clause
- Tier 3 — Raise, document, watch (flags 8–10)
- Flag 8: No reference clients you can call
- Flag 9: An overpromised demo
- Flag 10: No SOW (or a one-page SOW)
- The print-and-keep worksheet
- Frequently Asked Questions
- Book a 30-minute idea review
How to use this list
Print this page. Sit with the proposal. For each of the ten flags, find the clause in the SOW that addresses it — or confirm the clause is missing. Mark pass, fail, or unclear. A flag is unclear only if the SOW gestures at the topic without answering the question. “We are committed to quality” is unclear. A numeric eval threshold is pass. Silence is fail.
The severity ordering matters. Flags 1–3 are deal-breakers in 2026: their presence in a SOW is the most reliable signal that the engagement will overrun, drift, or quietly fail acceptance. Flags 4–7 are recoverable if the partner inserts the missing clause within 72 hours. Flags 8–10 are softer — they raise questions about how the partner operates but rarely sink an otherwise-defensible engagement on their own.
McKinsey’s State of AI and Gartner’s pilot-to-production research put AI-pilot failure rates between 70% and 85%. The flags below are the contractual fingerprints of the projects in that failure bucket. Treat each flag as a clarification request first, a renegotiation second, a walk-away third.
Tier 1 — Walk now (flags 1–3)
These three flags distinguish an eval-first AI partner from a 2022-era dev shop that bolted on AI work. None should appear in a defensible 2026 AI MVP contract. If any of the three is present, treat it as walk-away material unless the partner rewrites the clause same-day.
Flag 1: No eval discipline in the methodology
Diagnosis. The methodology describes how the team will build, demo, and ship — but says nothing about how they evaluate model quality on a fixed test set. No mention of an evaluation set, regression budgets, or failure-mode tracking. The team treats evals as a QA afterthought rather than the central contract.
Why it kills the project. AI features are delivered to a quality bar on a distribution of inputs. Without eval discipline, the team has no measurable definition of done — only a definition of shipped. You discover at acceptance that the bar you assumed was implicit is not the bar they built to, and the fix runs on your budget.
Walk-away verdict. Walk. If a 2026 AI partner cannot articulate their eval practice in two sentences, they are not building the product you think you are paying for.
Cost of getting it wrong. Missing eval discipline correlates with a 25%–40% overrun ($20k–$70k on a typical $80k–$180k MVP) and a four-to-eight-week slip — half a year of runway, gone, with no defensible product at the end.
Flag 2: No fixed eval acceptance threshold
Diagnosis. Even partners that mention evals often fail to encode them as contract artifacts. The acceptance section reads “deliverables to be reviewed and accepted by client” rather than “delivery passes when the system scores ≥92% on a 150-case evaluation set, with no failure mode below 85%.” No numeric threshold, no fixed test set, no per-failure-mode breakdown.
Why it kills the project. Without a numeric acceptance criterion, every milestone payment becomes a negotiation. The team delivers, you ask “does it work?”, and the conversation collapses into qualitative back-and-forth. Each round is a week of slipped timeline and another week of burn.
Walk-away verdict. Walk if the SOW lacks an acceptance appendix with one row per AI-bearing deliverable. A defensible partner has the appendix in a template and produces it inside 24 hours.
Cost of getting it wrong. A slow-creep overrun of $15k–$40k and a two-to-four week milestone-dispute window where neither side can prove the deliverable was met. The runway clock keeps ticking through every dispute call.
Flag 3: Hourly billing without a weekly milestone schedule
Diagnosis. Pricing quotes an hourly rate ($150–$300/hour is typical in 2026 for senior AI engineers) and an estimated hour band (“400–600 hours”), without a week-by-week milestone schedule that ties hours to outcomes. The invoice format is “hours billed × rate”, not “milestone delivered, week N”.
Why it kills the project. Pure hourly billing blows solo-founder budgets because the work that produces evals, observability, and recovery code is invisible to a non-engineer reviewing a weekly invoice. The team logs 220 hours, you see a number, and no shared artifact maps hours to outcomes. By week six you have spent half the budget without a way to tell whether week six was on track.
Walk-away verdict. Walk unless the partner converts to milestone-based fixed price with hourly overage above a defined cap. A defensible engagement names a milestone per week and ties at least 70% of contract value to milestone acceptance.
Cost of getting it wrong. Overruns on pure hourly typically run 30%–60% for solo founders ($25k–$100k on a $100k–$170k engagement). The structure invites scope drift; the structure is the red flag.
Tier 2 — Renegotiate or walk within 72 hours (flags 4–7)
These four flags are recoverable, but recovery is binary: the partner either inserts the missing clause inside three business days, or they have told you the omission was structural. The proposal is not signable until the clauses appear.
Flag 4: No IP and weights clause
Diagnosis. The IP section says “all source code is the property of the Client.” It does not mention fine-tuned model weights, the evaluation set, prompt-version history, or vendor-hosted model artifacts. That was a defensible SaaS clause in 2018; it is not a defensible AI clause in 2026.
Why it matters. The codebase is the wrapper. Fine-tuned weights and the evaluation set are the bulk of what makes the system perform. If the partner retains either, you cannot meaningfully migrate, renegotiate, or operate independently. Six months later, you are paying ongoing fees for access to artifacts you already funded.
Walk-away verdict. Insert a five-item IP-and-artifacts schedule: source code with history, fine-tuned weights to a client-owned registry, evaluation set as a versioned dataset, prompt-version history in a structured export, vendor-hosted model artifacts with a 30-day handover. If the partner declines item two or three, the engagement is a managed service with switching costs.
Cost of getting it wrong. Annual lock-in fees of $10k–$40k post-launch, or a forced rebuild at 40%–80% of the original budget if you eventually need to leave.
Flag 5: No named technical lead
Diagnosis. The team section names roles (“Senior AI Engineer”, “Full-Stack Engineer”, “PM”) and a headcount. It does not name a specific senior engineer, by name, as the architect of record — the one who signs the architecture document, attends every weekly demo, and is on the page for the first production incident.
Why it matters. Without a named lead, accountability diffuses. The engineer who designs the architecture is rarely the one who writes the implementation, rarely the one on the demo call. Six weeks in, you cannot tell who is responsible for an eval regression because no one in particular is. The “staffing-reservation” pattern — sell one engineer to win the deal, assign another to do the work — depends on this ambiguity.
Walk-away verdict. Demand a named-lead clause: one engineer, named in the SOW with LinkedIn link, calendar invites for the first six demos, on-call commitment for the first 30 days post-launch. If they leave mid-engagement, the partner notifies you in writing within 48 hours and proposes a successor you can interview. Of every clause in this list, this one delivers the largest reduction in operational risk per word added.
Cost of getting it wrong. Mid-engagement re-staffing costs two to three weeks of context-loss and frequently a $10k–$25k catch-up bill. For screening the named engineer’s prior work, see how to evaluate an idea-to-product partner’s prior work.
Flag 6: No on-call window past launch
Diagnosis. The timeline ends at “MVP live”. Transition is vague — “knowledge transfer” or “documentation handover” — and the SOW names no rotation, no response-time SLA, no kill-switch protocol. The implicit message: operations are your problem from day one.
Why it matters. The default “we build, you operate” shape collapses on day eight of production traffic. A solo founder hits an eval regression, a cost spike, a model-deprecation notice, or a prompt-injection probe — and has no one to call. Remediation either runs through paid change orders at the partner’s premium rate, or you cobble together a fix that breaks something else, or the system silently degrades.
Walk-away verdict. Require a 30-60-90 on-call window. Days 1–30 partner-led with the named lead on call. Days 31–60 joint, with your nominated operator shadowing. Days 61–90 founder-led with the partner on backup. Day 91, transition to a documented support contract or a clean exit.
Cost of getting it wrong. A single post-launch incident with no on-call costs anywhere from $5k (one paid emergency change order) to $30k+ (silent degradation, customer churn, eight weeks of rebuild).
Flag 7: No graceful-exit clause
Diagnosis. The contract has a termination clause naming a 30-day notice period, and nothing else. No exit appendix specifying what is delivered on termination, in what state, by what date, at what capped remediation cost.
Why it matters. A 30-day notice is only useful if both parties want a clean exit. If the engagement degrades, those 30 days are spent disputing scope rather than transitioning artifacts. The graceful-exit clause caps your downside if any of flags 1–6 was misjudged.
Walk-away verdict. Require an exit schedule as a numbered appendix: code to a named repository, evaluation set exported, weights to a client-owned registry, prompt-version history exported, architecture document and runbook delivered, capped remediation budget ($5k–$15k typical) for follow-on questions in the first 30 days post-exit. A partner who has done this before describes the exit schedule in two minutes.
Cost of getting it wrong. A bad exit without this clause routinely costs $15k–$40k in remediation, three to six weeks of stuck transition, and a meaningful chance of losing access to the weights and eval set entirely.
Tier 3 — Raise, document, watch (flags 8–10)
These rarely sink a deal on their own, but each tells you something about how the partner operates. Raise them in the proposal-review call. Document the answer in writing. Two or more tier-3 fails coexisting with one tier-2 fail escalates to tier-2-equivalent severity.
Flag 8: No reference clients you can call
Diagnosis. The proposal lists logos and case studies. None is accompanied by a reference client willing to take a 20-minute call under NDA. The partner offers “we can connect you with happy clients” without committing to specific names before signing.
Why it matters. Logos are wallpaper. A 20-minute call with a founder who shipped with this team — and is now operating the product alone — is the highest-information artifact in vendor selection. Refusing to surface that call before signing means either (a) the partner does not have happy operating clients, or (b) the named team that shipped those projects has since left.
Walk-away verdict. Raise. Ask for two reference calls with named operating clients before signing. A partner with a track record produces both inside a week. A partner with friction here is signaling something about the post-launch experience their clients had.
Cost of getting it wrong. Lower-magnitude than tier-1 or tier-2, but discovering at week six that the partner has high client churn is the cost of the whole engagement. The asymmetry justifies insisting on the calls.
Flag 9: An overpromised demo
Diagnosis. The pitch demo is impressive — too impressive. Latency is sub-second on a chat surface the proposal says will route through three model calls. Accuracy is presented as 99% on a domain the partner has worked in for six weeks. The demo elides which prompts, which evaluation set, which model version.
Why it matters. The demo-versus-production gap is the single largest source of solo-founder disappointment in AI MVPs. A demo runs on cherry-picked inputs, a frozen model version, and no rate limits. Production runs on real-user inputs, model versions that change quarterly, and rate-limit pressure. The gap between demo accuracy and production accuracy is routinely 15–30 percentage points without a disciplined evaluation regime.
Walk-away verdict. Raise. Ask three questions: which evaluation set was the demo accuracy measured against, what was the model and prompt version, what is the production-realistic regression budget. A defensible partner answers all three in one minute. An evasive partner has confirmed the demo was theatre.
Cost of getting it wrong. Mid-engagement disappointment when you discover at week four that the demoed accuracy was a best case rather than a contractual bar — a renegotiation of acceptance criteria that adds $10k–$25k and three to five weeks.
Flag 10: No SOW (or a one-page SOW)
Diagnosis. The proposal is a Notion page, a Google Doc, or a one-page summary. There is no Statement of Work appendix with clauses on deliverables, acceptance criteria, IP, payment terms, change-order process, and termination. The partner says the SOW will be drafted “after kickoff”.
Why it matters. Drafting the SOW after kickoff means drafting it during the weakest-bargaining-position phase of the engagement. You are already paying. You have announced the partnership. The negotiating room to insist on the missing clauses has evaporated.
Walk-away verdict. Raise sharply and refuse to countersign until a proper SOW exists. A reputable partner has a SOW template they fill in for every engagement. Drafting from scratch post-signature is the procurement equivalent of going on a road trip without a destination.
Cost of getting it wrong. Without a proper SOW the rest of this list is unenforceable. The cost is the cumulative cost of every other flag that subsequently trips, with no contract to anchor your position. See how to read an idea-to-product SOW and what to negotiate.
The print-and-keep worksheet
Print this table. Sit with the SOW. For each flag, mark pass, fail, or unclear.
| # | Flag | Tier | Where in the SOW | Result | Action |
|---|---|---|---|---|---|
| 1 | No eval discipline | T1 — walk now | Methodology / scope | Pass / fail | If fail: walk unless same-day rewrite |
| 2 | No fixed eval acceptance threshold | T1 — walk now | Acceptance appendix | Pass / fail | If fail: walk unless appendix appears in 24h |
| 3 | Hourly without weekly milestones | T1 — walk now | Pricing / schedule | Pass / fail | If fail: convert to milestone-based or walk |
| 4 | No IP and weights clause | T2 — 72h to fix | IP section + artifacts appendix | Pass / fail | If fail: insert 5-item artifacts schedule |
| 5 | No named technical lead | T2 — 72h to fix | Team section | Pass / fail | If fail: require name, LinkedIn, demo attendance |
| 6 | No on-call window past launch | T2 — 72h to fix | Timeline + support | Pass / fail | If fail: insert 30-60-90 on-call schedule |
| 7 | No graceful-exit clause | T2 — 72h to fix | Termination + exit appendix | Pass / fail | If fail: require exit appendix |
| 8 | No reference clients you can call | T3 — raise, document | References / case studies | Pass / fail | If fail: insist on two calls before signing |
| 9 | Overpromised demo | T3 — raise, document | Discovery / demo notes | Pass / fail | If fail: ask the three demo-versus-production questions |
| 10 | No SOW (or one-page SOW) | T3 — raise, document | Document itself | Pass / fail | If fail: refuse to countersign without a proper SOW |
Scoring is simple. Any T1 fail with no same-day rewrite is a walk. Two or more T2 fails without rewrite-within-72-hours is a walk. T3 fails alone are negotiation items; two or more T3 fails together with any T2 fail escalates to walk-away material.
Nothing on this list is exotic. Every clause is in the template of a competent partner. Their absence is information about how the partner operates, and your job — alone on a Friday with a Monday signature expectation — is to read the information.
Frequently Asked Questions
Is this list specifically for solo founders, or does it apply to any AI MVP buyer?
The list applies to any AI MVP buyer, but the print-and-keep, severity-ranked framing is calibrated for solo founders with no technical reviewer. Founders with a fractional CTO can lean on richer frameworks like the 7 SOW signals deep-dive. Solo founders need a checklist that executes in 30 minutes against the proposal on the desk, and that is what this list is.
What is the single most predictive red flag a solo founder should check first?
Flag 2 — no fixed eval acceptance threshold. It is the contractual encoding of “we’ll iterate to quality” and the most reliable predictor of the slow-creep overrun that pushes a $100k engagement to $140k. If a proposal is missing only one clause, this is the clause to fight for.
How long should I give a partner to fix a tier-2 flag before walking?
Seventy-two business hours. A competent partner has each tier-2 clause in a template and rewrites the SOW the same day; the 72-hour window covers a single internal review cycle. A partner that takes a week, or pushes back substantively on more than one clause, is signaling that the omission was structural. Walking is the right call.
What if the partner pushes back on every clause?
Pushback on one or two clauses is a normal negotiation. Pushback on every clause means the partner’s standard contract is built to retain optionality at the founder’s expense. Walk. The cost of finding a different partner is two to four weeks of vendor-selection time; the cost of signing with this one is the cost of every flag they refused to fix.
Should I share this list with the partner before they send the proposal?
Yes — reputable partners welcome it. Sharing the list reframes the conversation from “trust us” to “here is the operating contract”, and partners with mature templates find the screen easy to pass. Partners that resist sharing scope clauses in advance have told you something useful about the asymmetric bargaining position they expect.
How much does a defensible AI MVP cost in 2026, with all ten flags handled?
A typical solo-founder AI MVP that clears all ten flags lands in the $80k–$180k range for a 6–12 week build — roughly $20k–$30k for a PRD phase, $60k–$110k for the MVP build, and $25k–$40k for hardening and a 30-day on-call window. Quotes below $50k almost certainly omit one or more tier-1 clauses. For the broader cost picture, see how much does an AI MVP cost in 2026.
What is the right tone when raising a red flag with a partner?
Professional, not adversarial. “I read every AI MVP proposal this way; these are clauses I expect in the SOW; I would like to add them before signing.” A reputable partner agrees, often with relief that the founder is asking the right questions. The tone is the procurement equivalent of a pilot reading the pre-flight checklist aloud — routine, expected, non-personal.
Can I get this list reviewed against my SOW if I do not have a technical reviewer?
Three workable options. A fractional CTO with one or two paid days. An engineer-friend with the SOW shared under NDA. A 30-minute structured review with the SFAI Labs team, walking the ten flags against your proposal — checklist-based, yes/no per flag, with a written summary you can take back to the partner. Not a substitute for legal review of the contract, which a business attorney should perform in parallel.
What if my favored partner fails a tier-1 flag but every other signal is strong?
A single tier-1 fail without a same-day rewrite is still a walk in 2026. Fast email, a sharp demo, a smart founder on the partner side — none of it compensates for the structural risk of a SOW with no eval discipline, no fixed acceptance threshold, or no milestone schedule. Tell the partner directly: “I want to sign with you. Here is the clause I need rewritten by end of day, or I am walking.” Reputable partners take that frame as a useful constraint.
Book a 30-minute idea review
If you have an AI MVP proposal on your desk this week and any of the ten flags tripped, the next useful step is a 30-minute review with a senior engineer who has run this exact screen on dozens of SOWs. Bring the proposal. We walk the ten-flag checklist. You leave with a yes/no on each flag, a list of clauses to insert, and a written summary you can take back to the partner. The session is no-charge and confidential. Book a 30-minute idea review with SFAI Labs.
For the upstream work — turning the hunch into a PRD worth proposing against — start with the idea-to-product manifesto. For the positive inverse of this list — what good looks like — see the 11 must-haves in a founder-friendly AI partner. For the full guide to running the engagement week by week once you sign, see the founder × AI partner operating manual.
Arthur Wandzel