Most six-figure AI-build engagements are signed on momentum, not diligence — and the standard “things to ask a software agency” checklist is calibrated for pre-AI work. It covers references, portfolio, pricing, and NDA. It misses eval discipline, prompt and weight ownership, change-order mechanics for AI scope drift, and what happens when the model the build was scoped around gets deprecated in week 6. The result is a checklist that looks rigorous and is operationally hollow. This article is the 12-item list I hand founders in the final 5 working days before they sign.
The checklist here builds on the idea validation playbook and the broader idea-to-product manifesto. It picks up where how to pick an AI development partner when you have never built software and the contract anatomy in what a defensible idea-to-product SOW looks like, with examples leave off. Reference-call mechanics are covered in the AI agency reference call: 11 questions that surface real client outcomes.
Table of Contents
- Why a Pre-Signing Checklist
- The 12-Item Checklist
- How to Use the Checklist in 5 Days
- What a Passing Score Looks Like
- Frequently Asked Questions
- Closing
Why a Pre-Signing Checklist
Shortlisting filters 20 candidates to 3 — testing for archetype, price band, and chemistry. Pre-signing is the last gate before the contract is countersigned. Shortlisting asks “is this partner plausible?” Pre-signing asks “is this engagement, as currently structured, signable as-is?”
Vendor pressure peaks in the final week. The proposal is in, the kickoff date is on the calendar, switching costs feel high. This is when the wrong items get waved through — the eval clause that “we will finalize in week 1,” the IP carve-out that “is standard language,” the on-call window that “we always cover, we just have not written it down yet.” A pre-signing checklist converts every reassurance into a contract line before signature.
McKinsey’s State of AI 2025 found 78 percent of organizations use AI but only a small minority capture enterprise value. BCG’s 2024 analysis found 74 percent of companies see limited value from AI. Most builds ship and then fail to produce ROI because the quality bar was never contracted. Pre-signing is the cheapest moment to fix it.
The 12-Item Checklist
Each item follows the same structure: what to ask, what good looks like, what red flag looks like. Score each 1 or 0.
1. Team identity
What to ask: “Send LinkedIn profiles of the specific people on this engagement with their last three projects, and add a clause to the SOW naming each of them with founder approval required for any swap.”
What good looks like: A response within 48 hours with named senior people, real profiles, three recent shipped projects, and willingness to add the named-team clause.
What red flag looks like: “We will assign the right team at kickoff.” Standard bait-and-switch — the senior people from sales are not the ones executing.
2. Prior work artifacts
What to ask: “Send me a redacted eval rubric, a redacted post-mortem, and a runbook from your most comparable past engagement.”
What good looks like: Three real artifacts within a week, client names removed but structure preserved. The eval rubric names cases and a pass-rate threshold. The post-mortem names what went wrong and how it was fixed. The runbook is operationally usable.
What red flag looks like: A case-study deck instead of artifacts. A deck describes the engagement after the fact; artifacts are what got produced during it.
3. Eval discipline in the contract
What to ask: “Will you attach an eval suite with a pass-rate threshold to the SOW as a milestone-payment gate, and what happens financially if you miss the threshold?”
What good looks like: Yes — 80 to 150 cases seeded in planning, a contracted pass rate at the build milestone, and a payment-withholding or rework provision if missed.
What red flag looks like: “We will define quality collaboratively as we go.” The single highest-impact item on the checklist. A partner who refuses contracted evals shifts all quality risk onto you.
4. AI-specific IP ownership
What to ask: “Walk me through the IP-assignment clause for each of these four categories: application code, prompt library, eval suite, and any fine-tuned weights or adapters.”
What good looks like: A clean assignment clause covering all four categories on milestone payment, plus a schedule of the partner’s pre-existing tools with a royalty-free license back to the partner on those tools only.
What red flag looks like: Silence on prompts and evals, or a broad carve-out for “frameworks, methods, and templates.” A founder who signs away the prompt library has paid full price for half the asset.
5. Pricing model and milestone structure
What to ask: “Show me each milestone, its fixed price, its deliverable list, and the go/no-go gate that triggers the next payment.”
What good looks like: Three named milestones — planning (roughly $30K, weeks 1–3), build (roughly $80K, weeks 4–10), hardening (roughly $40K, weeks 11–13) — with fixed prices, deliverables, and gates naming a specific artifact and a kill criterion.
What red flag looks like: T&M with an estimate range, sprints with “scope refined collaboratively,” or a single all-in price. A founder cannot exit a runaway engagement without milestone gates.
6. Change-order mechanics
What to ask: “What counts as a scope change versus normal iteration, and what is the change-order process — pricing, approval, timeline impact?”
What good looks like: A written clause distinguishing (a) discovery-driven adjustment within planning, (b) AI-specific drift like a model swap or threshold lift, and (c) genuinely new features — each with a written change order, founder signature, and a dollar threshold above which the build pauses for approval.
What red flag looks like: “We handle change orders flexibly,” or no change-order language at all. AI builds reliably encounter scope drift; without contracted mechanics, every drift becomes a contentious negotiation.
7. On-call commitment post-handoff
What to ask: “What is your on-call commitment in the 30 days after handoff — in the SOW, with specific hours, response time, and a transition plan?”
What good looks like: A specific window — 30 days of weekday business-hours coverage with same-day response on production incidents — in the SOW with named contacts and a transition to either a maintenance retainer or an in-house engineer.
What red flag looks like: “We stand behind our work” without hours, response times, or duration. If “on-call,” “response time,” and “post-handoff” are not in the SOW with numbers, the commitment does not exist.
8. Kickoff agenda
What to ask: “Send me the agenda for the first 5 working days, hour by hour where it matters, with the deliverable at the end of day 5.”
What good looks like: A written agenda covering stakeholder interviews, workflow mapping, success-criteria definition, eval-seed brainstorming, and a day-5 deliverable like a workflow map or feasibility memo. The partner has run this enough times to have a templated agenda. See anatomy of a great AI agency kickoff for the benchmark.
What red flag looks like: “We will design the kickoff together at the start.” A partner who has not pre-formatted week 1 is improvising the most expensive week of planning.
9. Week-1 deliverable as a kill gate
What to ask: “What is the specific deliverable at the end of week 1 that I could cancel the engagement against if it is not credible?”
What good looks like: A named artifact — feasibility memo, workflow map, or PRD outline — with a contractual provision that the founder can terminate within 5 business days of receipt at a pre-negotiated kill fee (typically 30 to 50 percent of the planning milestone), with all artifacts to date transferring to the founder.
What red flag looks like: “The planning milestone is a single unit, we do not break it up.” A partner unwilling to expose a week-1 gate is asking the founder to commit the full planning fee on faith.
10. Escalation path
What to ask: “When the engagement is off-track and the weekly standup is not fixing it, who do I escalate to, how fast does that path respond, and what is the dispute-resolution language?”
What good looks like: A named senior person above the engagement manager — typically a partner or principal — with a 48-hour response commitment, and a written escalation clause in the MSA naming a mediation step before litigation.
What red flag looks like: “Talk to your engagement manager” with no named alternative. Routing the escalation back to the person whose performance you are escalating is a non-process.
11. Termination economics
What to ask: “If I terminate at the end of planning, at week 6, or at week 10, what do I owe, what do I take with me, and what does the IP transfer look like in each case?”
What good looks like: A termination matrix in the SOW — per-milestone exit fees (often a percentage of the next unstarted milestone), full IP transfer of work to date on payment of fees due, and a defined handoff window (5 to 10 business days) for the partner to deliver a partial-build runbook.
What red flag looks like: “We do not really do termination — our clients always finish.” Translation: no exit, and the partner intends to use schedule and switching-cost pressure to keep you in. A founder who cannot exit cleanly is structurally captive.
12. Reference clients
What to ask: “Send me three references — PMs or product owners, not executive sponsors — with calendly links to schedule this week. I want to ask what was contracted versus delivered, where the engagement ran over, what broke in the first 30 days, and what they would contract differently.”
What good looks like: Three names within 24 hours, calendly links live, references answer substantive questions in under 20 minutes with specifics that match what the partner described. The 11-question reference script is in the AI agency reference call: 11 questions that surface real client outcomes.
What red flag looks like: Stall, only-executive references, or coached responses (“they were great to work with”) without specifics.
How to Use the Checklist in 5 Days
The checklist is designed to run inside the final week — long enough to surface signal, short enough not to lose the engagement window.
- Day 1 — Send the diligence email. One email, 12 questions, a 48-hour deadline for written responses on items 1, 2, 3, 4, and 6.
- Day 2 — Schedule reference calls. Send references the question set in advance. Two calls per day across days 3 and 4.
- Day 3 — Contract red-line pass. Read the SOW with the checklist in front of you. For each item, find the corresponding contract line. Items with no contract line are sales reassurance, not commitments.
- Day 4 — Reference calls. Inconsistencies between the references’ description and the partner’s pitch are signal.
- Day 5 — Score and decide. Score each item 1 or 0. Apply the rubric below.
For a $150K engagement, the cost is a few hours of founder time plus the $1,000 to $3,000 paid proposal review covered in the partner-selection rubric — roughly 1 to 2 percent of contract value as insurance against a 30 percent probability of a six-figure mistake.
What a Passing Score Looks Like
Twelve items, scored 1 or 0 each.
10 to 12 of 12 — Sign. Structurally credible engagement.
8 to 9 of 12 — Renegotiate the missing items. Send a single amendment request covering the 0-scored items. A partner who fixes 2 to 3 of them is signable. A partner who refuses is sending a second signal.
6 to 7 of 12 — Pause. Gaps that compound during execution. Insist on amendments, or restart the search. Signing into a 6-of-12 engagement is the modal path to a runaway build.
5 or below — Pass. The partner has not built engagements like this enough times to have the contract muscle or the operating discipline. Restart the search.
Kill-the-deal items — where a 0 score is a pass regardless of total — are item 3 (eval discipline), item 4 (AI-specific IP), and item 7 (on-call commitment). No compensating strength closes those three risks.
Frequently Asked Questions
How is this different from the standard “10 things to ask a software agency” list?
The standard list was written for pre-AI work — references, portfolio, pricing, NDA, IP-as-code-only. This list adds the items specific to AI builds in 2026: eval discipline as a contracted artifact with a pass-rate threshold, four-category IP covering prompts and eval suites and fine-tuned weights, change-order mechanics for AI scope drift, and on-call language with concrete hours and response SLAs.
When should I run this checklist?
In the final 5 working days before signing, after shortlisting is complete and the proposal is in. Running it earlier surfaces noise. Running it later means the contract has already been signed and the negotiating position is gone.
What if the partner refuses to add a contract clause?
Refusal is the signal. A partner with the operating discipline to deliver an item is willing to write it into the contract — the language already exists in their template library because they have signed it before. Treat refusal as a 0 score.
Are the three kill-the-deal items always non-negotiable?
Yes. Eval discipline (item 3), AI-specific IP (item 4), and on-call commitment (item 7) protect against the three highest-frequency, highest-magnitude failure modes — ship-without-quality, lose-the-prompt-library, no-coverage-when-production-breaks. No compensating strength on the other 9 items closes those three risks.
How does this checklist relate to the SOW anatomy piece?
The SOW anatomy describes the document structure — 9 sections and the language that should be present. This checklist is the diligence pass against that SOW — for each section, is the substance real, and does the partner have the operating discipline behind it.
Can a partner pass the checklist and still be a bad fit?
Yes — the checklist covers engagement structure, not chemistry, domain fit, or working style. Run the checklist first to filter for structural credibility, then apply the soft factors.
What if I am running the checklist with multiple finalists at once?
Score each 1 or 0, then compare totals. Where two partners tie, break the tie on the three kill items first, then on the four artifact items (1, 2, 8, 10), then on the five commercial items (5, 6, 9, 11, 12).
Does SFAI Labs pass its own checklist?
Yes — we publish it because we structure engagements this way. Three milestones with fixed prices and gates, eval suite attached to the contract, four-category IP assignment on payment, named team in the SOW, written change-order process, 30-day on-call window, week-1 kill gate, and a referenced 11-question reference call. We score 12 of 12 because we engineer the engagement to clear it.
What is the next step if my current partner is failing the checklist?
If you are mid-evaluation, restart the search. If you have already signed, read the 10 rules of working with an AI agency without losing leverage and bring in a fractional CTO. If you have not yet signed and the partner is willing to redline, send a single amendment request covering all 0-scored items and rerun the checklist in 5 business days.
Closing
Pre-signing is the cheapest moment in the engagement to fix the structure. A clause added in red-line costs a few hours; the same clause negotiated mid-build costs the schedule, the relationship, and frequently the contract. A founder who runs this 12-item pass before signature is doing the work that would otherwise show up in the post-mortem.
The checklist converts the partner choice from a momentum decision into a structural one. When the engagement clears 10 of 12 with the three kill items green, the signature is safe. When it clears 6 of 12 with one kill item red, the engagement is not signable as-is and the partner needs to fix the structure or lose the deal.
When you are ready to walk a real proposal through the 12 items, the idea review is the next step — a 30-minute structured conversation where we run your draft SOW against the checklist and tell you, in plain language, whether it is signable as-is, signable with redlines, or not signable. If it is the third, we will say so even when the partner under review is us.
Arthur Wandzel