Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 15 min read

How to do due diligence on an AI MVP partner's prior work

How to do due diligence on an AI MVP partner's prior work

Due diligence on an AI MVP partner is a sales-call procedure, not a portfolio review. Once a partner hands you two or three “similar prior projects” to inspect, the diligence is a seven-question script run live against each project — with explicit thresholds for what a strong answer sounds like, what hand-waving sounds like, and what should kill the deal. The portfolio is the hypothesis. The script is the test. Non-technical founders who walk into this conversation without the script end up evaluating presentation skill rather than engineering discipline, and they sign a contract whose downstream cost overruns are visible in the sales call if anyone is listening for them.

This builds on the AI MVP economics playbook. It runs upstream of the partner’s-prior-work artifact request procedure and pairs with the pre-signing engagement checklist. For background on why the marketing portfolio itself is unreliable, see decoding the AI agency case study. This piece gives you the seven questions to run in the sales conversation before any of those followups become necessary. It belongs to the broader idea-to-product manifesto.

Table of Contents

Why a Sales-Call Script Beats a Portfolio Review

A portfolio review measures what the partner chose to show. The seven-question script measures what the partner has actually built. Only the second predicts engagement quality.

The AI agency category since 2023 produced an oversupply of partners whose marketing materials are indistinguishable — each lists “production AI deployments,” “eval-driven development,” and “post-launch support.” McKinsey’s State of AI 2025 reported 78 percent of organizations using AI with value concentrated in a small fraction of deployments, while 2024–2025 Gartner CIO surveys put AI pilots stalling before production at roughly 85 percent. Most projects in any portfolio are pilots that never reached production, and marketing rarely says so. A script run live against two or three named prior projects forces the partner to commit to facts that you can verify in followups.

The script has seven questions. Each maps to one engagement risk: scope ambiguity, demo-vs-shipped, false references, knowledge transfer failure, cost dishonesty, post-launch abandonment, IP entanglement. For each: what to ask, what good sounds like, what hand-waving sounds like, what to walk on.

Question 1 — Eval Set, Threshold, and Contract

The question. “For each of these two prior projects, did you have an eval set with a documented pass threshold? Was the threshold attached to the contract as an acceptance criterion or to a payment milestone?”

What good sounds like. A specific answer. “On project A we had 240 cases across three task categories, GPT-5 as judge, threshold of 92 percent on category one and 85 percent on category two, and the threshold was a payment milestone — final 20 percent unbilled until we hit it. On project B we used a smaller set of 80 cases because the task was narrower.” Good answers come from people who have lived in eval rubrics. They name model judges, case counts, and thresholds without consulting notes.

What hand-waving sounds like. “We tested thoroughly throughout.” “We had a comprehensive QA process.” “A mix of manual review and automated testing.” These are sentences from people who did not run an eval-first build — they describe testing as a generic activity rather than a contracted gate. A partner who built an eval set names cases, scoring function, and threshold within thirty seconds, because the artifact exists in their working memory.

What to walk on. Two things. First, inability to name a numeric threshold for any prior project — if the threshold does not exist, the engagement was negotiated post-hoc rather than contracted up-front. Second, refusal to attach a threshold to your contract. Eval-first partners accept threshold-gated milestones; non-eval partners refuse or stall, citing “different shape of work.” Either is a walk. Eval discipline is the most predictive single behavior in AI partner selection.

Question 2 — Deployed URL and Live Users

The question. “Of the projects on your portfolio page, which ones are currently in production with paying users or active internal use, and which were prototypes or pilots that did not ship? Can you give me a deployed URL or login-walled screenshot for each shipped one?”

What good sounds like. A clear honest split. “Three are in production — here are URLs and rough user counts. Two are pilots; one stalled at the eval stage because the client’s data quality was insufficient, the other was a feasibility build that did not advance to production.” Honest partners volunteer the pilot/shipped distinction without prompting because it is core to how they understand their own work.

What hand-waving sounds like. “All of these were successful engagements.” “We delivered against scope every time.” “Most are still being used internally.” The pattern is to assert success without naming operational status. Pilots and shipped products are different categories. A partner who treats “worked in a demo” as equivalent to “in production serving users” will misrepresent your build to your investors.

What to walk on. Three or more “deployed but I cannot share the URL” answers across the portfolio. One or two NDA-restricted projects is plausible. A whole portfolio under NDA usually indicates a very limited body of actually-shipped work or work too recent to have finished. Either way, you cannot verify the production claim and should not pay production prices.

Question 3 — Client Reference With a Direct Number

The question. “Can you put me on a thirty-minute call with the technical buyer at one of those projects? The CTO, head of engineering, or product lead — not their marketing contact. I will not record the call, and I will share my questions in advance.”

What good sounds like. A name and a number within 48 hours. Strong partners have warm relationships with prior clients and broker reference calls quickly. The reference is a technical buyer who can speak to the engineering relationship — not a CEO who saw quarterly reports. The call should yield three answers: did the partner ship on time, did the budget hold, would the buyer hire them again. A reference who hedges on any of the three is signal.

What hand-waving sounds like. “The client is very busy.” “Most of our clients have NDAs that restrict reference calls.” “We can give you a written testimonial instead.” These deflect without naming a specific blocker. A strong reference relationship survives a thirty-minute call. A partner who cannot produce one across a portfolio either ended past projects badly or has overstated the client relationship.

What to walk on. No live reference within seven days, or a reference call that contradicts the partner’s claims about timeline, budget, or scope. A reference saying “they shipped late but it was fine” while the partner claimed on-time delivery is a walk — the partner is now demonstrated to misrepresent prior projects to prospects, and will do the same about your project to their next prospect.

Question 4 — Handoff Artifact From the Prior Project

The question. “Send me a redacted version of the handoff artifact from one of these projects. The runbook, ops doc, or post-mortem you gave the client at the end. I want to see what knowledge transfer actually looks like with you.”

What good sounds like. A document arrives. It contains a system diagram, an inventory of which prompts do which work, named eval cases the client can rerun, operational risks with mitigations, and a recommended on-call procedure with thresholds. The shape says: this is how we leave a client. Production-grade partners have this artifact ready because it exists for every project — redaction removes the client name but preserves the structure.

What hand-waving sounds like. “We do extensive knowledge transfer.” “We have a comprehensive offboarding process.” “Our documentation is detailed but client-specific.” Statements about the process rather than the artifact. Partners who have shipped to production have a handoff document; partners who have shipped only pilots do not — pilots end with a slide deck, not a runbook.

What to walk on. No handoff artifact across the whole portfolio, or an artifact that is a sales deck dressed as documentation (presentation rather than operational runbook). Pilots-as-runbooks signal a partner who has not run a production handoff and will leave you holding a system you cannot operate. Handoff is the second-most-predictive document after the eval rubric; its absence is a walk.

Question 5 — A Project They Overran and What They Changed

The question. “Tell me about a project where you overran the budget or timeline. What did you change in how you scope or run engagements afterward?”

What good sounds like. A specific story. “On project X in 2024 we underestimated data cleaning because we accepted the client’s claim that the dataset was production-ready. We were 30 percent over on timeline and absorbed the cost. Since then we run a one-week paid data audit before any contract above $75K, and we name dataset readiness as an acceptance criterion in the SOW.” Specific failure, specific lesson, specific process change.

What hand-waving sounds like. “We have not had any significant overruns.” “We always deliver on scope.” “If we ever go over it is because the client added requirements.” The first two are statistically unlikely across more than five projects. The third is a partner who externalizes every overrun onto the client. Real engagements overrun — a partner who claims none either has few engagements to draw from or is unwilling to discuss failure with prospects.

What to walk on. The “we never overrun” answer, or any answer that blames every prior overrun on the client. The first signals dishonesty. The second tells you exactly how your own overrun will be framed mid-project: as your fault. Either is a walk.

Question 6 — Post-Launch On-Call History

The question. “After you shipped these projects, did you have an on-call window? For how long? What was your response-time SLA, and what did you actually do during the window?”

What good sounds like. A clear scope. “On project Y we ran a 30-day on-call window on the client’s PagerDuty. Four pages during the window — two prompt-drift issues from a model update, one client-side data ingestion failure they fixed themselves, one token-cost anomaly we traced to a stuck retry loop. Average response time 22 minutes.” Specific window, incident counts, response data. Partners who have operated their own builds in production have this data because they lived it.

What hand-waving sounds like. “We are always available.” “We have a strong post-launch relationship.” “We typically respond within a business day.” These describe an availability posture, not an on-call discipline. On-call has a start, an end, a paging mechanism, an SLA. A partner who has run one describes it as a contract; one who has not describes it as a feeling.

What to walk on. No defined on-call window across any prior project, or a window that ended on launch day. Both signal a partner who treats launch as the end of the engagement rather than the beginning of the riskiest two weeks. An AI MVP produces its first production failures inside 21 days — prompt drift, cost anomalies, edge cases absent from the eval set. A partner not present during that window leaves your team to fight the failures alone.

Question 7 — IP and Source-Code Ownership Terms

The question. “What do your standard contract terms say about source-code ownership, prompt templates, eval rubrics, and any internal tooling you used to build the project? Show me the IP clause from one of these prior contracts.”

What good sounds like. A clear answer with the clause. “The client owns all source code, prompts, and eval rubrics generated during the engagement. We retain ownership of our internal scaffolding — build tooling, project-management templates, and any open-source contributions we make. Here is the redacted clause from project Z.” Strong partners ship their work as fully transferable IP because they have no business model that requires retaining it.

What hand-waving sounds like. “We can discuss IP in negotiation.” “Our standard terms are negotiable.” “We typically retain some IP rights for our framework.” The first two stall. The third is the dangerous one — it usually means the partner has built a proprietary framework that runs inside your application and to which they hold the rights, creating a lock-in that surfaces months later when you try to switch partners or take development in-house.

What to walk on. Any retained IP that runs in production. A partner who keeps ownership of code that runs in your product has an option on your future engineering budget that the contract does not name. The walk is not “IP rights are negotiated” — those negotiations are fine. The walk is a partner who treats their proprietary framework as non-negotiable and will not show you what is in it. The clause must let you move the code, eval rubrics, and prompt assets in-house or to another partner without renegotiation.

How to Score the Conversation

Run all seven questions in a single 60-minute call against two prior projects. Score each 0 (walk), 1 (acceptable), or 2 (strong).

Question Walk-on-zero Threshold
1 — Eval set, threshold, contract Yes Must be 2
2 — Deployed URL, live users Yes Must be 2
3 — Reference call Yes Must be at least 1
4 — Handoff artifact Yes Must be at least 1
5 — Overrun honesty No At least 1 across portfolio
6 — Post-launch on-call No At least 1 across portfolio
7 — IP and ownership Yes Must be 2

A partner scoring 1 or 2 on every question with no zeros is a strong candidate. A single zero on questions 1, 2, 3, 4, or 7 is a walk regardless of the other six. The eval question is the heaviest — if the answer is hand-waving, end the conversation politely. The script makes three or four partners comparable; without it, sales conversations diverge as each partner steers to their strongest topic, and the founder chooses on rapport rather than discipline.

Frequently Asked Questions

How is this different from “check the portfolio and call references”?

The portfolio shows what a partner chose to publish. The script forces commitment to specific facts (case counts, thresholds, URLs, reference names, document shapes, on-call incidents, IP clause language) that you can verify in followups. It moves diligence from marketing review to fact-checking.

Can a non-technical founder really run this script?

Yes. The questions are designed so answer quality is detectable without engineering background — specific numbers and named artifacts indicate competence, vague process language indicates the opposite. Pattern recognition is in the shape of the answer, not its technical content.

How long does the full diligence loop take?

The script runs in 60 minutes per partner. Followup artifact requests (eval rubric, handoff document, reference call) take 5 to 8 business days. Total elapsed across three partners: about two weeks, ideally run in parallel with broader procurement.

What if a partner refuses to answer Question 5 (overruns)?

Walk. A partner unwilling to discuss prior failures with a prospect will not discuss failures with you mid-engagement either. Willingness to name a specific overrun is the cheapest signal of professional maturity available in the diligence.

What about partners with limited prior work (one or two projects)?

The script still applies. Run it against what they have and add a feasibility engagement ($5K to $15K, two weeks) before the full build. Newer partners can still score strongly on eval discipline and IP clarity; what they cannot offer is portfolio breadth, and the feasibility window substitutes for that.

What if the partner has shipped many small projects but none at our scale?

Ask Question 1 (eval discipline) and Question 4 (handoff artifact) particularly hard. Scale-up risk is largely a project-management risk, not an engineering-discipline risk. Eval and handoff discipline at small scale usually transfers up; lack of it at small scale compounds at larger scale.

How do strong partners react to the script?

They welcome it. The seven questions favor partners with disciplined engagement practice and disadvantage partners running on marketing. Strong partners gain a competitive advantage from buyers running the script; weak ones lose deals they were going to lose anyway, just later and more expensively for the founder.

Does SFAI Labs go through this script with prospective clients?

Yes. The first conversation typically includes the script. We name prior projects with eval thresholds and case counts, share redacted handoff artifacts where NDA permits, and walk the IP clause line by line. The diligence runs identically against us as against any other partner.

Closing

The marketing portfolio is a hypothesis. The seven-question script is the test. Run it in one hour, against two named prior projects, with the walk thresholds in hand. A partner who survives has run the engagement model that ships AI products at quality. A partner who hand-waves has revealed exactly what the cost overrun will look like in week six — visible in the sales call, if anyone was listening. The script costs one hour. The cost of skipping it is the next engagement gone wrong.

Last Updated: Jul 25, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles