A non-technical founder picking an AI development partner is in a structurally weaker position than a CTO running the same procurement, and the standard “how to pick a dev shop” checklist does not close the gap. The CTO can read the proposal, audit the eval suite, and notice when a stack choice is wrong. The non-technical founder cannot — and is still expected to sign a $100K to $300K engagement that defines the product surface for the next two years. This article is the rubric I give non-technical founders who do not have a CTO, do not want one yet, and need to make this call in the next three weeks.
It builds on the idea validation playbook and the broader idea-to-product manifesto. The solo-founder partner guide covers archetypes; the CTO-versus-idea-to-product piece covers the buy-or-hire fork. This piece is a non-technical buyer’s evaluation rubric, scored in 30-minute blocks, with explicit rules for what a non-engineer can and cannot verify alone.
Table of Contents
- Why the Standard Checklist Fails Non-Technical Buyers
- Five Signals a Non-Technical Buyer Can Verify Without Code
- What You Cannot Verify Alone — and the Cheap Insurance Pattern
- The 7-Question Vendor Diligence Checklist
- Three Red Flags That Should Kill the Engagement Before Signing
- What Good Looks Like at Each Milestone
- A 30-Minute Self-Diligence Sprint
- Frequently Asked Questions
- Closing
Why the Standard Checklist Fails Non-Technical Buyers
The generic “how to pick a software development company” article was written for a buyer who can read a SOW, interpret a Git history, and ask informed questions in a technical interview. A non-technical founder cannot. The checklist still asks them to “verify AI expertise,” “review the technical proposal,” and “assess the architecture” — exactly the skills the buyer is trying to rent. The result is a checklist that looks rigorous and is operationally hollow.
Two things changed in 2025 that make this gap larger. First, the AI build category matured into a specific shape — milestone-based fixed pricing, evals attached to the contract, a hardening milestone before handoff. The weakest engagements still sell time-and-materials sprints with no quality bar. A technical buyer reading the SOW can tell the difference inside ten minutes; a non-technical buyer often cannot tell at all.
Second, AI quality is hard to verify on demos. McKinsey’s State of AI 2025 found that 78 percent of organizations use AI, but only a small fraction capture meaningful enterprise value. The 2024–2025 Gartner CIO surveys put AI pilots stalling before production at roughly 85 percent. The common post-mortem: the demo looked good, the production system did not, and there was no eval suite contractually attached to define “good” before the build began.
Five Signals a Non-Technical Buyer Can Verify Without Code
These are the five signals that — verified together — give a non-technical buyer roughly 80 percent of the confidence a CTO buyer would get from a full technical diligence pass. Each is scoreable in 30 minutes by a non-engineer.
Signal 1 — Eval Rubric Template
Ask for a real eval rubric from a past engagement, with the client name redacted. Strong: a one- to three-page document listing test cases, inputs, expected outputs, scoring method, and the pass-rate threshold attached to the contract — in your inbox within 48 hours. Weak: a generic “QA approach” deck or vague reassurance about “thorough testing.”
You do not need to understand the test cases — you need to confirm the artifact exists, names cases, names a threshold, and was attached to a real contract. This signal predicts more about engagement success than any other. A partner who cannot produce one has not run an eval-first build, and the post-mortem corpus is clear about what happens when AI builds ship without one.
Signal 2 — Reference Call Substance
Ask for three references — specifically PMs or product owners, not executive sponsors. Executives say the polite, brand-safe version; PMs say what actually happened. The substance signal is whether the reference can answer five questions in under 20 minutes: what was contracted versus delivered, where the engagement ran over schedule, what broke in the first 30 days, what the total bill was versus the estimate, and what they would contract differently. A reference who deflects or stays at “they were great to work with” is coached, a friend, or did not actually engage the work. The full version is in the AI agency reference call: 11 questions that surface real client outcomes.
Signal 3 — SOW Clarity at the Milestone Boundary
Strong SOWs name each milestone, its fixed price, its deliverables, and a go/no-go gate: planning weeks 1–3 at $30K (PRD, feasibility memo, workflow map, eval suite seed of 80 to 150 cases, unit-economics worksheet, risk register, kill criterion); build weeks 4–10 at $80K (working application against the eval suite at the contracted pass rate, observability, deployment runbook); hardening weeks 11–13 at $40K (production deployment, 30-day on-call, knowledge transfer). Weak SOWs read as “agile engagement, two-week sprints, scope refined collaboratively” with a total estimate range. Read it to a non-technical friend; if they cannot describe what gets delivered and when, the SOW is too vague to sign.
Signal 4 — IP Ownership Terms
The contract should transfer four IP categories on milestone payment: application code, prompt library, eval suite, fine-tuned weights. The first three are non-negotiable for any modern AI build. Strong: a clean assignment clause naming all four, with a one-page schedule of the partner’s pre-existing tools and a royalty-free license back to the partner on those tools. Weak: assignment of “application code” only, silence on prompts and evals, or an open-ended carve-out (“partner retains rights to all frameworks, methods, templates”) — a back-door claim on most of what you paid for. A founder who signs away the prompt library and eval suite has paid full price for half the asset.
Signal 5 — On-Call Commitment Post-Handoff
Strong: a specific on-call window in the SOW — 30 days of weekday business-hours coverage with same-day response, then a defined transition to a maintenance retainer or an in-house engineer. Weak: vague reassurance (“we stand behind our work”) without hours, response times, or duration. Search the SOW for “on-call,” “response time,” and “post-handoff.” If those three terms are not in the contract with concrete numbers, the commitment does not exist.
What You Cannot Verify Alone — and the Cheap Insurance Pattern
Three things a non-technical founder genuinely cannot verify alone: whether the proposed architecture is sensible (model choice, orchestration depth, data store defaults), whether the eval suite is real or theater (right inputs, scoring substance, threshold calibrated above the contract floor), and whether the unit economics will hold (inference cost per query, latency budgets, fallback patterns).
The cheap insurance pattern: pay a senior engineer $1,000 to $3,000 for a 4- to 8-hour proposal review before signing. Sourcing channels — a fractional CTO via introduction or an hourly platform, a senior engineer in your network at a credible company, or a paid second opinion from a different partner archetype than your shortlist. The brief is one page:
Read this proposal. Tell me three things. One: is the architecture appropriate for the scope? Two: is the eval suite real, or theater? Three: are there three to five questions I should send back to the vendor before signing?
A $2K spend on a 6-hour review of a $150K engagement is a 1.3 percent insurance premium against a 30 percent probability of a six-figure mistake. The expected-value math is overwhelming. Most non-technical founders skip this step because nobody told them it was available; it is the single highest-return purchase in the diligence process. The reviewer does not need vertical specialization — they need to be a senior engineer who has shipped AI in production and will tell the founder the truth without trying to win the engagement themselves.
The 7-Question Vendor Diligence Checklist
A non-technical founder cannot run a 30-question diligence call — they will run out of time and the vendor will run out the clock. Seven questions, asked precisely in a 45-minute call, surface roughly 80 percent of the relevant truth. Score each 1 or 0.
- “Send me the eval rubric from your most comparable past engagement, redacted.” Strong: arrives within 48 hours with cases, scoring method, and pass-rate threshold. Weak: generic, late, or absent.
- “Walk me through the milestones — deliverable, fixed price, and go/no-go gate at each.” Strong: three named milestones, fixed prices, gated. Weak: T&M, sprint-based, undifferentiated.
- “If planning shows the build is not feasible at our budget, do we exit clean?” Strong: yes, no further obligation, founder owns the planning artifacts. Weak: contract forces a buildout, or the founder forfeits planning IP on exit.
- “What does the IP-assignment clause say about prompts, eval suites, and any fine-tuned weights?” Strong: a clean paragraph naming all categories. Weak: silence on prompts and evals, or a broad partner carve-out.
- “Who specifically will be on the engagement? Send me their LinkedIn profiles and last three projects.” Strong: named senior people, real profiles, three recent projects. Weak: “to-be-staffed” or “we will assign at kickoff.”
- “What is your on-call commitment in the 30 days after handoff — in the SOW, with hours and response times?” Strong: yes, named. Weak: vague reassurance, no SOW language.
- “Give me three references — PMs or product owners, not executive sponsors. I want to call them this week.” Strong: three names with calendly links within 24 hours. Weak: stall, push back, or only-executive references.
Six or seven is a credible partner. Four or five is workable with a paid proposal review. Three or below is a pass.
Three Red Flags That Should Kill the Engagement Before Signing
Most red-flag lists online are generic — “they were hard to reach during sales,” “their portfolio is thin,” “they pushed a particular technology too hard.” These are signals, but they are not kill-the-deal signals. The three below are.
Red Flag 1 — They refuse to attach an eval suite to the contract. Any version of “we will define quality collaboratively as we go” means the engagement has no quality bar. Mid-build, you have no contractual basis to ask for a fix. At handoff, the vendor’s defense is that no specific bar was contracted. A partner who refuses contracted evals is shifting all the risk to you. There is no compensating strength.
Red Flag 2 — They will not name the people on the engagement. “We will assign the right team at kickoff” is almost always a bait-and-switch. The senior people from the sales process are not the ones who execute. The juniors who get assigned are competent, but the engagement was priced as if seniors were running it. The fix is to name specific people in the SOW with a clause requiring founder approval for swaps.
Red Flag 3 — They cannot articulate failure modes for AI products. Ask: “What are the three most common ways an AI product like ours fails in production in the first 60 days, and how does your engagement protect against each?” A strong partner names specific modes — quality drift on a model update, cost-per-query bloat from a retrieval-side bug, prompt-injection risk in a customer-facing surface, eval-gap discovery after launch — and names structural protections. A weak partner generalizes (“we use best practices”). They have not run enough AI builds to have a failure-mode taxonomy; the engagement is their on-the-job training, paid by you.
These three are kill signals. No archetype, brand, price point, or reference compensates.
What Good Looks Like at Each Milestone
A founder who only learns the engagement was off-track at the post-mortem has lost the ability to course-correct. The checklist below is the one a founder can run unaided at each milestone boundary.
End of Planning — Week 3. Strong: a 15–25 page PRD a senior engineer can estimate from, eval suite seed of 80–150 cases, unit-economics worksheet at three usage tiers, risk register with mitigations, and a written kill criterion. Weak: a “discovery deck” of slides, eval cases promised for build, no unit economics, no kill criterion. If planning is weak, do not approve the build — it compounds the gaps at five times the cost.
Mid-Build — Week 6 or 7. Strong: a working demo against the eval suite with current pass rate visible, observability dashboards queryable by the founder, weekly written status naming what shipped, blocked, at risk, and the vendor proactively flagging eval failures. Weak: happy-path screenshots, marketing-shaped status, observability deferred, failures hidden until asked. The last cheap intervention point.
End of Build — Week 10. Strong: contracted eval pass rate hit in a recorded run, all four IP categories transferred, deployment runbook a competent operator can follow without the vendor, known-issues list with a hardening plan. Weak: pass rate “in progress,” code still on the vendor’s GitHub, runbook requiring vendor execution.
End of Hardening — Week 13. Strong: production deployment running on real traffic for at least seven days, on-call window in effect with runbook and SLA, 20–40 page handoff document, and a written post-mortem. Weak: a “we will fix later” backlog, informal on-call, handoff as a screenshot deck, no post-mortem. A founder running these four checklists has roughly the same mid-engagement signal a technical co-founder would provide.
A 30-Minute Self-Diligence Sprint
For the founder in the middle of a vendor evaluation, the operational sequence is:
- Minute 0–5: Open the SOW and grep for “eval,” “on-call,” “intellectual property.” Any missing or vague is a Tier 1 signal.
- Minute 5–15: Read the milestone section. Can you describe each deliverable in one sentence, with a fixed price and a date? If not, the SOW is too vague to sign.
- Minute 15–25: Send the seven-question diligence email. Set a 48-hour deadline for the eval rubric and the named-team response.
- Minute 25–30: Schedule the three reference calls and book the $2K proposal review.
Thirty minutes at this stage produces more signal than the eight hours of vendor calls that typically follow. A founder who skips it and signs on vibes is the modal source of the post-mortems behind the 85 percent pilot-failure number.
Frequently Asked Questions
How is a non-technical buyer’s diligence different from a CTO’s diligence?
A CTO can verify engineering directly — read the SOW, audit architecture, spot a weak eval suite. A non-technical buyer cannot do any of these alone. The rubric narrows diligence to what a non-engineer can verify (commercial terms, process artifacts, operating cadence, reference substance) and names the one place they cannot self-rely — the technical proposal review.
How much should a non-technical founder budget for the proposal review?
$1,000 to $3,000 for a senior engineer to spend 4 to 8 hours on the SOW, eval rubric, and architecture. Pay through a fractional-CTO platform, a network introduction, or a different partner archetype than your shortlist. On a $150K engagement that is roughly a 1.3 percent insurance premium against a 30 percent probability of a six-figure mistake.
How long should the partner evaluation take?
Three to four weeks. Week 1: shortlist 4 to 6 candidates. Week 2: 30-minute intros and the diligence email. Week 3: working sessions with the top two or three and proposals back. Week 4: reference calls and the paid proposal review. A founder who signs in week 1 is reacting to vendor pressure; a founder still evaluating in week 8 is procrastinating.
What is the single most predictive signal of engagement quality?
Whether the partner will attach an eval suite to the contract with a pass-rate threshold. The highest-confidence predictor of whether the build ships at quality, and the only criterion where no compensating strength exists. A partner who refuses contracted evals is a pass.
Can I verify the partner’s AI expertise without writing code?
Indirectly, yes. You can read the eval rubric, the SOW, a redacted post-mortem from a past engagement, and the partner’s public technical writing. A partner with real AI-build experience produces these artifacts as a matter of course; one without it produces marketing decks.
What if the partner refuses to send a redacted eval rubric?
That is a signal in itself. A partner who has run eval-first builds has redacted rubrics they can share within a day. A partner who delays, deflects, or sends a generic QA approach has not run the playbook they are pitching.
What is the minimum-viable engagement for a non-technical founder?
A 1- to 2-week paid feasibility engagement, around $5,000 to $15,000, producing a feasibility memo, workflow map, and risk register. The lowest-stakes way to test a partner before committing to the planning milestone. A partner unwilling to scope below the $30K planning tier is not calibrated to a non-technical founder’s risk tolerance.
What if I cannot find a partner that clears all five signals?
Optimize for the eval rubric, SOW clarity, and on-call commitment — those three protect a non-technical founder mid-engagement. Signals 4 (IP) and 5 (references) are negotiable in contract revision or repairable through extra diligence. A partner clearing three of five with strong scores is signable with a paid proposal review.
Is the SFAI Labs engagement structured this way?
Yes. Three milestones (planning, build, hardening), fixed price per milestone, eval suite attached to the contract with a pass-rate threshold, named team in the SOW, IP assignment covering code, prompts, and evals on milestone payment, and a 30-day on-call window post-handoff. The first conversation is a 30-minute idea review where the founder walks through the seven-question checklist and we name whether the engagement is a fit. If it is not, we say so.
Closing
The non-technical founder is not at a structural disadvantage because they cannot read code. They are at a structural disadvantage because the standard diligence checklist was written for buyers who can. The rubric above inverts that — it scopes diligence to what a non-engineer can verify, names the one purchase that closes the gap, and gives the founder a procedure that runs in 30 minutes per vendor.
A partner who clears the rubric is signable. A partner who does not is wrong for this buyer — not necessarily bad, but wrong for a founder operating without independent technical verification. Keep looking, or bring in a fractional CTO before signing. The expected-value math on three more weeks of search beats the cost of a six-figure rebuild every time. When you are ready for a structured intake conversation, the idea review is the next step.
Arthur Wandzel