A non-engineer founder owns five things in an AI MVP build: the product spec, customer access, domain expertise, brand voice, and go-to-market. The partner owns five: engineering, eval discipline, model selection, observability, and on-call. Most engagements fail not because either party is incompetent, but because that split is never written down — so the founder either over-engages on engineering choices they cannot defend, or disappears for three weeks and lets the product drift away from the customer it was meant for. This article is the role split, written for a non-technical founder about to sign a 6-to-12-week Statement of Work.
This article extends the founder-AI-partner operating manual, part of the idea-to-product manifesto. Companion reading: how an idea-to-product engagement actually works week-by-week, what is an AI development partnership in plain English, and the companion piece on the first 14 days of an AI agency engagement.
Why role clarity matters more than role talent
A high-talent founder paired with a high-talent partner can still produce a missed launch if the role split is implicit. BCG’s analysis of enterprise AI projects points to organisational clarity — who decides what, when — as a more common failure mode than technical capability. The same pattern shows up at the founder-partner scale.
McKinsey’s The state of AI in 2025 reports 78% of organisations now use AI in at least one business function. Procurement pressure means founders are signing AI build contracts faster than ever, often with their first ever AI development partner. Without a written role split, the first thing that breaks is the founder’s confidence — they cannot tell whether a Monday demo is on track or off-track, and they over-correct in either direction.
The fix is not “trust the partner” or “stay close to the work”. It is a written split, agreed before W1.
The two-column role split at a glance
| Founder owns | Partner owns |
|---|---|
| Product spec (what the system does for users) | Engineering (how the system is built) |
| Customer access (who the partner can talk to) | Eval discipline (how quality is measured) |
| Domain expertise (the rules of the field) | Model selection (which model, what fallback) |
| Brand voice (how the system talks, feels) | Observability (logs, dashboards, alerts) |
| Go-to-market (launch plan, comms, pricing) | On-call (incident response after launch) |
Five responsibilities each. The asymmetry to notice: the founder owns what and who; the partner owns how and with-what. Founders who try to own the how over-engage; partners who try to own the what over-build. Each side has a column they cannot delegate.
What the founder brings
Product spec — 6 to 10 hours total
The founder owns the answer to “what does this system do for the user, and how do we know it did it well?” The partner writes the PRD; the founder signs the PRD. The hour-budget across a 12-week engagement is roughly 6 hours of PRD authoring conversation (W1–W2) plus 4 hours of edge-case PRD review at the W5 eval gate.
The PRD is not a wishlist. It is a contract. A founder who cannot point to the PRD and say “this paragraph is the launch criterion” has not done their job here.
Customer access — 4 to 8 hours total
The founder owns the relationships. The partner needs access to 5–10 real customers over the engagement — for W2 prompt-grounding interviews, W5 eval-set validation, and W9 launch-readiness usability checks. The founder cannot delegate the introduction. A partner cold-emailing the founder’s customers is a sign the founder has dropped this column.
Hour-budget: 4 hours of introductions plus 4 hours of “sit-in” time on the early customer calls so the founder can hear what the partner heard.
Domain expertise — 3 to 6 hours total
The founder owns the rules of the field — what counts as a correct answer in their domain, what counts as a dangerous answer, what the regulator or the buyer will not accept. The partner cannot generate this; they can only encode it. The hour-budget is 3 hours of domain-rules walkthrough in W2 plus 3 hours of eval-set rule-review in W5.
For regulated industries (legal, medical, financial), this column expands. A founder who underweights domain expertise produces a system that passes the eval set the partner wrote and fails the eval set a regulator would write.
Brand voice — 2 to 4 hours total
This is the column most agencies silently drop. AI products do not just behave; they talk. Tone, register, formality, humour, refusal style — these are brand choices that shape the system prompt, the eval rubric, and the user experience from W2 onwards. The founder owns brand voice the way they own the logo and the landing-page copy.
Hour-budget: 2 hours of voice-and-tone briefing in W1–W2 plus 2 hours of voice-review on the W5 prompt outputs. Founders who skip this find themselves at W10 trying to fix tone in production — the most expensive place to do it.
Go-to-market — variable, parallel work
The founder owns the launch plan, customer comms, pricing, and the first-month revenue motion. None of this is on the partner’s calendar. The partner ships a system at W10; the founder ships a business. The hour-budget here is whatever the founder’s other GTM work demands — typically the founder’s largest parallel-track effort outside the engagement.
A partner who folds brand identity, marketing site, or sales onboarding into the SOW is mispricing the engagement and overstepping the column line.
What the partner brings
Engineering — the whole stack
The partner owns the build: code, infrastructure, deployment, repository structure, CI/CD, secrets management, dependency hygiene. A non-engineer founder should not be reviewing pull requests. The partner produces a deploy-ready codebase at W12 with a README the founder can hand to the next engineer.
The diagnostic: the partner can answer “what would it take to add a second engineer to this codebase next month?” in under five sentences. If they cannot, the engineering column is weak.
Eval discipline — the quality system
The partner owns the eval set, the rubric, the baseline run, the threshold definition, and the rerun cadence. Anthropic’s published case studies are clear that the largest quality lift in production AI work comes from eval discipline — looking at failures, not averages. The partner is the one who must hold this discipline week after week.
The founder reviews the eval set (because the rules of the domain are the founder’s column), but the partner runs it. Founders who try to own the eval rubric typically end up with too few items, too narrow a distribution, and a launch threshold that production traffic beats easily.
Model selection — and the fallback plan
The partner owns the choice of foundation model (GPT-5, Claude Opus 4.8, Gemini 2.5 Pro, Llama 3.3, or an open-weights alternative), the choice of provider, the prompt-engineering approach, and the fallback model if the primary fails. The founder does not pick the model. They sign off on a budget (per-call cost, per-month cost) and on a behaviour (latency, refusal posture).
A founder who insists on a specific model name without budget or behaviour grounding is over-engaging.
Observability — the dashboards and the alerts
The partner owns the production observability layer: structured logs, dashboards, alerting thresholds, on-call rotation, and the runbook. By W10 the founder should be able to open a dashboard and see the eval-set pass rate against the last 24 hours of production traffic. If that dashboard does not exist, the partner has under-delivered this column.
On-call — the first 30 days after launch
The partner owns the on-call rotation for at least W11–W12, and typically a 30-day post-launch window. Every production incident gets a written post-mortem within 48 hours. The founder is not on-call; the founder is informed of incidents and reviews the post-mortems.
The four founder failure modes
Most founder-side engagement failures fall into one of four named patterns. Each has a one-line diagnostic and a one-line fix.
1. Under-engaging
Diagnostic: the founder skips the W2 PRD review or sends a delegate to the W5 eval review.
What it looks like: by W8 the partner has built a system that is technically correct against the W1 brief but no longer reflects the customer reality the founder learned in parallel. The product drifts off-target without anyone noticing.
Fix: protect the five show-up days (W1 kickoff, W2 PRD review, W5 eval review, W9 launch-readiness, W10 launch) as non-substitutable founder time. Three hours each. The other 7 weeks can be asynchronous.
2. Over-engaging
Diagnostic: the founder is on every standup, reviewing every commit, and opening Slack threads about model-selection choices.
What it looks like: the partner team slows down because every decision needs a founder sign-off. Engineering velocity drops by half. The founder burns out around W6 and then over-corrects into under-engagement for W7–W12.
Fix: name the quiet weeks (W3, W7, W11) in the SOW. The founder steps back. The partner does deep work. Trust is the operating mode; review is the cadence.
3. No-decisions
Diagnostic: the partner asks “do we ship feature A or feature B in W8?” and the founder responds “let’s circle back next week.”
What it looks like: scope ambiguity compounds. The partner builds both half-way; the eval set fragments; the W10 launch slips. The most expensive failure mode of the four because it looks like prudence and behaves like sabotage.
Fix: set a 48-hour decision SLA in the SOW. Any partner question that needs a founder decision gets one within two business days. If the founder cannot decide in 48 hours, the partner defaults to the option that keeps the calendar — and the founder lives with the choice.
4. Scope creep — founder-side
Diagnostic: the founder adds “and it should also do X” at the W4 Loom review.
What it looks like: the W6 eval threshold becomes unreachable because the surface area grew 30% mid-engagement. The partner either misses the threshold or cuts quality elsewhere to absorb the new scope. Both produce a W10 the founder regrets.
Fix: every scope change goes through the SOW’s change-order clause. New scope means new weeks, new fee, or a swap (the new thing in, an old thing out). Verbal “and also” requests never make it into the build.
Weekly cadence expectations for the founder
The 12-week shape clusters founder time into 5 show-up days (3 hours each) and 7 async weeks (1–2 hours each). Total: ~25 hours of founder time across the engagement.
| Week | Founder mode | Time | What the founder owes |
|---|---|---|---|
| W1 | Show-up | 3h | Kickoff attendance, data-access decisions |
| W2 | Show-up | 3h | PRD sign-off, eval-set rule-review |
| W3 | Quiet | 1h | Async Loom review only |
| W4 | Async | 2h | Loom walkthrough + edge-case feedback |
| W5 | Show-up | 3h | Eval review meeting (the diagnostic week) |
| W6 | Async | 1h | Hardening-report review |
| W7 | Quiet | 1h | Async Loom only |
| W8 | Async | 2h | One integration sign-off call |
| W9 | Show-up | 3h | Launch-readiness checklist walkthrough |
| W10 | Show-up | 3h | Production launch sign-off |
| W11 | Quiet | 1h | Post-mortem review (async) |
| W12 | Async | 2h | Handoff doc + 90-day roadmap review |
A founder who treats the show-up days as immovable and the quiet weeks as protected gets the operating rhythm right. A founder who flips that — skips show-ups, fills quiet weeks with check-ins — produces the engagement most agencies complain about behind closed doors.
How to know your role is being held up well
Three diagnostic questions a founder can ask themselves at the end of each month.
Month 1 (W1–W4): “Can I summarise the PRD in 500 words from memory, and name the 10 hardest eval items?” If yes, the founder is holding the product-spec and domain-expertise columns. If no, the founder is under-engaging.
Month 2 (W5–W8): “Have I made every scope decision the partner asked for within 48 hours, and have I introduced the partner to 5 real customers?” If yes, the founder is holding the customer-access and decision-rights columns. If no, the founder is no-deciding or under-engaging.
Month 3 (W9–W12): “Can I open the production dashboard and read it without help, and have I shipped at least one customer comm under my own brand voice?” If yes, the founder is holding the brand-voice and GTM columns. If no, the founder is over-delegating to the partner — which always comes back as a brand-voice repair bill in Month 4.
Three months. Six diagnostic questions. If any answer is “no,” fix the column before W10, not after.
FAQ
How much time should a non-technical founder budget across a 12-week AI MVP build?
About 25 hours of active founder time across the engagement, clustered into 5 show-up days of 3 hours each (W1, W2, W5, W9, W10) plus around 10 asynchronous review hours across the other 7 weeks. Founders who budget “1 hour a week” underestimate the show-up days; founders who budget “10 hours a week” burn out by W6 and start under-engaging in the back half of the build.
Should the founder pick the AI model the partner uses?
No. The founder signs off on a budget (per-call cost, per-month cost) and a behaviour (latency, refusal posture, tone). The partner picks the model. A founder who insists on a specific model — GPT-5, Claude Opus 4.8, Gemini 2.5 Pro, or an open-weights alternative — without budget or behaviour grounding is over-engaging on a decision they cannot defend. The partner owns model selection and the fallback plan.
What if the founder is also a domain expert — does that change the split?
The domain-expertise column gets larger and the customer-access column gets denser, because the founder is also customer-zero. The partner-side columns do not change. The most common mistake here is the domain-expert founder drifting into the engineering column — reviewing pull requests, opining on libraries, picking the model. That is over-engagement, not strength.
Who owns the marketing site and the launch comms?
The founder, every time. The partner ships the AI system; the founder ships the business. Brand identity, marketing site, sales onboarding, and customer-support tooling are not on the 12-week calendar. A partner who folds them into the SOW is either mispricing or trying to expand the engagement scope at the founder’s expense.
What does “scope creep” mean on the founder side?
Scope creep is the founder adding new requirements mid-engagement — typically at the W4 Loom review when the prototype is real enough to spark new ideas. Every new requirement competes for the same fixed budget of engineering hours. The fix is the SOW’s change-order clause: new scope means new weeks, new fee, or a clean swap (one thing in, one thing out). Verbal “and also” requests never become part of the build.
How does the founder know if the partner is doing their job in the engineering column?
Three diagnostics, none of which require reading code. First, the partner can demo the system end-to-end on a Loom every Friday with the eval-set score visible. Second, the production dashboard exists by W9 and shows the eval-set pass rate against live traffic. Third, the handoff document at W12 names the architecture, the deploy process, and the on-call escalation in plain language a non-engineer can summarise.
Can the founder skip the W5 eval review?
No. W5 is the diagnostic week of the entire engagement — the eval baseline run reveals the three biggest quality gaps and the rewrite plan. A founder who skips W5 loses the chance to catch a misaligned product before the W6 hardening sprint locks the direction. If the founder must miss W5 for an unavoidable reason, the eval review reschedules — it does not delegate.
What if the founder has never worked with an AI development partner before?
That is the expected case for a first-time AI MVP build. The role split in this article is the calibration point. Read the SOW with the role split in hand and ask the partner to map each clause to either the founder column or the partner column. Any clause that does not map cleanly is the negotiation point worth resolving before W1.
How does the role split change for a 6-week engagement instead of 12?
The columns do not change; the time budgets compress. A 6-week shape demands the founder show up for W1 kickoff, W3 mid-eval review, and W6 launch. Quiet weeks compress to half-weeks. The customer-access column gets tighter — 3 customers instead of 5–10. Anything that has to drop usually drops from the eval-set rigor, which is why 6-week engagements typically ship prototypes rather than production systems.
What is the one role the founder cannot delegate?
Decisions. Every other column can flex — domain expertise can be supplemented, brand voice can be reviewed asynchronously, customer access can be batched. Decision-rights cannot. If the founder cannot decide between feature A and feature B within 48 hours of being asked, the engagement stalls. A founder who delegates decisions to a co-founder, an advisor, or a board has not actually delegated — they have inserted a serial dependency that breaks the 12-week shape.
For the founder role split as a one-page printable PDF plus monthly editorial on the idea-to-product engagement model, subscribe to the SFAI Labs newsletter. One email a month, written for non-technical founders about to sign their first AI development SOW.
Arthur Wandzel