The weekly demo is the founder’s primary reading instrument on an AI MVP engagement — and most founders read it backwards. They walk in hoping for reassurance and leave with vibes. The instrument is not designed for reassurance. It surfaces five specific things to question and five specific things to celebrate, and the meeting closes on one written commit list and one written de-scope decision. Thirty minutes, every week, same agenda. Run it correctly and the engagement converges. Run it as a status update and the engagement drifts until week six, when the founder discovers the eval suite was scoped to the wrong rubric.
This piece extends Founder Operating with an AI Partner, part of the idea-to-product manifesto. It zooms into the 30-minute Wednesday block of the weekly founder-partner cadence and prescribes the dual-list checklist a non-engineer founder uses to read it.
Thesis: the demo is a reading instrument
A weekly demo is not a status update. A status update is one-way reporting; the founder receives, the partner performs. A reading instrument is a calibrated meeting that produces a decision before the end of the call.
Most founder-side demos drift into status-update theatre because nobody hands the founder an agenda. The partner runs the meeting in the shape that flatters their week. Engineers walk through what shipped, the founder nods, somebody asks “any questions?”, everyone agrees to talk next week. By week four the founder has thirty pages of demo recordings, zero written decisions, and no read on whether the eval suite is converging or diverging.
The fix is a buyer-side agenda the founder owns: five questions, five celebrations, two written outputs, three anti-demos to refuse. The agenda does not require the founder to be technical — only that they read evaluation outputs, not source code.
The 30-minute demo at a glance
| Block | Length | Who leads | What it produces |
|---|---|---|---|
| Live demo of the highest-risk capability | 8 min | Partner engineer | Live behaviour on a fresh input, not a recorded clip |
| Eval suite results vs threshold | 7 min | Partner technical lead | One slide: pass rate by rubric dimension, this week vs last |
| Five questions (the question list below) | 8 min | Founder | Five recorded answers, one per question |
| Scope check + cost-per-task | 4 min | Founder + partner | Written commit list + any de-scope decision |
| Celebrations (the celebration list below) | 3 min | Founder | Named recognition of work that held |
Thirty minutes, no slip, every Wednesday. If the demo runs long the partner is presenting and the founder is not driving. The block lengths are deliberate: only eight minutes of live demo because the founder is not buying a feature show — the founder is buying calibrated signal.
Five things to question
The founder asks these every Wednesday, in this order. Eval-first, scope-second, cost-third — the order in which an AI MVP engagement quietly fails.
Question 1 — What is the eval pass rate this week against the threshold we set in week one?
The founder is not asking “did it work in the demo?” — the founder is asking for a number against a number. If the rubric was scoped at 85% across seven dimensions with a 50-prompt golden set, the partner walks in with a slide reading “this week 87 / 86 / 84 / 88 / 79 / 91 / 82” and last week’s row underneath. The founder reads which dimensions cleared the threshold and which regressed. A partner who cannot produce that slide on demand has not built the eval discipline this engagement requires.
Question 2 — What new failure modes did the eval suite catch this week, and what did we do about them?
A converging engagement adds failure modes faster than it removes them — then crosses over. A diverging engagement stops finding new failure modes around week three, not because the system is fixed but because no one is looking. The founder asks: “name three failure modes the suite surfaced this week that were not in the rubric two weeks ago.” If the partner cannot name three, the suite is being treated as a passing test, not a discovery instrument.
Question 3 — Did anything shipped this week change the scope without a written de-scope decision in the log?
Scope creep rarely arrives as a request. It arrives as a quiet decision the partner made because the original scope was harder than expected. The founder asks for the diff: “show me one thing in the rubric that was true on Monday and is no longer true on Wednesday, or confirm none.” Softening D4 may be the right call — softening D4 without telling the founder until week six is the failure mode. The question forces the diff to surface weekly.
Question 4 — What is our current cost-per-task and is it on the trajectory we forecast?
Unit economics is load-bearing, and most engagements defer it to a finance review that never quite happens. The founder asks for one number weekly: cost-per-task in cents at the median, with the 90th percentile next to it. If week-one forecast was $0.08 median and week-four reality is $0.31 median with a $1.10 90th percentile, the founder decides this week whether to switch the routing layer, change the model mix, or expand the budget.
Question 5 — What customer input from this week did we miss in the eval seed set?
The suite is only as good as the inputs it was built from, and inputs go stale within three weeks. The founder asks: “name two messages, tickets, or transcripts from customers this week that the seed set would not have generated.” The partner who names them and shows where they were added is running the discovery loop. The partner who answers “we’ll check Slack” has not closed the customer-to-eval path — results are converging toward yesterday’s customer.
Five things to celebrate
Celebration is not a soft skill. Naming the work that held is how the founder calibrates which engineering moves to reinforce. Specific buyer-side recognition is one of the cheapest engagement-quality interventions available — most founders skip it because nobody told them which five things to look for.
Celebration 1 — A specific regression catch, named
When the suite catches a regression before it ships to pilot, the work is paying for itself. The founder names it: “the suite caught the JSON-schema break on D3 between Tuesday’s and Wednesday’s commit — a four-engineer-day save.” Named catches reinforce the eval-first discipline; unnamed catches dissolve into general competence.
Celebration 2 — An eval threshold cleared on a dimension below it last week
When D5 moves from 79% to 86% and clears the 85% threshold, the founder calls it and asks “what changed?” — and writes the answer into the decision log. The engineering move that lifted D5 is now an artifact, not a memory.
Celebration 3 — A fallback path tested under failure
Production-ready quality is not the happy path. It is what happens when the model returns garbage, the API rate-limits, or the parser breaks. The founder celebrates the week the partner deliberately broke the primary path and verified the fallback handled it gracefully. Tested fallbacks are the difference between a demo and a product.
Celebration 4 — Observability dashboards in the green
The founder celebrates the week the partner showed a live dashboard — latency percentiles, error rates, token cost per request, drift signals — and it held green. No dashboard is a question, not a celebration. A green dashboard is the third leg of the production-ready stool after evals and fallbacks.
Celebration 5 — Customer feedback applied within the week it arrived
When a customer message on Monday becomes a seed-set prompt on Tuesday, an eval re-run on Wednesday, and a behaviour change shipped on Thursday, the engagement is running discovery-to-deployment in days, not months. The founder names it: “the Acme Corp ticket from Monday is prompt 47, ran red on D2 Tuesday, fix shipped this morning.” Naming the loop makes it repeatable.
The three anti-demos a founder should refuse by name
Three demo shapes look like progress and are not.
Anti-demo 1 — The prerecorded demo
A partner who walks in with a polished screen recording is showing the happy path on chosen inputs. The founder asks for one live run on a fresh input supplied on the call. A partner who cannot run live is hiding either an eval gap, a latency problem, or a brittle prompt. Not adversarial — the buyer’s correct read.
Anti-demo 2 — The engineer-only demo
A demo where only the partner’s engineering lead speaks and the founder’s product or customer-success colleague is absent is missing the buyer-side feedback loop. The founder insists the team member closest to the customer is in the room. A partner who treats the demo as engineering-to-engineering has misread which side is the buyer.
Anti-demo 3 — The moving-target demo
A demo where the rubric, threshold, or scope quietly shifted between this week and last is the most dangerous of the three because it usually arrives wrapped in good news (“we hit 92% on the new rubric!”). The founder asks: “what changed in the rubric between last Wednesday and this Wednesday?” Anything that changed without a written decision in the log is the anti-demo. Refuse by writing the de-scope or scope-expansion decision into the log before the demo ends.
The two written outputs every demo produces
A weekly demo that ends without writing two things down is not the meeting this article describes. Both outputs are written, dated, one-paragraph each — they become institutional memory and the basis for next week’s agenda.
Output 1 — The commit list for next week. Three to five items the partner commits to ship by next Wednesday. Each item names the rubric dimension or failure mode it targets. Vague items (“improve quality”) do not belong. Specific items do: “lift D7 from 79% to ≥85% by adding fifteen Spanish-language prompts to the seed set and changing the routing layer to prefer Claude Sonnet 4.6 on Spanish detection.”
Output 2 — The de-scope decision (or explicit confirmation of none). One sentence either confirming “no rubric changes this week” or recording a specific de-scope (“D7 paused for v1; ships in week 8 phase 2”). The discipline of writing “none” when there are no changes makes the entry that records a real change credible six weeks later.
What good looks like by week 4
By the end of week four, the founder reads the four demo recordings end-to-end and sees one thing: the agenda held. The same five questions were asked, the same five celebration categories tracked, and the decision log has roughly four entries. The cost-per-task line moved deliberately, not surprisingly. The customer-feedback-to-eval loop closed at least twice.
The founder who runs this agenda discovers what the absentee founder and the over-meddler never do: the weekly demo is the cheapest, highest-signal hour the engagement produces — the only meeting where evaluation, scope, and unit economics converge in a single calibrated read.
That is what to question. That is what to celebrate. That is the weekly AI MVP demo, read correctly.
Frequently asked questions
What is the weekly AI MVP demo in plain English?
A 30-minute Wednesday meeting where the partner shows live behaviour on a fresh input, walks through this week’s eval results against the week-one threshold, and the founder runs a five-question buyer-side agenda that closes on a written commit list for next week and a written de-scope decision (or confirmation of none).
Who runs the demo — the founder or the partner?
The partner runs the live walk-through (the first eight minutes). The founder runs the agenda — five questions, scope check, cost-per-task, celebrations. A demo where the partner runs the entire meeting is a status update, not this agenda.
What if my partner cannot produce eval pass-rate numbers against a threshold?
Structural finding, not a slow week. By week two an engagement should have a named rubric, a written threshold, and a seed set the partner can re-run and report against in a slide. A partner in week four without the slide is operating without the discipline this article assumes — the conversation moves from “let’s improve the demo” to “let’s recover the eval baseline.”
Should the demo be recorded?
Yes — a Loom of the live walk-through and the eval slide preserves the artifact. The five-question discussion does not need to be recorded; the written commit list and de-scope decision capture what matters.
How is this different from a sprint review or a customer-facing demo?
A sprint review is internal and uses feature-shipping vocabulary. A customer-facing demo is outward-facing and uses narrative storytelling. The weekly AI MVP demo is buyer-side and uses evaluation vocabulary (pass-rate vs threshold, regression, fallback, observability) — the difference between calibrating an engagement and reporting on one.
What if the partner refuses to take a fresh live input on the call?
Refuse the demo. Politely ask why and listen. Honest answers — “forty-five second cold start,” “staging shares a database with production” — are themselves findings worth recording. A partner who routinely cannot run live on fresh inputs is signalling brittleness or theatre.
How do I run question 4 (cost-per-task) without being technical?
Ask for one slide weekly: median cost-per-task in cents and the 90th percentile in cents, with last week’s row underneath. If the partner cannot produce it, observability is not in place — itself one of the five things this agenda surfaces.
How does this demo connect to the rest of the founder’s weekly rhythm?
The Wednesday demo sits inside the 90-minute weekly founder-partner cadence: Monday eval review, Wednesday demo + scope check, Friday customer review. It converts Monday’s eval read into next week’s scope and feeds Friday’s customer scan. For deeper context on the eval discipline behind question 1, see the role of evals in your weekly partner relationship; for the broader critique of partner-side demo theatre, see the AI agency demo problem and inside the SFAI Labs operating cadence.
If your weekly demo today is a recorded walk-through with an “any questions?” close, the gap is not effort — it is the agenda. Adopt the dual-list version next Wednesday: five questions in order, five celebrations by name, two written outputs before the meeting ends. Read the engagement, do not just attend it.
Arthur Wandzel