A founder forwards a $185K AI MVP quote to their CFO Sunday night with a one-line note: “wanted to flag before Monday’s board prep.” Monday morning the calendar invite comes back with a single agenda item — “AI MVP scope conversation, 25 minutes.” That meeting decides whether the project ships in 2026 or gets pushed to Q1 2027. The deciding factor is rarely the headline number. It is whether the founder can decompose the quote into lines the CFO can grade, defend each line in twelve seconds, and answer the three objections the CFO is already drafting. This piece names the seven lines a 2026 AI MVP budget runs on, sizes each, and shows what to say when the CFO asks — about each one — “can we cut this?”
It builds on the AI MVP economics playbook, which lays out the full economics of an AI MVP build, and sits within the broader idea-to-product manifesto. The playbook makes the case that the build economics shifted between 2018 and 2026. The seven-line worksheet below is the artifact a founder hands to a CFO to prove the shift on a specific quote.
Why the standard SaaS budget shape fails a 2026 AI MVP review
Most AI MVP proposals reach a CFO in a shape inherited from 2018 SaaS: design, development, QA, deployment, sometimes a generic “AI integration” bucket. That template was calibrated against a system whose behavior was deterministic, whose runtime substrate was stable for years at a time, and whose QA was a phase rather than an instrument. None of those conditions hold for a frontier-model build in 2026.
McKinsey’s State of AI has documented for two reporting cycles that roughly four out of five AI pilots stall before reaching production at scale. The failure mode is not the engineering itself. The failure mode is that the budget did not fund the lines a production AI system needs — the eval suite was unfunded, the hardening window was hand-waved, and infra was point-estimated when it should have been banded.
A 2026 CFO has seen this pattern enough times to test for it. The CFO does not want a 4-line breakdown. The CFO wants a 7-line decomposition that separates the work upstream of code (planning and PRD engineering), names eval engineering as a discrete category, and treats post-launch hardening and on-call as distinct from build. The 7 lines below are what survives that test.
The 7-line worksheet at a glance
| # | Line | 2026 range | % of $150K–$250K budget | What skipping costs | One-line CFO justification |
|---|---|---|---|---|---|
| 1 | Planning | $6K–$12K | 4–5% | Idea ships without ICP validation; 60-day post-launch pivot is 3× this cost | Validates the ICP and the workflow before any engineer writes code |
| 2 | PRD engineering | $10K–$18K | 6–8% | Scope creeps in week 6; change orders run 15–25% of build | Produces the eval contract the build is graded against |
| 3 | Eval engineering | $25K–$45K | 18–22% | No threshold to ship against; launch slips 4–8 weeks or ships unsafe | The instrument that proves the system meets the spec |
| 4 | Build | $55K–$95K | 38–42% | Cannot be skipped; under-staffing here surfaces as Lines 6 and 7 doubling | Implements the capabilities to clear the eval threshold |
| 5 | Infra | $4K–$10K build + $2K–$8K/mo post-launch | 3–5% build / $24K–$96K Y1 | Production stalls on month-2 throughput; emergency migration costs $15K–$35K | Hosting, inference cost, observability, secrets management |
| 6 | Hardening | $12K–$22K | 8–10% | Production incidents in week 1–3 post-launch; trust collapse, refund cycle | The two-week window that takes the build from “demo works” to “production safe” |
| 7 | On-call retainer | $18K–$30K (3 months) | 10–12% | First model-alias update or distribution shift causes a silent regression | Funded engineering capacity for the first 90 days of real users |
The percentages are calibrated to a $150K–$250K MVP. On a smaller $75K–$120K build the percentages shift — eval engineering and hardening hold their share while build and on-call compress. On a larger $300K–$500K build, eval and infra grow as a share while planning compresses.
Line 1: Planning
What it is. Five to ten calendar days of paired founder-engineer work before any PRD is written. Customer interviews to confirm the ICP. A workflow audit that names the manual sequence the AI MVP replaces. A constraint inventory — regulatory, data residency, latency, integration surfaces, budget ceiling. The output is a 2-page brief naming the workflow, the ICP, three to five candidate capabilities, and the constraints any build must respect.
2026 range. $6K–$12K. 25–45 hours at a blended $200–$300 per hour. Senior product or engineering time, not junior research.
% of budget. 4–5%.
What skipping costs. The most expensive class of AI MVP failure is shipping a working system to the wrong ICP. The mid-2024 wave of internal-tooling AI MVPs that pivoted to external customer use cases in 2025 cost their sponsors three to four times this line — because the eval suite, the integration surface, and the on-call shape all had to be rebuilt against a different distribution of users.
One-line CFO justification. “Five days of founder time and senior engineering time that validates the ICP and the workflow before we authorize a $150K build.”
Line 2: PRD engineering
What it is. A 6–12 page Product Requirements Document, written by a senior engineer paired with the founder, with an explicit eval-contract appendix. The eval contract names the test set the build will be graded against — typically 250–600 prompts or workflow inputs — the rubric for each capability, the pass threshold required to ship, and the failure modes the system must not produce.
This is the line that distinguishes a 2026 AI MVP from a 2018 SaaS MVP. A SaaS PRD names features. An AI PRD names capabilities and the evaluation contract those capabilities will be graded against. Without an eval contract, the build has no shipping definition and the team negotiates “done” in week 8.
2026 range. $10K–$18K. 40–70 hours of senior engineering paired with the founder, over 7–14 calendar days.
% of budget. 6–8%.
What skipping costs. Scope creep in week 6. The pattern is predictable: without a PRD that names the eval contract, the team converges on “we’ll know it when we see it” — and the founder pays for change orders, in cash, every two weeks. Industry benchmarks across agency engagements put scope-creep cost at 15–25% of original build budget. On a $150K build that is $22K–$37K of change orders that PRD engineering would have prevented.
One-line CFO justification. “The contract the build is graded against. Without it, we negotiate scope every two weeks for the next ten weeks.”
Line 3: Eval engineering
What it is. The instrument the system is graded against — and the only AI-specific budget line a CFO has not seen before. Eval engineering produces a graded test suite: a representative input set, a rubric for each capability, automated and human-in-the-loop scoring, and a CI integration that runs the suite on every prompt change, model change, or retrieval change. The suite is the production-readiness contract.
A funded eval suite on a 2026 AI MVP typically contains 250–600 graded inputs, a rubric covering 4–8 capabilities, a target pass threshold (commonly 80–90% with no critical-class failures), and an annotation methodology — typically domain-expert review at $150–$300 per hour for the first pass, with downstream LLM-as-judge or rule-based grading for ongoing CI.
2026 range. $25K–$45K. 80–140 hours of dedicated eval engineering, plus 20–40 hours of domain-expert annotation time.
% of budget. 18–22%. This is the line that breaks the SaaS-shape budget — a 2018 MVP allocated 5–10% to QA. A 2026 AI MVP allocates 18–22% to the eval discipline that replaces QA.
What skipping costs. Launch slips four to eight weeks while the team improvises a pass/fail definition under pressure, or the project ships without a threshold and produces a silent quality regression in week 3 that no one notices until a customer complains. The Stack Overflow Developer Survey and post-mortem write-ups across 2024–2025 converge on the same diagnosis: the AI projects that shipped on time and stayed shipped funded their eval suite at the same depth as their build.
One-line CFO justification. “The instrument that proves the system meets the spec. Without it, we have no shipping definition and no regression detection.”
Line 4: Build
What it is. The implementation — prompts, agents, retrieval, integrations, user interface, deployment. A typical 2026 AI MVP build runs 6–10 weeks of senior engineering against the eval contract produced in Line 2, instrumented by the eval suite produced in Line 3. Build velocity is measured against eval-suite pass rate, not feature checklists.
2026 range. $55K–$95K. 220–380 hours of senior engineering at a blended $250–$300 per hour, plus 40–80 hours of design and frontend work. The wide band reflects three drivers: number of integration surfaces (one channel vs three), architecture complexity (prompt-only vs RAG vs agent), and the depth of the user interface (admin console only vs full end-user UI).
% of budget. 38–42%. This is the single largest line and the line CFOs accept most readily — it maps cleanly to their existing mental model of “the engineering cost.”
What skipping costs. Build cannot be skipped. The failure mode here is under-staffing — running the build with one senior engineer and two juniors when the work needs two seniors and one junior. Under-staffed builds surface their unpaid debt downstream: Lines 6 (hardening) and 7 (on-call) typically double on under-staffed builds because the build shipped without absorbing its own complexity.
One-line CFO justification. “The engineering that implements the capabilities to clear the eval threshold.”
Line 5: Infra
What it is. Two sub-lines that finance often wants combined and engineering wants kept separate: build-window infra (development environments, staging, observability tooling, initial inference cost during eval runs) and post-launch recurring infra (production hosting, inference cost at user scale, observability storage, secrets management, error tracking).
Build-window infra is fixed because the 6–12 week window is short enough that variance is bounded. Post-launch recurring infra is a band, not a point estimate, because it scales with user volume and inference depth. Present it as a band tied to a named usage projection rather than as a single number.
2026 range. $4K–$10K build-window. $2K–$8K per month post-launch for a typical MVP-scale deployment (under 1,000 monthly active users with moderate-depth queries). Frontier-model inference costs at 2026 prices range from $1–$30 per million tokens depending on tier — published pricing pages from Anthropic, OpenAI, and Google. A back-of-envelope for a 500-MAU product running 20 model calls per user per month at $5 per million blended cost is approximately $1.5K monthly inference.
% of budget. 3–5% of the build budget for build-window. Post-launch infra is presented separately as a 12-month projection — typically $24K–$96K Year 1 — and is named as an operating expense, not capitalized.
What skipping costs. Under-budgeting build-window infra surfaces as a stalled staging environment in week 4, which costs engineering velocity. Under-budgeting post-launch infra surfaces as an emergency migration in month 2 — $15K–$35K of unplanned engineering time when a 200-MAU spike exposes a hosting choice calibrated for 50 MAU.
One-line CFO justification. “Hosting, inference, and observability. Build-window is fixed; post-launch is a band tied to a named user projection and reviewed quarterly against actuals.”
Line 6: Hardening
What it is. A discrete two-week window between “build complete” and “production launch” dedicated to taking the system from “demo works” to “production safe.” Hardening means: red-teaming the eval suite for adversarial inputs, instrumenting fallback behavior on every model call, validating data residency and access controls, load-testing against the projected traffic, and producing the runbook the on-call engineer will use in Line 7.
A SaaS MVP rolls hardening into QA. An AI MVP separates it because the failure modes are different — model refusals, distribution shifts, prompt-injection attempts, silent quality regressions, and inference-cost spikes are all production risks the build phase does not surface.
2026 range. $12K–$22K. 50–80 hours over two weeks. Senior engineering plus a focused security or red-team pass.
% of budget. 8–10%.
What skipping costs. Production incidents in weeks 1–3 post-launch. Trust collapse is the expensive failure here — a public AI product that hallucinates or refuses in the first week reaches social media before it reaches the on-call inbox, and the recovery cycle absorbs 4–8 weeks of engineering plus an unquantifiable hit to reputation. BCG’s 2024–2025 “AI at Work” surveys document that the single most-cited reason CFOs withhold further AI investment is a public failure mode in the first 90 days post-launch.
One-line CFO justification. “The two-week window that takes the system from demo-grade to production-grade. The line the audit committee asks about first.”
Line 7: On-call retainer
What it is. Funded engineering capacity for the first 90 days post-launch. Typically 8–15 hours per week of senior engineering, retained against a known monthly fee, with a written response-time SLA and an explicit list of incident classes covered (model regressions, distribution shifts, inference-cost spikes, integration breakage, customer-reported quality issues).
This is the line every founder is tempted to skip and every CFO has learned to insist on. The first 90 days of real user contact are when the eval suite gets its first real-world calibration. The model alias the build was tested against often updates within those 90 days — frontier model release cadence in 2026 averages a major model update per provider every 8–14 weeks, per Artificial Analysis tracking — and the production system needs an engineer who can re-run the eval suite, identify regressions, and ship the fix without a 4-week scope conversation.
2026 range. $18K–$30K for a 3-month retainer. Senior engineering at retained rates ($150–$250 per hour, 8–15 hours per week, for 12 weeks).
% of budget. 10–12%.
What skipping costs. The first model-alias update produces a silent regression no one catches until a customer complains. The integration breakage from a partner’s API change waits in a queue for two weeks. The distribution-shift incident that the hardening window predicted ships to production unmitigated. Each of these costs more than the retainer would have — and they all happen.
One-line CFO justification. “Funded engineering capacity for the first 90 days of real users. Three months is the minimum window for the first model update, the first traffic spike, and the first regression.”
The 3 CFO objections and how to respond
Across hundreds of board conversations on AI MVP budgets, three objections recur. Every founder defending a 2026 AI MVP budget should rehearse the answer.
Objection 1: “Why is QA-equivalent (eval engineering) 18–22% when our SaaS budget allocates 5–10% to QA?”
Response. Eval engineering replaces QA on an AI MVP — it is not an extension of it. SaaS QA verifies deterministic behavior against fixed test cases. AI eval engineering produces a graded suite that detects quality regressions against a probabilistic system whose behavior shifts under model updates. The discipline is closer to bench science than to traditional QA. Industry data from McKinsey’s State of AI and post-mortem write-ups across 2024–2025 converge on the same finding: AI projects that under-fund evals stall in production. Funded at 18–22%, eval engineering is the line that determines whether the other 78% ships.
Objection 2: “Can we defer hardening and on-call to post-launch as opex?”
Response. Hardening cannot defer because production-launch is the trigger for the failure modes it prevents. On-call can technically defer to a month-by-month decision, but every month of deferral compounds risk — the first model-alias update, the first traffic spike, the first regression are all in the first 90 days. The cleanest CFO framing is: hardening is capitalized as part of the launch project, on-call is operating expense for 12 months and reviewed at the end of Q1 post-launch. That treatment keeps the capex bucket tight and gives the CFO the quarterly review point.
Objection 3: “These ranges feel wide — give me a single number.”
Response. The single number is the midpoint of each band — for a $185K build, $9K planning, $14K PRD engineering, $35K eval engineering, $75K build, $7K build-infra, $17K hardening, $24K on-call retainer, plus $36K Year-1 recurring infra presented as opex. The bands are wide because three structural inputs drive the variance — number of integration surfaces, architecture complexity, and user interface depth. Naming those three inputs in the budget memo lets the CFO understand why a band exists and what would move the project to the low or high end.
Worked example: defending a $185K AI MVP
A founder receives a $185K AI MVP quote from a vendor for an 8-week build. The vendor’s proposal arrives in four lines: design ($25K), development ($110K), QA ($25K), and deployment ($25K). The founder rewrites the quote into the 7-line worksheet before the CFO meeting:
| Line | Vendor allocation | Worksheet allocation | Notes |
|---|---|---|---|
| 1. Planning | $0 (assumed inside design) | $9K | Carved out from design |
| 2. PRD engineering | $0 (assumed inside design + dev) | $14K | Eval contract appendix is the load-bearing artifact |
| 3. Eval engineering | $25K (labeled QA) | $36K | Up-funded; QA-shape allocation under-funds eval |
| 4. Build | $110K | $75K | Lower because PRD and hardening are separated out |
| 5. Infra (build) | $0 (assumed inside deployment) | $7K | Build-window only |
| 6. Hardening | $0 (assumed inside deployment) | $17K | The two-week production-readiness window |
| 7. On-call retainer | $25K (labeled deployment) | $27K | 3 months of retained senior engineering |
| Total build | $185K | $185K | Same total, defensible decomposition |
| Year-1 recurring infra (opex) | Not named | $36K | Presented separately as a 12-month band |
The total is unchanged. The decomposition is what survives the CFO conversation. Each line has a 2026 range, a percentage benchmark, a skip-cost, and a one-line justification. The CFO does not need a follow-up meeting.
From worksheet to 24-month TCO
The 7-line worksheet is build-only. The CFO will ask — sometime in the meeting, often near the end — “what is the 24-month total cost of ownership?” The answer is the worksheet plus recurring infra plus the post-launch TCO categories named in the decoding AI project TCO breakdown: eval test-set maintenance, model-upgrade re-evaluation, regression triage time, inference-cost variance, prompt registry maintenance, observability storage growth, and the renewed post-launch retainer beyond Month 3.
A defensible 24-month projection for a $185K build is approximately $260K–$340K — the build, plus 12–18 months of recurring infra, plus the post-launch TCO categories at typical run rates. Naming this number in the build memo, with the post-launch breakdown linked, is the move that closes the meeting in one round rather than two.
For founders who want the worksheet as a fillable artifact, the AI MVP scoping worksheet is a one-page version with the 7 lines, the % benchmarks, and a worked example.
Frequently asked questions
Why seven lines and not five?
Five lines is the founder-defense decomposition — the minimum a founder can hold in their head when arguing the budget. The 5-line decomposition is right for board-conversation framing. Seven lines is the CFO-approval decomposition — it separates planning from PRD engineering (because the CFO tests for the latter as the eval contract), and it separates hardening from on-call (because those are budgeted differently — hardening is capex inside the build, on-call is opex over 90 days). The two are companions, not alternatives. Use the 5-line frame to argue the budget. Use the 7-line worksheet to win the CFO sign-off.
How does this differ from a generic SaaS MVP cost breakdown?
A SaaS MVP breakdown follows a 2018 template: design, development, QA, deployment, with QA at 5–10% of budget. A 2026 AI MVP runs eval engineering at 18–22% (replacing QA), adds explicit hardening and on-call retainer lines, and treats post-launch recurring infra as a separate operating-expense band. A comparable-scope SaaS MVP runs $80K–$160K in 2026; an AI MVP runs $115K–$255K. The 30–60% delta is the cost of the eval suite, the hardening window, and the on-call retainer.
What if my vendor’s proposal arrives in only 3 or 4 lines?
Ask the vendor to redecompose into the 7 lines before you sign. A 3-line proposal is not necessarily wrong — it may be a presentation convention. But if the vendor cannot produce a 7-line decomposition with a per-line skip-cost on request, the proposal is almost certainly under-funding eval engineering, hardening, or on-call. The under-funding surfaces as scope creep in weeks 6–10 and as production incidents in weeks 1–3 post-launch. Both cost more than asking the question upfront.
Can a $75K or $100K AI MVP be defensible against this worksheet?
Yes, with structural compressions. A $75K AI MVP is defensible when the architecture is prompt-only (no retrieval, no agents), the integration surface is one channel, the eval set is 150–250 inputs (not 400–600), and the on-call retainer is 60 days (not 90). The trade-off is a narrower capability surface and a thinner buffer for surprises. Below $60K, the math typically breaks at 2026 senior engineering rates of $200–$300 per hour blended.
How do I know if eval engineering is genuinely funded or just labeled “eval” on a line item?
Three checks. Ask for the eval test set size and annotation methodology — a defensible suite has a named input count (e.g., “400 inputs, domain-expert annotated at $200 per hour”) and a written rubric. Ask which engineer owns the eval harness — a build engineer who runs evals on the side rarely produces a defensible suite; the funded shape is a dedicated eval engineer for the first 6 weeks. Ask for the eval pass threshold the build must clear before launch — “85% pass on the test set with no critical-class failures” is a real threshold; “evals look good” is not.
How should I present post-launch infra to a CFO if it is a band, not a point estimate?
Present it as a band tied to a named user projection. “$2K–$8K per month depending on whether we hit 100, 300, or 500 monthly active users by month three, reviewed quarterly against actuals.” CFOs accept bands tied to external inputs. They reject bands that read as “we do not know.” The reviewing cadence — quarterly, against actuals, with a named forecast input — is what makes the band defensible.
Where do the 2026 dollar ranges come from?
Three public primary sources. The Stack Overflow Developer Survey for senior engineering compensation bands (US-based 7–10-year seniors at $150K–$220K base, mapping to $200–$300 blended hourly on agency engagements). Public pricing pages from Anthropic, OpenAI, and Google for inference cost ranges. Artificial Analysis for frontier model release cadence, which informs the on-call retainer length in Line 7.
How does this connect to the broader idea-to-product manifesto?
The 7-line worksheet is the CFO-facing operational form of the manifesto’s claim that AI products are built by funding evaluation as the spec, not features as the spec. Line 2 produces the eval contract. Line 3 builds the suite. Lines 4–7 execute against the contract. A founder who accepts the manifesto’s principle but cannot defend the 7 lines to a CFO has only half of what the manifesto promises.
What is the most common founder failure pattern when presenting an AI MVP budget to a CFO?
Two failures, one upstream of the other. Presenting a single-number quote (“$185K”) instead of a decomposition — the single number invites the question that ends the meeting. Presenting a 4-line SaaS-shape decomposition that fails the CFO’s test because it does not name eval engineering, hardening, or on-call as discrete lines. The fix is to arrive with the 7-line worksheet filled in, a one-line CFO justification per line, and rehearsed answers to the three objections above.
Key takeaways
- The CFO meeting is decided by decomposition, not by the headline number. A $185K quote in 7 defensible lines closes in one meeting. The same quote in 4 SaaS-shape lines takes three.
- Eval engineering at 18–22% is the load-bearing line. It replaces QA on a 2026 AI MVP and is the line the project ships against.
- Hardening and on-call are separate, not combined. Hardening is the two-week production-readiness window — capex inside the build. On-call is funded engineering capacity for the first 90 days — opex over 12 weeks.
- Infra is a band, not a point. Present build-window infra as fixed. Present post-launch infra as a band tied to a named user projection and reviewed quarterly.
- Rehearse the three CFO objections. Why is eval 18–22%, can hardening defer to post-launch, and give me a single number. The founders who win the meeting have the answer queued.
- The 7-line worksheet is build-only. Name the 24-month TCO separately, with the post-launch TCO decomposition linked, so the CFO does not have to ask.
Next step. Pull the AI MVP scoping worksheet into a board memo, fill the 7 lines against your vendor’s quote, draft the one-line CFO justification per line, and arrive at the meeting with the three objections answered. The Monday-morning calendar invite will be a single meeting, not three.
Arthur Wandzel