Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 13 min read

The Eval Budget rule: what % to spend on evaluation engineering

The Eval Budget rule: what % to spend on evaluation engineering

The Eval Budget Rule: spend 20–30% of a fixed-price AI MVP build on eval engineering. Below 15%, the product ships without a defensible quality bar. Above 35%, the team over-instruments a product that has not found market fit, burning runway on calibration nobody is asking for. The rule is a band, not a number, and the band has both a floor and a ceiling — most founders only hear about the floor. The 20–30% lands at $20K–$30K on a $100K MVP, $30K–$45K on a $150K MVP, and $50K–$75K on a $250K MVP. Pricing outside the band is fine — what is not fine is pricing outside without naming which side and why.

This article builds on the AI MVP economics playbook and the broader idea-to-product manifesto. It is the technical companion to the AI feature sizing framework and the dollar-bracket twin how much does AI eval engineering cost on a fixed-price MVP. The total-program lens sits in the hidden cost of AI evals, which lands at 30–40% — this article is narrower, scoped to the MVP build line only.

What eval engineering actually covers

Before naming a percent, name what the percent is paying for. Eval engineering on a 2026 fixed-price AI MVP is the labor across five sub-lines:

Sub-line What it is Typical share of the eval line
Test-set authoring 100–200 labeled cases sampled from the actual workload distribution ~25%
Rubric design Written grading criteria; LLM-as-judge prompt or human rubric ~13%
Harness setup Eval runner in CI; one-shot and scheduled execution ~22%
Regression suite Pass/fail gates; blocks merges on quality drop ~25%
Observability hookup Eval results joined to production traces; per-call grading ~15%

The five sub-lines are senior AI engineer labor, not commoditized QA. The Stack Overflow Developer Survey 2025 places senior AI engineers at $180–$220 per fully-loaded hour (Stack Overflow Survey 2025). Tooling is roughly 5–10% — Anthropic Inspect, OpenAI Evals, and Promptfoo are free; Langfuse, Braintrust, and Arize Phoenix run $100–$1,000/month at MVP scale. The other 90–95% is people.

What this percent does not cover: post-MVP eval ops, monthly observability SaaS fees beyond the build, calibration drift after launch, and model-swap re-runs against new frontier releases. Those land on a different P&L line — the retainer — and they are why the program-scale eval cost is higher than the build-scale percent here. Conflating the two is the most common pricing error in 2026 AI proposals.

Why a percent rule beats a dollar rule

A dollar rule fails on AI MVP budgets because build sizes vary too much — $25K of eval is rigorous on a $100K MVP and a rounding error on a $400K MVP. A percent normalizes across tiers and makes proposals auditable on one ratio. Two reasons the percent holds.

Total build budget already reflects feature complexity. A larger AI MVP is larger because the workload is harder, the surface area broader, or the integrations heavier — all of which raise the eval surface in proportion. Anthropic’s Building Effective Agents treats eval scope as a function of agent surface area, not as a fixed overhead. Holding the band at 20–30% scales the discipline with what it has to grade.

The percent makes off-band proposals legible. A $150K SOW with a $12K eval line is 8% — below the floor — and the founder has a clean question to ask: which of the five sub-lines is silently missing. A $150K SOW with a $60K eval line is 40% — above the ceiling — and the question flips: which post-MVP work is being pulled forward and why. The dollar number alone provokes neither conversation.

The 20–30% band — and why it has a ceiling

Most eval content asserts a floor. The interesting half of the rule is the ceiling.

The floor at 15% is the line below which one of the five sub-lines gets dropped or compressed beyond defensibility. Below 15% the regression suite typically disappears (the most common silent compression), or the test set collapses to fewer than 50 cases, or the rubric becomes informal and the LLM-as-judge runs without calibration. McKinsey’s State of AI reports 78% organizational AI adoption but a small fraction capturing material value — the pilot-to-production gap that defines the data is, in practice, an eval discipline gap.

The ceiling at 35% is the line above which the team instruments a product that has not earned the instrumentation. Above 35%, three failures recur:

Reviewer training and calibration depth exceed what pre-PMF traffic justifies. A 5-rater rubric calibrated against a 500-case gold set is excellent eval engineering; it is $30K of work on a product nobody has shipped. Calibrating that deep before the workload distribution stabilizes calibrates against assumptions.

Multi-model judge ensembling lands in the build. Cross-checking single-judge bias by running two or three frontier judges in parallel is right for an audit-ready production system. It doubles or triples eval inference and labor; a pre-PMF MVP is not where that audit posture pays off.

Eval ops platform work creeps in. Self-serve dashboards, per-stakeholder rubric reports, eval result versioning across feature branches — platform-grade deliverables that belong in the post-MVP retainer, not the build.

The band runs 20–30% so both edges have headroom. 20% is the floor with margin: five sub-lines instrumented, regression suite gating CI, test set sampled from the workload. 30% is the ceiling with margin: calibrated rubric work and broader sampling, no platform-grade eval ops yet.

Founder signals the spend is too low

A founder reading a fixed-price SOW with an eval line below 15% will not see the compression on paper — the SOW will still say “eval suite included.” The compression shows up in the operating cadence of the build. Three signals to watch for in weeks 2–6.

Demos work; eval scores are not shown. Every Friday demo passes a few hand-picked cases. No per-merge eval score in standup notes, no regression dashboard. When the founder asks “what’s the pass rate” the answer is qualitative (“looking good”). The eval line was compressed to a smoke test.

A model swap is treated as a workshop, not a re-run. A new frontier release lands — Anthropic, OpenAI, and Google ship updates on quarterly-or-faster cadence per Artificial Analysis — and the team’s response is to “evaluate it manually next sprint” rather than re-run the harness. There is no harness to re-run.

Quality complaints cannot be reproduced. A customer flags a wrong output; the team cannot replay the trace, cannot tell whether the case was in the test set, cannot quote a regression date. Observability was dropped to save the percent.

Any one of the three is a signal the build is under-invested. Two together is a signal the MVP will ship without a quality bar a founder can defend.

Founder signals the spend is too high

The over-invested case is rarer but real, and the symptoms are different.

Eval polish gates shipping. The team is on week 7 of the harness, the feature was working in week 3, and merges keep getting blocked on calibration disagreements between human raters. The product is being graded more carefully than it is being used.

Eval scope creeps into SOW change-orders. Change-orders name dashboards, reviewer-training sessions, eval ops tooling, multi-judge ensembling. Each is defensible work; none is MVP work. The retainer is being pulled forward.

Engineering velocity inverts. Commits skew toward the eval directory rather than the feature directory in weeks 4–6. A defensible eval discipline shows up as ~25% of commits, not 60%.

The cure is not “spend less on evals” — that destroys defensibility. It is “defer the platform-grade work to the retainer.” Draw a line between MVP-scale evals (the five sub-lines at 20–30%) and post-MVP eval ops (calibration drift, multi-judge, platform tooling), and push the second category into the run-rate.

Worked dollar splits at $100K, $150K, $250K

The band lands at concrete dollars at each MVP tier. The split below is the canonical 2026 shape — 5 sub-lines summed to the midpoint of the band — at three common fixed-price tiers.

Sub-line $100K MVP (25% mid) $150K MVP (25% mid) $250K MVP (25% mid)
Test-set authoring $6K $9K $16K
Rubric design $3K $5K $8K
Harness setup $5K $8K $14K
Regression suite $6K $9K $16K
Observability hookup $5K $6K $9K
Eval line total $25K (25%) $37K (25%) $63K (25%)
Band floor (20%) $20K $30K $50K
Band ceiling (30%) $30K $45K $75K

Three reading rules for the table:

Shape is constant; dollars are not. Regression suite and test-set authoring stay the largest sub-lines at every tier. Observability hookup grows slowest — the work is bounded by integration count, not feature complexity.

The $250K tier earns a bigger workload sample, not a sixth sub-line. A common error is assuming the $250K MVP gets a “richer” category. It does not — it gets broader workload sampling, deeper calibration, and a stronger regression baseline across the same five sub-lines. Adding a sixth category here means pulling post-MVP work forward.

At $100K the floor matters most. The temptation to compress eval to $10K (10%) is highest because the absolute dollar number looks like a small win. $10K is half the floor; the SOW review should treat it as such.

How the rule changes after PMF

The 20–30% band is an MVP-build rule. After PMF — once the workload distribution is stable and a regression is a revenue event — the eval percent rises and the cost shifts off the build line.

Post-PMF, the percent typically lands at 30–40% of total ongoing AI engineering cost, the band reported in the hidden cost of AI evals. The increase reflects continuous calibration drift work, ongoing model-swap re-runs against new frontier releases, and platform eval ops investment that pays off at production scale. What changes is the P&L line: from the build budget (one-time, fixed-price) to the operating budget (monthly retainer).

Treat these as two numbers on two clocks. MVP build sits at 20–30% on a 6–12 week clock. Post-PMF operating sits at 30–40% on a monthly clock. Confusing the two — by pricing the build at 35% to “get evals right from day one” — burns runway that should have funded distribution. Confusing them the other way — by pricing the retainer at 20% to “keep ongoing costs down” — degrades the production quality bar within two model swaps.

The Eval Budget Rule is narrow on purpose. Used inside its scope, it makes proposals legible and SOWs auditable. Outside its scope, it is the wrong rule.

CTA

If you are scoping a fixed-price AI MVP and want a second read on whether the eval line sits in the 20–30% band — or want the AI MVP scoping worksheet that puts this rule alongside the other four cost lines — both are available from SFAI Labs at no cost.

FAQ

What is the Eval Budget Rule? Spend 20–30% of a fixed-price AI MVP build budget on eval engineering — test-set authoring, rubric design, harness setup, regression suite, and observability hookup. Below 15% the MVP ships without a defensible quality bar. Above 35% the team over-instruments a product that has not yet found market fit.

Why isn’t there a single number — why a band? Because workloads vary in how much grading they need. A deterministic-check feature lands at the floor of the band; a graded-generation feature lands at the ceiling. A band lets the SOW reflect the workload without compressing the discipline. A single number forces a misfit on one side or the other.

How is this different from the “35% of project budget” figure I have seen elsewhere? The 35% figure refers to total project cost over a longer window — including post-MVP operations, calibration drift, and model-swap re-runs. This article scopes the rule to the fixed-price MVP build only, where the percent sits structurally lower at 20–30%. The two numbers describe different P&L lines and should not be conflated.

Does the rule apply to non-LLM AI products — e.g., a classical ML classifier? The shape holds but the sub-line mix shifts. Classical ML evaluations need labeled data, train/test/holdout splits, and metric thresholds — different artifacts, similar labor share. The 20–30% band is a defensible starting point; the floor reason (“regression suite gates CI”) is identical.

What if my MVP is a single deterministic-check feature — do I still need 20%? Likely closer to the floor at 18–22%. A deterministic-check feature tolerates a smaller test set and simpler rubric. The regression suite and observability hookup still sit at the same share — those sub-lines do not shrink with workload complexity.

Can I defer evals to v2 and ship faster? No. Retrofitting an eval discipline onto a shipped AI product reliably costs 1.5x–2x the inline cost, because the workload distribution has to be reconstructed from production traffic and the rubric authored against outputs that already shipped. Deferred evals also forfeit model-swap optionality.

Where does eval-time LLM inference cost sit — inside the 20–30%? No, under infrastructure — but name it explicitly. The 20–30% is labor for the five sub-lines. Eval-time inference (re-running the harness daily or per-merge) is a recurring infra cost in the $200–$500/month range at MVP scale per Anthropic pricing and OpenAI pricing. Folding it into a combined line hides the most volatile number in the budget.

How do I tell if a vendor’s eval line is padded toward the ceiling? Ask which sub-line lands at the top of its sub-band, and why. A defensible 30% line names a workload reason (“graded-generation feature with five-rater calibration to land at audit-ready”). A padded 30% line names tooling or dashboards — a platform reason that belongs in the retainer.

What is the rule for an in-house team versus a fixed-price partner? For an in-house team, the equivalent is 25–35% of senior AI engineering capacity allocated to evals during the MVP build window. The slightly higher band reflects retro and ramp work that a fixed-price partner does not absorb. The floor-and-ceiling principle is identical.

What does the rule say about agent-based MVPs with multiple tools and a planner? Agent MVPs sit at the ceiling, typically 28–32%, because the eval surface multiplies — per-step, end-to-end, and tool-use grading are three surfaces, not one. Anthropic’s agent design guide treats this as a first-order property. 35% remains the over-investment line; the band’s center shifts up.

Last Updated: Jul 22, 2026

DJ

Dirk Jan van Veen, PhD

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles