Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 21 min read

Fixed-price AI feature scoping: what's included

Fixed-price AI feature scoping: what's included

A fixed-price AI scoping engagement is only defensible when the proposal names the eight deliverables the fee buys, drops the three things vendors routinely call “scoping” that aren’t, and quotes a pricing bracket that lines up with the named deliverables. This article is the counter-signing inclusion list for non-technical founders, heads of product, and procurement leads reading a fixed-price scoping proposal in 2026: the eight required deliverables, what each one looks like on the page, what missing it costs in the downstream build, the three anti-patterns to strike out, and the three pricing brackets that map to the scope shape.

This is the procurement-edge companion to the eval-first build playbook, the guide to AI-feature scoping inside the broader idea-to-product manifesto. The playbook explains why an eval-first scoping method ships at quality. This article goes inside the scoping engagement itself and asks: when the partner sells a $40K fixed-price scoping sprint, what does that line item actually buy?

Why fixed-price AI scoping engagements fail

A fixed-price AI scoping engagement fails for one reason: the proposal sells the phase without naming the artifacts. The partner quotes a number, the founder signs, and four weeks later the engagement ends with a slide deck, a verbal “we think this is feasible”, and a build estimate that has no contractual relationship to the scoping work.

Three patterns recur in 2026 scoping proposals:

  1. Scope-by-phase, not by-artifact. The proposal names “discovery”, “feasibility”, and “PRD” as phases with weeks against each — but not the artifacts each phase outputs, who owns them, or what they look like. The founder pays for a phase narrative and receives no artifact the build phase can be priced against.
  2. The eval set is absent. BCG’s 2024 “Where’s the Value in AI?” study reported that 74% of corporate AI investments fail to scale beyond proof-of-concept. McKinsey’s 2024 State of AI work named scope discipline and eval acceptance as the two largest gaps between scaling and non-scaling programs. Most fixed-price scoping proposals exclude the eval set entirely, leaving the downstream build with no acceptance criterion.
  3. The go/no-go memo is a verbal recommendation. Most engagements end with a meeting and a follow-up email that softens the verdict, which means the founder cannot kill the project without losing political capital.

The fix is structural: bind the fixed price to eight named artifacts. Anything not in the eight is an explicit out-of-scope line. The cross-cluster companion piece — AI MVP fixed-price contracts: what’s in scope vs what’s not — does the same for the build phase. This article does it for the scoping phase that precedes the build.

The 8 deliverables a fixed-price scoping engagement must include

Every defensible fixed-price AI scoping engagement in 2026 produces eight artifacts. Below is the inclusion list, the cost of omitting each, and the SOW section it should feed into.

#DeliverableWhat it isCost of missing itMaps to build-SOW section
1PRD draftThe product requirements document in eval-first form — feature anatomy, persona, task, success criteriaSubjective build acceptance; partner sets the barScope section
2Eval set v130–50 graded test cases with rubric, frozen and version-controlledNo acceptance gate at M2; engagement converts to T&MAcceptance criteria section
3Task taxonomyThe named tasks the feature performs, with inputs, outputs, and dependenciesOut-of-scope tasks added during build; 30%+ overrunFeature decomposition
4Capability mapThe list of model and stack capabilities the tasks requireWrong-model lock-in or build-buy mistakesArchitecture section
5Cost modelThe unit-economics worksheet for inference, infra, ongoing evalInference reimbursables inside fixed fee, partner has no incentive to optimizeCost & reimbursables section
6Fallback designThe deterministic behaviour when the AI feature fails or refusesNo graceful degradation; first hallucination is a P1 incidentNon-functional requirements
7Risk registerThe named risks ranked by likelihood × cost, with mitigation ownersSurprise blockers in week 5 of build; engagement stallsRisks & assumptions
8Go/no-go memoThe written recommendation with named alternatives and decision criteriaFounder cannot kill the project without political loss(Triggers the build SOW or not)

A proposal that names fewer than eight is structurally underspecified. A partner who refuses to name all eight in writing is either inexperienced with AI scoping or pricing in the discretion to deliver less than was implied.

Deliverable 1: PRD draft (eval-first form)

What it looks like: a 6–10 page document naming one feature, one persona, one task, one primary frontier model, and one acceptance threshold expressed against the eval set. The format follows the eval-first pattern set out in from idea to PRD: what the first 14 days should produce.

What’s inside: the feature anatomy in six scope layers (feature shell, eval set reference, prompt library, model contract, observability stack, on-call window), the user task in plain English, the input distribution, the rubric, and the threshold below which the feature is not shipped.

Cost of missing it: the build phase has no specification to be priced against. The partner sizes the build off a Slack thread and the founder discovers in week 4 that what they thought was implied is being billed as a change order.

Deliverable 2: Eval set v1

What it looks like: a CSV or JSONL of 30–50 graded test cases — inputs the feature will see at runtime, the expected outputs or rubric dimensions, and the passing threshold. Version-controlled, owned by the founder post-delivery, and used as the build-phase acceptance gate. The eval test set every MVP needs before launch is the sibling reference.

What’s inside: an input-distribution sample, a graded rubric across 3–5 dimensions (faithfulness, completeness, format compliance, refusal correctness, latency), and a threshold (e.g., ”≥ 0.78 weighted average on 30 inputs, no rubric dimension below 0.65”).

Cost of missing it: build acceptance is subjective. The single most-omitted scoping artifact in 2026 proposals and the single most expensive omission downstream — without it, the M2 acceptance review is a debate, not a decision, and the engagement converts to time-and-materials the first time quality is contested.

Deliverable 3: Task taxonomy

What it looks like: a 1–2 page table naming every distinct task the AI feature performs at runtime, with inputs, outputs, dependencies, and the model invocation pattern (single-shot, multi-step, agentic).

What’s inside: a row per task. For an AI customer-support agent the rows might be classify-intent, retrieve-knowledge, draft-response, refuse-out-of-scope, escalate-to-human, log-resolution. Each row names triggers, data touched, returns, and whether the task has its own eval rubric.

Cost of missing it: the partner builds task 1, the founder discovers tasks 2–6 are out of scope. The “and more” clause covers a lot of expensive structure. We cover this anti-pattern in the case for boring AI features in MVP-1.

Deliverable 4: Capability map

What it looks like: a one-page matrix of the model and stack capabilities the named tasks require — function calling, structured output, retrieval, code execution, image input, long-context handling, fine-tuning, on-premises hosting.

What’s inside: rows are capabilities. Columns are candidate vendors or open-weights models (e.g., GPT-5, Claude Opus 4.8, Gemini 2.5 Pro, reasoning, Llama 3.3). Cells are pass/fail or score-ranked. The map produces a defensible primary and secondary model. AI model selection 101 is the sibling reference.

Cost of missing it: wrong-model lock-in. The partner picks a model based on familiarity rather than fit, and the founder discovers at M3 that the chosen model cannot produce structured output reliably. Switching is non-trivial because the prompt library, evals, and integration glue all bind to a model contract.

Deliverable 5: Cost model

What it looks like: a unit-economics worksheet with three layers — per-call inference cost, monthly recurring infra cost, and ongoing eval-engineering cost — projected against three usage tiers (low, expected, high).

What’s inside: input tokens × input price + output tokens × output price per call, multiplied by expected monthly calls, plus retrieval and reranker costs, plus observability, plus the founder’s expected ongoing eval cadence (typically 1–2 days of engineering effort per month for the first six months). The worksheet should answer “what is the marginal cost per resolved customer-support ticket” or its equivalent.

Cost of missing it: inference is bundled into the fixed fee and the partner has no incentive to optimize prompt length. The buyer discovers in production that the per-call cost is 3–5× what was implied. How much does AI eval engineering cost on a fixed-price MVP is the sibling reference.

Deliverable 6: Fallback design

What it looks like: a 1-page design document naming what the feature does when the model refuses, hallucinates, returns malformed output, exceeds latency budget, or hits a rate limit. Distinguishes graceful degradation from hard failure and names user-visible behaviour for each.

What’s inside: the refusal pattern, the malformed-output handler (retry policy, structured-output schema validation, fallback to deterministic logic), the latency-budget handler, and the human-in-the-loop trigger. The hallucination budget — the worst tolerable rate of unfaithful outputs — is named here and referenced in the eval rubric. The hallucination budget: how to write it into your spec is the sibling reference.

Cost of missing it: the first production hallucination becomes a P1 incident. The partner adds emergency guardrails as a change order. Fallback should be designed at scoping, not retrofitted at incident.

Deliverable 7: Risk register

What it looks like: a ranked list of 8–15 named risks with likelihood, cost-if-realized, and a mitigation owner. The format mirrors a board-level risk register; what changes is the row content.

What’s inside: model-deprecation risk, eval-distribution drift, latency-budget breach, cost-blowout, data-availability risk (the knowledge base is incomplete), regulatory risk (data residency constraints), and integration risk (rate limits or API constraints in the target system).

Cost of missing it: surprise blockers in week 4–5 of the build. The buyer learns at M2 review that the retrieval source was incomplete or the latency cannot be hit at the chosen context length. Each is a named scoping risk if the register was produced; each is a change order if it was not.

Deliverable 8: Go/no-go memo

What it looks like: a 2–3 page written recommendation, signed by the partner’s lead, naming (a) the recommended path, (b) the rejected alternatives and why, (c) the cost-benefit, (d) the conditions under which the recommendation flips to no-go, and (e) a build-phase estimate range with stated assumptions.

What’s inside: a one-sentence verdict, the eval-set baseline (how does an out-of-the-box GPT-5 or Claude Opus 4.8 score on the eval set today, without custom work — if the baseline is above threshold, the buyer probably doesn’t need a build), the named alternatives (use a vendor SaaS feature, defer, narrow the feature, change the persona), the cost-benefit table, and the flip conditions (e.g., “if the eval baseline exceeds 0.85 on a free-tier API, the recommendation flips to no-build”).

Cost of missing it: the founder cannot say no without political loss. The point of a defensible scoping engagement is to give the buyer permission to walk. A written memo with explicit flip conditions does that. A verbal recommendation does not.

Three things vendors call “scoping” that aren’t

In 2026 procurement reviews, three patterns recur where a vendor uses the word “scoping” to label work that does not produce the eight deliverables. Treat each as an anti-pattern and refuse it on counter-signing.

1. The free 60-minute discovery call as “scoping”

A pre-sales call where the vendor’s solutions team asks about the use case, then sends a 1-page proposal and a build-phase estimate. The buyer is told the scoping is “complete”.

What’s wrong: nothing in the eight-deliverables list has been produced. The build-phase estimate is anchored on solutions-team intuition, which is the most expensive estimate-anchor available.

How to refuse: ask the vendor to produce the eight artifacts as paid scoping deliverables before the build SOW is countersigned. A partner who claims the work is “included for free in the build” is pricing the scoping discretion into the build margin — the buyer pays for it without seeing the artifacts.

2. The “discovery workshop” without artifacts

A 2–5 day on-site or remote workshop with stakeholders, ending in a slide deck. The vendor calls this “scoping” and charges $15K–$30K for the week.

What’s wrong: a slide deck is a workshop narrative, not an artifact in the eight-deliverables sense. No eval set, no task taxonomy, no fallback design, no risk register — nothing that binds the build phase to a specification.

How to refuse: counter-sign only if the eight artifacts are explicit deliverables of the workshop. The slide deck is fine as an executive summary; it is not fine as the only output. If the partner cannot produce the eight artifacts in five days, that is useful signal.

3. The “ChatGPT POC” without an eval set

The vendor builds a working prototype using ChatGPT, Claude, or a similar end-user product in 1–2 weeks, then proposes a build phase to “productionize” the POC. The POC is offered as scoping evidence.

What’s wrong: a POC without an eval set is a demo. The vendor cherry-picks inputs that work, the buyer sees a confident demo, and the build phase is sized off the POC’s apparent success rather than its actual eval performance.

How to refuse: require the POC to be graded against a frozen eval set the buyer owns. The companion piece stop scoping AI features in user stories, scope them in evals covers the same anti-pattern from the vocabulary side.

2026 pricing brackets

Fixed-price AI scoping in 2026 falls into three brackets. The bracket is set by the shape of the scope (number of named features, data complexity, regulatory constraint), not by the partner’s seniority or geography. A proposal whose price does not match the bracket’s deliverable shape is mispriced — either padded or under-quoted.

BracketRangeScope shapeNamed deliverables
Small$15K–$25K1 named feature, 1 persona, 1 task, ≤2-week engagement8 deliverables, eval set v1 with 20–30 cases
Mid$25K–$60K1 feature with regulated or multi-source data, or 2 closely-related tasks8 deliverables, eval set v1 with 30–50 cases, data-availability assessment
Enterprise$60K–$120K+Multi-feature decomposition, vendor-neutrality requirement, regulated industry8 deliverables per feature, multi-vendor capability map, procurement-ready risk register

What moves a proposal up a bracket: regulated data (HIPAA, FedRAMP, GDPR-strict EU residency), multi-tenant architecture from M1, vendor-neutrality requirement (the buyer wants two model contracts as live alternatives), or a multi-feature scope decomposition.

What does not justify moving up a bracket: the partner’s hourly rate card, the number of stakeholder interviews, or “we usually charge more”. A bracket is anchored by deliverable shape, not by partner cost.

A proposal in the small bracket should produce all eight deliverables. The most common 2026 mispricing is a small-bracket scope quoted at mid-bracket prices because the partner has bundled the build-phase estimating work into the scoping fee. Strike that and ask for the eight deliverables priced separately from the build-phase estimate.

For a longer breakdown of pricing-model alignment with outcomes, see the companion piece AI project pricing models ranked by alignment with outcomes.

The scoping-to-build handoff clause

The eight deliverables are useful only if they bind the downstream build SOW. Every scoping engagement should produce a one-paragraph handoff clause that reads:

Each scoping deliverable is the load-bearing input to exactly one build-SOW section. The PRD draft feeds the build-SOW Scope section. The eval set v1 feeds the Acceptance Criteria section. The task taxonomy feeds the Feature Decomposition. The capability map feeds the Architecture section. The cost model feeds the Cost & Reimbursables section. The fallback design feeds the Non-Functional Requirements. The risk register feeds the Risks & Assumptions section. The go/no-go memo triggers the build SOW or terminates the engagement. Any build-SOW section without a feeding scoping deliverable is out of scope until a change order is executed.

This clause is what makes the scoping engagement enforceable downstream. Without it, the build phase is renegotiated from scratch and the eight deliverables become reference documents the partner can ignore.

The clause should sit at the bottom of the scoping SOW and be repeated as a forward-reference at the top of the build SOW. The cross-cluster companion piece what a defensible idea-to-product SOW looks like, with examples shows the structural form the downstream SOW takes once the handoff clause is in place.

Counter-signing checklist

Read the fixed-price AI scoping proposal in front of you with this checklist next to it. If any line is missing, ask the partner to add it before counter-signing.

  • The proposal names eight artifacts as paid deliverables, not three.
  • The eval set v1 is an explicit deliverable, with cardinality (30–50 cases) and ownership (the founder, post-delivery) named in the SOW.
  • The PRD draft is described as eval-first — it references the eval set and an acceptance threshold, not as a vision document.
  • The capability map is multi-vendor by default. The partner names alternatives, not just the recommended model.
  • The cost model is decomposed into inference, infra, and ongoing eval. The partner does not bundle inference into the fixed fee.
  • The fallback design is a named 1-page deliverable, not a paragraph inside the PRD.
  • The risk register has named mitigation owners — buyer or partner — for each row.
  • The go/no-go memo is a named written deliverable, with flip conditions stated.
  • The proposal includes the scoping-to-build handoff clause binding each artifact to one build-SOW section.
  • The pricing bracket matches the scope shape. Small ($15K–$25K), mid ($25K–$60K), or enterprise ($60K+).

If three or more checkboxes fail, the proposal is not scoping in the eight-deliverables sense. It is a workshop, a demo, or a free discovery call dressed in a scoping price. Walk or rewrite.

FAQ

What is the single most-omitted deliverable in 2026 fixed-price AI scoping proposals?

The eval set v1. Most proposals treat eval as an internal-QA task during build, not a paid scoping deliverable owned by the buyer. Without an eval set produced during scoping, the build phase has no acceptance gate and the engagement converts to time-and-materials the first time quality is contested. Insist on a numbered eval set as a scoping deliverable with cardinality named in the SOW.

How long should a fixed-price AI scoping engagement take?

For a single named feature, 2–3 weeks. For a feature with regulated data or multi-source dependencies, 3–4 weeks. For an enterprise multi-feature decomposition, 4–6 weeks. Anything longer is either a workshop being padded or a build phase masquerading as scoping. The eight deliverables can be produced in two weeks by a competent partner working full-time on a small-bracket scope.

Is a free pre-sales discovery call ever sufficient as “scoping”?

No. A free discovery call produces no artifact in the eight-deliverables sense. Use it as qualification on both sides — the partner decides whether the engagement fits, the buyer decides whether the partner is competent — but do not let it substitute for the paid scoping engagement. A partner who claims the discovery call replaces scoping is selling the discretion to deliver less.

Should the eval set be owned by the founder or the partner after delivery?

The founder. The eval set is the load-bearing artifact that binds future build phases, future partners, and future model changes. If the partner retains ownership, the founder loses the ability to switch vendors or run the eval against a new model without renegotiating. The SOW should say “the eval set is delivered as a JSONL or CSV file, owned by the founder, with no IP encumbrance”.

How do I price a scoping engagement if the partner quotes outside the brackets?

The bracket is set by scope shape, not partner rate. Ask the partner to map their quote to a bracket and justify the price against the named deliverables. A $50K quote for a single-feature small-bracket scope is mid-priced and should produce mid-bracket deliverables — likely an expanded eval set, a multi-vendor capability map, or regulated-data risk work. If the deliverables are not expanded, the quote is padded.

What goes in the go/no-go memo’s flip conditions?

The conditions under which the recommended path changes. Typical flip conditions include: eval-set baseline (if a free-tier API already scores above the threshold, the build flips to no-build), capability-map outcomes (if no available model passes the structured-output requirement, the build flips to vendor SaaS), cost-model outcomes (if inference cost exceeds the unit-economics threshold, the build flips to deferred), and risk-register outcomes (if model deprecation is named likely within the build window, the build flips to a different vendor). The memo names each condition explicitly so the founder has written permission to walk.

Can I run a scoping engagement with one partner and the build phase with a different partner?

Yes — and the scoping-to-build handoff clause is what makes this possible. If the eight deliverables are owned by the founder and bound to one build-SOW section each, a second partner can price the build against the existing scoping artifacts. Most vendor lock-in in 2026 AI engagements comes from scoping deliverables that are retained by the partner; insist on founder ownership at scoping signing.

Should the partner produce a baseline eval score during scoping?

Yes. The single highest-value scoping output, after the eval set itself, is a baseline score — how well does an out-of-the-box frontier model score against the eval set today, without custom work. If the baseline is above the launch threshold, the engagement should recommend no-build and the founder saves the build phase fee. The baseline score is the empirical input to the go/no-go memo and the most useful single number the scoping engagement produces.

What happens to the scoping fee if the recommendation is no-go?

The scoping fee is paid in full. The engagement’s purpose is to produce a defensible go/no-go decision, and a no-go outcome is as valuable as a go outcome — arguably more, because it saves the build-phase fee. A partner who reduces fees on no-go outcomes has a structural incentive to recommend go, which is precisely the bias the engagement is designed to remove.

How does the eight-deliverables list interact with a discovery-week method?

Some 2026 partners run a 5-day discovery-week format rather than a 2–4 week scoping engagement. The eight deliverables still apply; only the timeline compresses. The companion piece the AI agency discovery week: a 5-day method that replaces 4-week scoping covers the compressed format. The counter-signing checklist is identical; what changes is the partner’s intensity.

Where to go from here

If a fixed-price AI scoping proposal is in front of you this week, run the counter-signing checklist before signing. If three or more checkboxes fail, send the eight-deliverables list back to the partner as the revised scope. If you have not yet selected a partner, the list is also useful as an RFP attachment — vendors who produce all eight as named deliverables filter themselves cleanly from vendors who treat scoping as a workshop.

For a second opinion on a proposal already on the table, request a paid idea review. The review is itself fixed-price, follows the eight-deliverables template, and returns a written recommendation you can use to counter-sign or walk away.

Last Updated: Jul 15, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles