Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 15 min read

AI MVP launch costs vs ongoing costs: what to expect month 1–12

AI MVP launch costs vs ongoing costs: what to expect month 1–12

An AI MVP in 2026 typically lands a $100–200K launch invoice across months 1–4, then settles into a $3–15K per month run-rate across months 5–12 — split roughly half infra (model API, hosting, vector store, observability) and half people (maintenance, evals, on-call). The two numbers behave nothing alike. Launch is a one-time capital event with a fixed scope and a fixed end date. Ongoing is a recurring operating line that scales with usage on two lines, stays flat on three, and crosses a cost cliff sometime between month 6 and month 10 if traffic ramps. This guide walks the first 12 months one month at a time, separates the cliff from the curve, and ends with a year-2 planning frame so the conversation with the CFO at month 11 is not a surprise.

This is a founder’s explainer that runs alongside the AI MVP economics playbook, which decomposes the build line in depth. This piece sits within the broader idea-to-product manifesto, which frames how non-engineers ship AI products in 2026.

Launch cost vs ongoing cost: the two-line shape

A 12-month AI MVP budget has two lines, and they do not behave alike.

Launch is the SOW line — what the agency or in-house build team invoices to take the product from PRD to a hardened MVP in production. It is one-time, scope-bounded, and lands across months 1–4 in our typical sequencing. The 2026 envelope we see most often is $100–200K, with the lower end reserved for a tightly scoped single-workflow build and the upper end covering multi-workflow products with non-trivial data ingestion and a hardening sprint.

Ongoing is the run-rate line — what it costs to keep the product alive, evolving, and getting better at the job it was built to do. This is months 5–12 and beyond. The 2026 envelope is $3–15K per month, split roughly half infra and half people at MVP-1 scale. The split is what most budget templates get wrong — they quote a single retainer number that hides which lever is which.

The shape matters because the two lines invert at different points. A CFO who sees only the launch number assumes ongoing is a small footnote. A CFO who sees only the ongoing number underestimates the upfront capital event. Year one needs both numbers visible, on the same chart, with the cliff drawn in. We saw the same pattern documented in decoding AI project TCO: 7 cost lines most CFOs miss — the lines most often missed are not the obvious ones; they are the ones that change shape between launch and ongoing.

The launch invoice, decomposed

The $100–200K launch envelope decomposes across four phases. Each phase is a discrete deliverable a CFO can tie a payment milestone to.

PhaseWeeksTypical 2026 costWhat lands
1. Planning & PRD1–2$8K–$20KProduct requirements doc, eval framework, success metrics, scope-cut decisions
2. Architecture & design1–2$12K–$30KSystem architecture, model selection, data model, infra plan, security review
3. MVP build4–8$60K–$120KOne or two workflows live, retrieval pipeline, model integration, internal eval pass
4. Hardening1–2$20K–$30KExternal eval suite, observability wired, on-call runbook, security pass, launch

The decomposition matters because the four phases have different risk profiles. Phase 1 and 2 are the cheapest and the most expensive to skip — every $1 skipped here typically costs $5–$10 in rework during phase 3. Phase 3 is the biggest line item and the one where most cost overruns originate, mostly because scope creeps after phase 1 was rushed. Phase 4 is the one founders cut first and regret second; the eval suite and observability that get built here are what make month 5 onwards cheap.

The full decomposition with engineer-day math is covered in how much does an AI MVP cost in 2026. The anti-pattern catalogue of overruns is in anatomy of a runaway AI project: 5 cost-side root causes.

The ongoing bill, decomposed

The $3–15K monthly ongoing envelope splits along two axes — infra vs people, and flat vs variable.

Infra side ($1.5K–$8K / month at MVP-1)

Five infrastructure lines run the product month over month. Three are flat, two scale with usage.

  • Model API spend (variable): $200–$2,000 / mo. Token-priced. Frontier-tier models (Claude Opus 4.8, GPT-5, Gemini 2.5 Pro) run roughly 5–10x the per-token cost of workhorse-tier models (Claude Sonnet 4.5, GPT-5 Mini, Gemini 2.5 Flash). The most common 3–5x saving in 2026 is routing evals to frontier tier and production traffic to workhorse tier.
  • Hosting (mostly flat at MVP-1): $150–$1,000 / mo. Compute, database, edge, secrets. Vercel + Supabase + Cloudflare is a defensible baseline. Tiers up once around 5K queries / day.
  • Vector store (mostly flat at MVP-1): $0–$500 / mo. Postgres with pgvector is the cheapest defensible default. Managed vendors (Pinecone, Weaviate, Turbopuffer, Qdrant) become defensible only once query latency or recall on the Postgres path becomes a bottleneck.
  • Observability (flat): $100–$500 / mo. Trace store with eval-replay retention. Langfuse self-hosted or a managed observability vendor. This is the line founders cut first; it is also the line that gates weekly iteration speed.
  • Security & secrets (flat): $50–$200 / mo. Secret manager, log retention, vulnerability scanning baseline.

The full decomposition of these five infra lines, with what scales and what does not, is the topic of AI infrastructure cost for an MVP: a non-technical founder’s guide.

People side ($1.5K–$7K / month at MVP-1)

Three roles consume the ongoing people line at MVP-1 scale, and the mix is what most retainers get wrong.

  • Maintenance engineering ($1K–$4K / mo): bug fixes, dependency upgrades, minor feature work, dev-prod parity. Usually a fractional engineer at 10–25% of capacity.
  • Eval & quality work ($300–$2,000 / mo): refreshing the eval suite as the product changes, running regression evals on model upgrades, investigating quality regressions. This is the line that grows fastest when the product is improving and the line that signals trouble when it shrinks.
  • On-call & incident response ($200–$1,000 / mo): retainer for off-hours coverage, incident review cadence, post-mortem write-ups. Founder-absorbed at the floor; fractional engineer at the ceiling.

The retainer-side decomposition lives in monthly AI development retainer costs. The hidden categories that creep into months 3–6 if not budgeted are catalogued in hidden costs of AI development.

Month-by-month: M1 through M12

The 12-month spend curve has three distinct shapes — capital-heavy months (M1–M4), stabilisation months (M5–M8), and optimisation months (M9–M12). The table below assumes a typical $150K launch and a $6K average run-rate; the shape generalises to the $100K and $200K endpoints by scaling the launch column.

MonthPhaseLaunch invoiceOngoing run-rateTotal monthly
M1Planning & PRD$15K$0$15K
M2Architecture & design start$25K$0$25K
M3Build (mid)$50K$0$50K
M4Build (late) & hardening$35K$1.5K (infra spin-up)$36.5K
M5Launch + first month live$25K (hardening tail)$4K$29K
M6First eval cycle$0$5K$5K
M7First user growth bump$0$6K$6K
M8Pre-cliff month$0$7K$7K
M9Cliff month (often)$0$9K$9K
M10Cliff response (model swap or self-host pivot)$5K (one-time engineering)$7K$12K
M11Stabilised post-cliff$0$6.5K$6.5K
M12Year-2 planning$0$6K$6K
Total~$150K~$52K~$202K

The story this table tells is the story most budget templates miss: year one is roughly 75% capital, 25% operating. Year two flips. The cost-curve mechanics behind the flip are the subject of the AI project cost curve: why year-2 spend should drop 40%.

A few things to call out month by month.

  • M1–M2 are deceptively quiet. No production cost yet, but the eval framework and architecture chosen here determine whether M9 is a cliff or a curve. Skip them and the cliff arrives early and steep.
  • M4–M5 are the transition months. Hardening tail overlaps with first-month infra spin-up. Budgets that treat these as separate months tend to double-count or, more often, leave out the overlap entirely.
  • M6 is the first month the eval cycle runs end-to-end. The bill jumps because eval calls hit frontier-tier models. This is the right place to spend; cutting it is the most expensive cut available in months 5–12.
  • M8–M9 is the cliff window. See the next section.
  • M10 is a working month — the response to the cliff. A model swap, a routing layer, or a self-host migration. One-time engineering cost, recovered inside 2–3 months at the new run-rate.
  • M12 is the planning month for year two. The conversation that should happen here, in writing, with the CFO.

The cost cliff: when token economics break

The cost cliff is the moment in months 6–10 when the model API line crosses 50% of the monthly ongoing bill. Once that line dominates, the economics of running the product change shape.

Before the cliff, the budget is dominated by flat lines (hosting, observability, retainer). Cutting cost means renegotiating contracts or absorbing work into the founder’s bandwidth — and the marginal dollar back is small.

After the cliff, the budget is dominated by a variable line that responds to engineering. Cutting cost means routing more traffic to workhorse-tier models, batching requests, caching, switching to a smaller fine-tuned model, or self-hosting one of the workhorse-tier models. A single engineering week here can move a five-figure annual line.

The break-even math is approximately this. A workhorse-tier API call in 2026 costs roughly $0.02–$0.10 per query. A self-hosted equivalent (a small open-weights model on a managed GPU endpoint) runs roughly $1,500–$3,000 / mo all-in for ~30K queries / day capacity. That makes self-hosting break-even somewhere between 50K and 150K queries / month — well above MVP-1 scale, but reachable inside year one if traffic ramps cleanly.

The cliff response usually plays out across two months. M10 is the engineering work (model swap, routing layer, eval re-validation). M11 is the stabilised state at the new economics. The framework for thinking about this break-even at the unit-economics level is in decoding cost per query: a defensible unit economics framework.

Two things that look like the cliff but are not:

  • A bad month from one runaway workflow. Usually a prompt that exploded in output length or a retry loop. Fixable in days, not a cliff.
  • A vendor price change. Frontier-tier prices have moved roughly -30% per year through 2025; a price drop is not a cliff event, it is a tailwind.

Year-2 planning: what changes after month 12

The conversation in month 12 is the one most retainers do not prepare buyers for. It has three load-bearing decisions.

1. Renegotiate the retainer. The retainer that made sense at M5 (one fractional engineer plus on-call) is rarely the right shape at M12. The eval suite is mature, the product is stable, and the work is more about feature velocity than maintenance. The retainer shape that fits is often half the maintenance line, double the feature line, same total.

2. Swap models where the cliff is unresolved. If M10 happened, this is just review. If M10 did not happen because traffic stayed flat, M12 is the moment to plan ahead — pick a workhorse-tier provider on a thin abstraction (OpenRouter or an internal router) before lock-in becomes expensive.

3. Internalise vs continue. The cheapest defensible answer to “should we hire a head of AI in year two” depends on whether the product is one-workflow or multi-workflow. One-workflow products almost always stay externally retained. Multi-workflow products with growing usage typically internalise between M14 and M18.

A defensible year-2 budget at MVP-1 scale, at the cliff-resolved state, lands roughly $45–80K all-in for the year — half year-one ongoing, less than a quarter of year-one total. This is the 40% drop the cost-curve thesis describes.

Frequently asked questions

What is the floor for a defensible AI MVP launch invoice in 2026?

Roughly $80–100K for a single-workflow MVP, tight scope, no enterprise compliance, founder absorbing PM. Below that, one of the four phases (planning, architecture, build, hardening) is being skipped. Hardening is the cheapest line to skip and the most expensive to skip — it sets months 5–12 run-rate.

What is the cheapest defensible ongoing bill after launch?

Roughly $3K per month at MVP-1 scale: fractional engineer at 10% capacity, founder-absorbed on-call, pgvector inside the application Postgres, workhorse-tier routing. Below that floor, the line being cut is usually evals or observability — the two cuts that punish year-2 cost the most.

When does the model API line dominate the monthly bill?

Typically between 5K and 15K queries per day at MVP-1 product complexity. At 1K queries per day, the API line is 10–20% of the bill. At 50K queries per day, it is 60–80%. The crossover is the cost cliff — months 8–10 in most launches that see linear user growth.

Can a non-technical founder negotiate a fixed monthly cap?

Yes, on the people side. Maintenance, eval, and on-call retainers are routinely fixed-fee at MVP-1. The infra side cannot be fixed without surrendering control over which models get called — a bad trade. The defensible structure is fixed retainer plus pass-through infra with a written monthly variance report.

What if usage stays flat for 12 months?

The bill stays roughly flat too. Flat-usage MVPs land at the floor end of every range — roughly $3–5K per month, with the cost cliff never triggered. Most common shape for internal-tool MVPs and B2B MVPs in design-partner mode.

What is the most under-budgeted line in months 1–12?

Evals. Typically priced at zero in the launch SOW (lumped into hardening) and then $300–$2K per month in the retainer — enough at M5, not enough at M9. The deeper argument is in stop paying AI agencies for documentation, pay them for evals.

Is a one-time hardening sprint defensible, or should it be ongoing?

One-time hardening is defensible below 5K queries / day. Past that, hardening becomes a continuous line item and folds into the eval and observability retainer — usually around M9, in step with the cost cliff.

How does the 12-month total compare to building the same product in-house?

Agency path: roughly $200–250K all-in for year one. In-house path: roughly $300–400K (two engineers plus a product manager, fully loaded), with longer time-to-launch. In-house becomes cost-competitive at year 2 — which is why most teams that start agency-built move toward internalisation between M14 and M18.

Do these numbers hold if my product runs an agentic workflow?

No. Agentic workflows call the model 5–10 times per user query, moving the cost cliff forward by roughly a factor of 5 — to 1K–3K queries per day, not 5K–15K. Same month-by-month frame; the cliff just lands two to three months earlier.

Where does this guide stop being applicable?

At 50K queries per day, multi-tenant, or with compliance overlays (SOC 2, HIPAA, EU AI Act high-risk). At that scale the numbers in this article are floors, not ranges, and the AI MVP economics playbook hands off to the scaled-product playbook.

Where to go next

The month-by-month walk-through above is the operating layer. The next step depends on which line is the bottleneck in the conversation with the CFO.

If this kind of operational decomposition is what you want monthly — defensible cost frames, 2026 benchmark numbers, and the math that runs underneath — subscribe to the SFAI Labs newsletter. One issue a week, founder-readable, no hype.

Last Updated: Jul 18, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles