Most AI vendor reference calls are unstructured chats with vendor-selected advocates and produce a uniformly positive narrative uncorrelated with actual production behavior. The structured alternative is a three-call protocol with eleven named questions across three roles; technical lead, procurement or vendor-management lead, and CFO or finance lead; plus one buyer-selected reference outside the vendor’s curated list. Each question is named with what a strong answer looks like and what a weak answer reveals. The output is a 22-point scorecard comparing vendor candidates on production behavior rather than demo polish. Run it on the shortlist of three from the RFI before any contract negotiation. Reference calls without an RFI produce shopping; an RFI without reference calls produces theory.
It builds on the AI build-vs-buy-vs-hire decision matrix for 2026. The matrix’s seventh principle — that most sourcing decisions get re-litigated on a quarterly cadence — depends on procurement evidence that is falsifiable. The reference call is the second falsifiable check (the first is the RFI eval-suite walkthrough).
Why most reference calls fail
A standard reference call is 30 minutes, vendor-arranged, with one happy customer who delivers a high-level positive narrative. The buyer takes notes, signs the contract, and three months in production behavior diverges from the impression.
What they missed is structure. The reference was selected by the vendor; the questions were unstructured; the reference’s role rarely matched what the buyer needed; the buyer rarely spoke to a customer the vendor didn’t pick. Each is fixable in a 90-minute reference protocol.
The structured alternative is sister-piece to the procurement RFI; see the AI vendor RFI template that surfaces real differentiation. The RFI compresses 8-12 vendors to a shortlist of 3; the reference protocol verifies the shortlist’s claimed answers in production.
The three reference roles
Three roles, three different signals.
- Technical lead. The engineer or engineering manager who runs the AI capability day-to-day. Reveals reliability, eval-set integration, incident time-to-resolution. Most important on engineering substance.
- Procurement or vendor-management lead. Negotiated the contract; manages the relationship. Reveals contract behavior, kill-clause acceptance, scope-creep, renewal dynamics.
- CFO or finance lead. Tracks total spend against plan. Reveals cost surprises, pricing friction, overage charges, contract-vs-actual variance.
Each role sees a different vendor. The vendor with excellent technical work can have terrible contract dynamics; the vendor with clean contracts can have an unstable engineering team. Three roles surface the multidimensional truth.
A fourth call; the off-list reference, picked by the buyer not the vendor; is the highest-signal call of the four.
The 11 questions
The 11 questions are distributed across the three calls; five for the technical lead, three for procurement, three for the CFO. The off-list reference gets a short subset (typically four questions, mixed across roles, depending on who they are).
Technical lead:
- How does the vendor handle eval regressions in production?
- Describe the most recent production incident and the vendor’s response.
- How often does the vendor deploy changes that affect your workload, and how is the change communicated?
- Who is your named technical lead at the vendor and how stable has the staffing been?
- Would you say the vendor’s product matches their pre-sales demo six months in?
Procurement:
- Did the vendor accept your kill clause? How did the negotiation go?
- Has scope crept since contract signature, and how was it handled?
- What does the vendor’s renewal posture look like; collaborative or extractive?
CFO:
- What did the vendor cost over the first 12 to 18 months versus the original contract value?
- Were there overage charges, integration cost surprises, or hidden line items?
- What is the renewal price increase the vendor has indicated?
Each question has a “strong answer” pattern and a “weak answer” pattern, named in the relevant section below. The buyer scores each answer 0-2.
Calling the technical lead
The most important call. Allocate 45 minutes; usually runs 60.
Question 1; eval regressions. Strong: specific recent regression named, time-to-detection in hours, engineer who fixed it, regression-prevention step shipped. Weak: process description without specifics. Score 2/1/0.
Question 2; most recent incident. Strong: detailed timeline, vendor’s engineer named, communication frequency, post-mortem shared. Weak: “they handled it well” without specifics, or “no incidents” (means eval is too thin to detect them).
Question 3; change cadence. Strong: specific cadence, channel where changes are pre-announced, eval-affecting changes flagged. Weak: “they push changes whenever” or “we find out from release notes after.”
Question 4; staffing stability. Strong: named technical lead, on the account 6+ months, accessible by Slack or email. Weak: “we work with the support team” or “depends on the issue”; both signal no named senior coverage.
Question 5; demo vs reality. Strong: “yes, with these specific caveats.” Weak: an enthusiastic “yes” without nuance (signals curated reference).
Pair this call with a field guide to evaluating an AI agency in under 90 minutes; same engineering signals apply.
Calling procurement
Allocate 30 to 45 minutes.
Question 6; kill clause acceptance. Strong answer: structurally-intact kill clause, reasonable threshold values, vendor accepted in good faith. Weak answer: “we have a standard SLA” (which usually means no real exit) or “they accepted but we rarely invoked it” without describing the structure. The signal: vendors who structurally accept kill clauses behave differently in renewal than vendors who don’t.
Question 7; scope creep. Strong answer: specific instance of scope drift named, change-order process described, total cost of changes. Weak answer: “no scope creep” (rare; usually means they aren’t tracking) or “minor changes here and there” without quantification.
Question 8; renewal posture. Strong answer: specific renewal price percentage change (named), specific term changes (named), collaborative vs extractive characterization with evidence. Weak answer: “we’re still negotiating” if past renewal date, or “they were fine” without specifics.
Procurement references reveal the contract truth that the technical reference cannot; and vice versa. Both perspectives are necessary. See the AI vendor of record decision: who carries SOC 2 risk for the adjacent contract questions that the procurement call can also surface.
Calling the CFO
Allocate 30 minutes; CFOs are typically efficient with reference calls.
Question 9; total spend vs original contract. Strong answer: a specific dollar number for actual 12- or 18-month spend, compared to original contract value, with the variance explained. Within 10 percent of plan is healthy. Weak answer: “around what we expected” or “I’d have to check.”
Question 10; surprise charges. Strong answer: itemized list of overages, integration costs, premium-tier upcharges, with each line traced to a contract clause. Weak answer: “no surprises” (rare; usually means the CFO isn’t auditing) or vague references to “extra costs.”
Question 11; renewal price posture. Strong answer: specific percentage increase the vendor has communicated, with comparison to the buyer’s expectation. Weak answer: “they haven’t told us yet” if within 90 days of renewal; signals vendor is gaming timing.
CFO references are uniquely important because they alone see the gap between the contract value the procurement team negotiated and the actual run-rate spend that lands on the P&L. The TCO checklist; see the AI capability TCO checklist; is the structural artifact that this call validates.
The off-list reference
The fourth call. The buyer picks a reference the vendor did not propose.
Process:
- Find the vendor’s customer list. LinkedIn customer connections, customer logos on the vendor’s site, case studies, conference talks. The vendor’s marketing usually surfaces 20 to 50 customers.
- Pick a customer of similar shape; same industry, same scale, same use case; that the vendor did not propose as a reference. Avoid the vendor’s biggest logo (over-coached) and the smallest (not representative).
- Ask the vendor for a warm intro, framed as part of the procurement process. A vendor who agrees and delivers within a week is signaling healthy customer relationships across the portfolio. A vendor who refuses or stalls is signaling that some customer relationships do not stand up to off-list inquiry.
The off-list call is the highest-signal call of the four because it bypasses the curated-advocate effect entirely. A vendor who has only happy customers will accommodate any reference request. A vendor whose customer base is split will arrange the curated three quickly and stall on the fourth.
Scoring the calls
Each of the 11 questions is scored 0-2. Plus the off-list call’s four questions, also 0-2. Total max is 22 points across the three primary calls; the off-list call adds an 8-point modifier (positive or negative).
- Above 16 on the primary calls with a positive off-list modifier: advance to contract negotiation.
- 12 to 16 with a neutral off-list modifier: advance conditionally with named follow-up calls or pilot work.
- Below 12 or any vendor where the off-list reference declines or stalls: decline.
Re-score quarterly for any vendor in the buyer’s portfolio. Reference dynamics shift as customer relationships age; a year-two CFO answer is different from a year-one CFO answer because year two is when overage and renewal pricing surface. The matrix’s quarterly re-litigation cadence applies here too.
Red flags and yellow flags
Red flags; any one declines the vendor: reference cannot name the vendor’s technical lead; CFO declines to discuss pricing vs original contract; off-list reference cannot be arranged; “I would not pick them again” in any framing.
Yellow flags; surface for follow-up: vague answers without specific incidents; procurement describes weak kill clause as “standard”; renewal price increase above 15% without justification; “they’re working on it” twice in one call.
Compare findings against the AI agency reference call: 11 questions that surface real client outcomes; same patterns recur.
Frequently asked questions
Why do most AI vendor reference calls produce no signal?
Most reference calls are unstructured 30-minute chats with a vendor-selected reference, who delivers a high-level positive narrative because the reference is a chosen advocate. The buyer leaves with an impression, not data. The fix is structured: 11 named questions, three different reference roles (technical lead, procurement, CFO), one buyer-selected reference outside the vendor’s curated list.
Who should the buyer ask to talk to?
Three roles, three calls. The technical lead who runs the AI capability day-to-day reveals reliability, behavior under load, and the eval-set integration story. The procurement or vendor-management lead reveals contract negotiation, kill-clause behavior, and renewal dynamics. The CFO or finance lead reveals total cost surprises, pricing-shape friction, and budget overruns. One buyer-selected reference outside the vendor’s curated list is the highest-signal call of the three.
What is the most important question to ask the technical lead?
How does the vendor handle eval regressions in production? The answer separates vendors with real eval discipline from vendors who rely on reactive fixes. Strong references describe a specific recent regression, the time-to-detection, the fix path, and the regression-prevention step shipped. Weak references give a process answer with no specifics.
What is the most important question to ask procurement?
Did the vendor accept your kill clause and how did the negotiation go? The answer reveals whether the vendor’s RFI commitments survived the contract phase or got negotiated away. Strong references describe a structurally-intact kill clause and reasonable threshold values. Weak references describe the kill clause being defanged or removed in negotiation.
What is the most important question to ask the CFO?
What did the vendor cost over the first 12 to 18 months versus the original contract value? The answer reveals scope creep, overage charges, integration cost surprises, and renewal increases. Strong references describe predictable spend within 10 percent of plan. Weak references describe overruns, surprise charges, or steep renewal increases that the buyer didn’t see coming.
How does the buyer get a reference outside the vendor’s curated list?
Search the vendor’s customer list (LinkedIn, customer logos on the vendor’s site, case studies). Pick a customer of similar shape; similar industry, similar scale, similar use case; that the vendor did not propose. Ask the vendor for a warm intro to that customer, framed as part of the procurement process. A vendor who refuses or stalls is signaling that not many customer relationships are healthy.
How long should each reference call be?
Thirty to forty-five minutes per role. Time is mostly the buyer asking, the reference answering specifically, and the buyer probing one or two answers further. Calls that run long because the reference is volunteering detail are higher-signal than calls where the buyer is filling silence with new questions.
What signals are red flags in a reference call?
Vague answers without specific incidents. Reference cannot name the vendor’s technical lead by name. Reference describes scope creep but says it “worked out fine.” CFO declines to discuss pricing comparison to the original contract. Procurement describes a weak kill clause as “standard.” Reference says they would not pick the vendor again “but it’s working out.”
How does the reference-call protocol integrate with the RFI?
The RFI compresses 8 to 12 candidate vendors to a shortlist of 3. Reference calls verify the RFI’s claimed answers in production. Run reference calls only on the shortlist; the RFI score determines which vendors are worth the buyer’s reference-call time. See the AI vendor RFI template for the upstream procurement step that produces the shortlist.
What is the right scoring of reference calls?
Score each role’s answers on a 0-2 scale per question, summed across the three calls. Total of 11 questions, max score 22. Above 16 advances to contract negotiation. 12 to 16 advances conditionally with named follow-up. Below 12 declines. Re-score quarterly for any vendor in the buyer’s portfolio because reference dynamics shift as customer relationships age.
Key takeaways
Standard AI vendor reference calls are unstructured chats with vendor-selected advocates and produce no falsifiable signal. The structured alternative is three roles (technical lead, procurement, CFO) with 11 named questions, plus one buyer-selected off-list reference.
Each role sees a different vendor. The technical lead reveals engineering substance. Procurement reveals contract dynamics. The CFO reveals total cost truth.
The off-list reference is the highest-signal call. A vendor who can arrange it has healthy customer relationships across the portfolio. A vendor who cannot is signaling the curated three are the only references that hold up.
Score on a 22-point scale plus an 8-point off-list modifier. Above 16 with positive off-list advances to contract negotiation; below 12 or stalled off-list declines. Run after the RFI, before contract negotiation, on the shortlist of three. Re-score quarterly.
Arthur Wandzel