What an AI pilot actually costs — a line-item breakdown
Most AI pilot budgets are wrong because they only price the build. Here's the honest line-item breakdown — discovery, data prep, evaluation, build, inference and the run cost nobody quotes — and where pilots quietly overrun.
- AI Engineering
- Cost
- Engineering Practice
Almost every AI pilot budget we're shown has the same shape: a number for "development", a number for "API credits", and a go-live date. It is wrong in a predictable way. The build is rarely the expensive part, the API credits are rarely the expensive part, and the go-live date assumes a step — proving the thing works well enough to trust — that hasn't been budgeted at all.
This is a line-item breakdown of what a pilot actually consumes. It is written for the person signing the purchase order, not the person writing the code, and the point is to make the invisible lines visible before they arrive as an overrun.
What a "pilot" has to prove
A pilot is not a small version of the product. It is an experiment with a decision attached, and the decision is always the same: is this worth building properly?
That framing determines the budget. If the pilot's job is to answer a question, everything that doesn't help answer it is waste — polish, edge-case coverage, a production-grade UI. But everything that does help answer it is mandatory, including the parts that feel like overhead. The most common way a pilot fails is not that the model performs badly. It's that it finishes and nobody can say whether it performed well, because nothing was ever measured.
So the budget has to cover measurement, not just construction.
The six lines
1. Discovery and scoping
Typically 5–10% of the pilot. One to two weeks. The output is a written problem statement, a defined success threshold, a named data source, and an agreed list of what the pilot will not do.
The success threshold is the part teams skip and the part that matters most. "The assistant should answer customer questions accurately" is not a threshold. "On a held-out set of 200 real support tickets, the assistant produces an answer a senior agent would send unedited at least 70% of the time, and never invents a policy" is a threshold. The first cannot be passed or failed. The second ends the argument on demo day.
If a vendor is willing to skip discovery to look cheaper, that cost hasn't been removed. It has been moved to the end, where it arrives as rework.
2. Data preparation
Typically 20–40% of the pilot, and the single most underestimated line. This is where budgets die.
The work is unglamorous and unavoidable: locating the data, getting access approved, extracting it from whatever system owns it, cleaning it, dealing with the PDFs that are scans rather than text, deduplicating the four versions of the same policy document, and deciding what to do with the records that contradict each other.
Two specific traps:
Access takes calendar time, not engineering time. A database credential that needs a security review can consume three weeks of a twelve-week pilot while costing nothing in effort. Start the access requests on day one, before anyone writes code.
Document quality is bimodal. A clean, well-maintained knowledge base is a few days of work. A shared drive with fifteen years of accumulated documents, no ownership, and three conflicting answers to every question is a month — and it will also surface an uncomfortable finding, which is that the organisation doesn't actually agree on its own policies. That finding is genuinely valuable. It is not free.
3. The evaluation harness
Typically 10–20%. This is the line that gets cut and shouldn't.
An eval harness is a fixed set of test inputs with known-good answers, plus a way to score outputs against them automatically. It's the difference between "it seems better now" and "quality went from 61% to 78% and latency held."
Without one, every change is a guess. You tweak a prompt, someone tries five queries, it feels better, you ship it — and you have no idea that it quietly got worse on the case that matters to the regulator. With one, you can change a model, a chunking strategy or a prompt and know within minutes whether it helped.
The build cost of a harness is modest: a curated test set (this is the real work — expert time, not engineering time), a scoring function, and a runner. The payoff is that every subsequent iteration gets cheaper, which is why cutting it makes the pilot look cheap and the production system expensive.
4. Build
Typically 25–35%. Less than people expect.
For a retrieval-backed assistant — the most common pilot shape — the build is: an ingestion job, a vector store, a retrieval step, a prompt, a model call, a thin interface, and logging. Modern SDKs have made this genuinely fast. A competent team gets a working end-to-end system in one to two weeks.
The reason this line is small is also the reason pilots mislead. Getting to "it works on a good day" is quick. Getting from there to "it works on a bad day, for a hostile user, with the document that's missing a section" is the rest of the project, and it lives in the next phase, not this one.
5. Inference cost during the pilot
Typically 2–5%. Usually negligible, occasionally not.
For a pilot with a handful of internal testers, model API spend is a rounding error — often tens to low hundreds of dollars. Teams worry about this line far out of proportion to its size.
The exceptions are real, though: anything that processes a large corpus repeatedly (re-embedding a million documents each time you change chunking), anything with a long context on every call, and anything agentic that loops. An agent that retries ten times per task multiplies your per-task cost by ten and is easy to build by accident. Put a hard spend cap on the pilot account on day one. Not because the expected cost is high, but because the tail is.
6. The run cost nobody quotes
This is not a pilot line at all — and that's the problem.
The pilot ends and the question becomes "what does it cost to keep this running?" If that number was never estimated, the decision gets made on the build cost alone, which is the wrong number.
Ongoing cost has four parts: inference at real volume (per-request cost × actual usage, which is usually 10–100× the pilot), infrastructure (vector store, queues, hosting), monitoring, and — the big one — maintenance. Source documents change. Models get deprecated on the provider's schedule, not yours. Retrieval quality degrades as the corpus grows. Someone has to own that, and "someone" is a fraction of an engineer, permanently.
Ask for the annual run cost before the pilot starts. A partner who can't estimate it hasn't run one to production.
Where pilots actually overrun
In our experience the overruns cluster in three places, and none of them is the model.
Data access latency. Covered above, and worth repeating because it's the most common. The fix is procedural, not technical: name the person who can approve access, in week one, in writing.
Scope drift after the first demo. The first working demo is the most dangerous moment in a pilot. It is genuinely exciting, and it generates requests — "can it also do X?" — that are individually reasonable and collectively fatal. The defence is the written not-doing list from discovery. Requests go on it, visibly, for the production phase. Nothing gets added silently.
Moving the success threshold. If accuracy lands at 74% against a 70% bar, and the response is "well, we'd really want 90% before we'd trust it," the pilot has failed to do its job — not because the system underperformed, but because the bar wasn't real. Agree the threshold with the people who will actually use the output, and have them sign it.
A realistic shape
For a mid-sized internal AI pilot — a retrieval assistant over company documents, a handful of test users, a decision at the end — a realistic shape is 8–12 weeks and a small team: one engineer most of the time, a second for parts of it, and meaningful time from a subject-matter expert on your side who can say whether an answer is correct.
That last resource is the one clients under-commit and the one that most determines the result. An AI pilot without expert time is a system nobody can grade.
We won't publish day rates here — they vary too much by region and engagement shape to be useful, and a number without context is worse than no number. What travels is the proportions above. If a quote allocates 80% to build and nothing to evaluation, you are being quoted for a demo, not a decision.
TL;DR
A pilot's job is to answer "is this worth building properly?" — so it must be budgeted to measure, not just to build.
Six lines: discovery (5–10%), data preparation (20–40%, the most underestimated), an evaluation harness (10–20%, the most commonly cut), build (25–35%, smaller than expected), pilot inference (2–5%, usually negligible but cap it), and ongoing run cost (not a pilot line, but the number the decision actually hinges on).
Overruns come from data access delays, scope drift after the first demo, and success thresholds that move. All three are contractual and procedural problems, not technical ones.
Commit expert reviewer time on your side. Define a numeric success threshold before anyone writes code. Ask for the annual run cost up front.
Planning a pilot and want the estimate stress-tested before you commit? Let's talk.
