Skip to content
9 min read·The TechKis team

Choosing an AI development partner — the questions that predict delivery

Case studies and logos don't predict whether an AI project ships. These questions do — about evaluation, data access, model choice, handover and run cost — plus the answers that should end the conversation.

  • AI Engineering
  • Engineering Practice
  • Enterprise
  • Cost

Every agency pitching AI work will show you a deck with the same things in it: logos, a capability matrix, a slide about "our proven methodology", and a demo that works. None of it predicts whether your project will ship, because none of it is falsifiable. The demo works because demos are selected to work.

What does predict delivery is how a partner answers a small number of specific questions — the ones where a team that has taken AI systems to production gives a concrete, slightly boring answer, and a team that has only built demos gives a confident abstract one.

This is that list. We're an agency that does this work, so read it with that in mind — but the questions are useful regardless of who you point them at, including us.

Why the usual signals fail here

Traditional software procurement signals transfer badly to AI projects for one structural reason: AI systems are probabilistic, so "it works" is a measurement, not an observation.

A conventional web app either renders the checkout page or it doesn't. You can verify it by looking. A retrieval assistant produces plausible-sounding output on every single input, including the ones where it is completely wrong. Looking at it tells you almost nothing. This means a demo has near-zero information content — and a portfolio of demos has near-zero information content at scale.

A demo is not evidence, an evaluation isA demo runs selected inputs with no baseline and looks convincing either way, while an evaluation scores a held-out set against expected answers and tracks the score across versions.DEMOSelected inputsLooks rightNo baselineEVALUATIONHeld-out setScored answersTracked per version
Figure. A demo runs selected inputs with no baseline; an evaluation scores a held-out set against expected answers and tracks the score across versions.

The second structural difference: the expensive risks sit in your data and your organisation, not in the model. A partner who has been through this will be interrogating you about data access, document ownership and reviewer availability in the first conversation. One who isn't asking hasn't hit the wall yet.

The questions

Signals that predict AI deliveryEvaluation evidence, clear ownership of code and prompts, and an honest run-cost estimate predict delivery better than case studies or demos.EvaluationOwnershipRun costASK FOR ARTEFACTS, NOT SLIDES
Figure. Evaluation evidence, code and data ownership, and honest run cost are the three signals worth weighting.

"How will we know if it's working?"

Ask it first, and weight it heaviest.

Good answer: a description of an evaluation set — real inputs from your domain, expected outputs graded by your experts, scored automatically, tracked across versions. They'll ask who on your side can grade answers, and they'll want that person's time written into the plan. They may also volunteer what their harness can't catch.

Bad answer: "we'll test it thoroughly", "we'll do a UAT phase", or a pivot to talking about the model's benchmark scores. Public benchmarks tell you nothing about performance on your documents.

Conversation-ending answer: "the model is very accurate."

"What happens when it gets an answer wrong in production?"

Good answer: a specific design — confidence signals, citation of source passages so a human can verify, an escalation path to a person, logging of every input and output so a bad case can be reproduced, and a process for turning that case into a new test. They'll talk about making failure visible and cheap rather than claiming to eliminate it.

Bad answer: an assurance that it won't happen, or "we'll add guardrails" with no detail about what a guardrail is in this context.

Any team that has run an LLM system in front of real users has a war story here. Ask for it. The absence of one is itself information.

"Which model would you use, and what would make you change it?"

Good answer: a provisional choice with reasoning tied to your constraints — cost per request at your volume, latency budget, context needs, where the data is allowed to be processed — plus an explicit statement that the choice is replaceable and that swapping models is a config change measured against the eval set.

Bad answer: deep commitment to one provider as an identity, or a refusal to name anything until "discovery is complete." Models change every few months. Architecture that treats the model as a swappable dependency is a sign of production experience; architecture built around one provider's SDK throughout is a sign of the opposite.

"Who owns the code, the prompts and the data?"

Get this in writing, and read the clause yourself.

Good answer: you own everything — repository, prompts, evaluation set, fine-tuned artefacts and embeddings. The partner may keep generic internal tooling, stated explicitly. Your data is processed under your accounts where possible, never used for training, and deleted on request.

Bad answer: vagueness, or hosting on the partner's infrastructure with no migration path. A prompt library and an evaluation set are the real intellectual property of an AI system — worth more than the code, and easier to quietly retain.

Red flag: the system only runs on a platform the partner controls, and the contract doesn't say what happens when you leave.

"What will this cost to run per month, at our volume, in year two?"

Good answer: a range with the arithmetic shown — requests per month × tokens per request × price per token, plus infrastructure, plus a named maintenance allocation. They'll flag which input dominates the estimate and how wrong it could be.

Bad answer: "that depends on usage" with no model offered. It does depend on usage — so ask them to price three usage scenarios.

Teams that have only delivered pilots genuinely don't know this number. That's the tell.

"What does handover look like?"

Good answer: documentation of the architecture and its failure modes, a runnable evaluation suite in your repository, a runbook for the common operational problems, a session with your engineers, and a defined support window. They'll be comfortable with you taking it in-house — a partner confident in the relationship isn't protecting it with opacity.

What a real handover containsA genuine handover leaves documented architecture and failure modes, a runnable evaluation suite in your repository, and an operational runbook.Architecture docsEval suiteRunbookHANDOVER IS A DELIVERABLE, NOT A FAVOUR
Figure. A real handover leaves documented architecture and failure modes, a runnable evaluation suite in your repository, and an operational runbook.

Bad answer: handover as an afterthought, or an ongoing retainer as the only way to keep the system alive.

"What did you get wrong on a previous project?"

Good answer: a specific story with a specific lesson — a chunking strategy that broke tables, a pilot that overran because of a three-week access approval, a model deprecation that forced a re-evaluation. Delivered without drama.

Bad answer: a non-answer dressed as a strength ("we're perfectionists, so sometimes we over-engineer").

This question is cheap and unusually diagnostic. Anyone who has shipped has scars, and people who have genuinely learned from them tend to be relaxed about naming them.

Engagement shape

Matching engagement shape to what is uncertainFixed price suits genuinely known scope, time and materials suits scope that will evolve, and a capped discovery suits unproven feasibility.Fixed priceKnown scopeTime & materialsEvolving scopeCapped discoveryUnproven feasibilityCOMPARE BY CALLER AND CONTRACT
Figure. Fixed price suits known scope, time and materials suits evolving scope, and a capped discovery suits unproven feasibility.

The commercial structure matters as much as the team, and the right structure depends on what's uncertain.

ShapeFits whenMain risk
Fixed priceScope is genuinely known and stable — an integration, a defined migrationChange requests become adversarial; quality gets squeezed to protect margin
Time and materialsScope will evolve as you learn, which is most AI workRequires trust and active management, or it drifts
Capped discovery, then decideFeasibility is unprovenAlmost none — this is usually the right opening move

Fixed price on a project whose scope depends on what the data turns out to look like is a trap for both sides. The partner prices in a large risk premium, and every discovery becomes a negotiation instead of a decision.

The structure that works most reliably for AI work: a small fixed-price discovery with a defined deliverable — a written problem statement, a data assessment, an evaluation set and a feasibility verdict — followed by a decision point where you can walk away, with the discovery artefacts, having spent a bounded amount. A partner who resists that structure is asking you to take all of the feasibility risk.

Practical checks that cost nothing

Ask to meet the engineers who would actually do the work. Not the account lead. If the people in the room are not the people on the project, ask why.

Give them a real problem in the first conversation. Describe an actual document your system would need to handle and ask how they'd approach it. Watch for whether they ask about its structure, volume and quality — or jump straight to a solution.

Check whether they'll tell you not to build it. The most useful thing a partner said to us, when we were on the buying side, was that the problem we described didn't need a model at all. A partner who never steers you away from AI is selling AI, not solving your problem.

Start small and structured. A bounded first engagement with a real deliverable tells you more about a team in four weeks than any reference call.

TL;DR

Demos and logos don't predict AI delivery, because a probabilistic system produces plausible output even when it's wrong. Evaluation discipline predicts it.

Ask: How will we know it's working? What happens when it's wrong in production? Which model, and what would change it? Who owns the code, prompts and data? What's the monthly run cost in year two? What does handover look like? What did you get wrong last time?

Weight the evaluation answer heaviest. Get IP ownership in writing, including prompts and evaluation sets. Treat "that depends on usage" without an offered model as a sign of pilot-only experience.

Prefer a capped discovery with a real deliverable and a genuine walk-away point over a fixed price on unknown scope.

Meet the engineers. Bring a real problem. Notice whether they're ever willing to tell you not to build it.

If you want to put these questions to us directly, that's what the contact form is for.

Back to all insights