Skip to content
10 min read·The TechKis team

Self-hosted LLMs vs provider APIs — when data residency actually forces your hand

Self-hosting an open-weight model is usually the wrong default and occasionally the only option. Here's what genuinely forces it — regulation, contracts, volume economics — and the hybrid split most teams land on.

  • AI Engineering
  • Architecture
  • Security
  • Cost
  • Enterprise

"We can't send our data to OpenAI" is the sentence that starts most self-hosting conversations. It is sometimes a binding legal constraint and more often an instinct that hasn't been checked. The difference is worth establishing early, because self-hosting an open-weight model is a permanent operational commitment, and taking it on for a reason that turns out to be negotiable is an expensive mistake.

This is about how to tell which situation you're in, and what each path actually costs once the novelty wears off.

What you're actually choosing between

Three options, not two:

Three hosting options, not twoProvider APIs give the best quality with no infrastructure, managed open-weight hosting keeps data in your region on someone else's operations, and self-hosting trades full control for full operational responsibility.Provider APINo infra, best qualityManaged open-weightYour region, their opsSelf-hostedFull control, full opsCOMPARE BY CALLER AND CONTRACT
Figure. Provider APIs give best quality with no infrastructure, managed open-weight hosting keeps data in your region on someone else's operations, and self-hosting trades full control for full operational responsibility.

Provider APIs — OpenAI, Anthropic, Google and the rest, called over the network. No infrastructure. Best available model quality. Per-token pricing. Your data crosses into their processing environment under their terms.

Managed open-weight hosting — Bedrock, Vertex, Azure AI Foundry, Together, Fireworks. Open-weight models served by someone else, often inside your existing cloud account and region. This is the option people forget, and it resolves a large share of residency objections without any of the operational cost of true self-hosting.

True self-hosting — open-weight models running on infrastructure you control, whether that's your cloud account or your own hardware. Full control over where the data sits. Full responsibility for GPUs, serving, scaling, upgrades and uptime.

The middle option matters. A lot of "we must self-host" requirements dissolve when the actual requirement is "the data must stay in the EU and be covered by our existing cloud agreement" — which a managed open-weight endpoint in an EU region satisfies.

What genuinely forces self-hosting

When data residency forces self-hostingRegulation that names a jurisdiction, a customer contract restricting sub-processors, or sustained very high volume push toward self-hosting; most other cases do not.Pick the shapeRegulated jurisdictionSelf-host or in-regionContract bars processorsSelf-hostSustained high volumeHybrid split
Figure. A decision path: regulated data, contractual bans on sub-processors, or extreme volume push toward self-hosting; most other cases do not.

Four reasons hold up under scrutiny.

1. Regulation that names a location or a processor. Some obligations are geographic and non-negotiable: certain public-sector workloads, some health data regimes, some financial supervisory requirements. If your data cannot leave a jurisdiction and no provider offers processing inside it, the question is settled. Note the precision required — "GDPR" alone doesn't force self-hosting; a data processing agreement with an adequate provider in an EU region is a normal, lawful arrangement. The forcing version is narrower and specific.

2. A contract you've already signed. Enterprise customer agreements sometimes prohibit adding sub-processors without consent, or prohibit them outright. This is the most common real constraint we see, and the most frequently missed — it originates in a sales contract, not in engineering, so nobody thinks to check until late.

3. Air-gapped or restricted environments. Defence, critical infrastructure, on-premise deployments at customer sites. If the system runs where there's no route to a public API, self-hosting is the only shape available.

4. Volume economics at genuine scale. At very high, steady, predictable throughput, dedicated hardware beats per-token pricing. Emphasis on all three adjectives — see below.

Reasons that usually don't hold up: "it feels safer" (provider enterprise tiers contractually exclude training on your data and are audited; your own misconfigured S3 bucket is the more likely breach), "we want to avoid lock-in" (the provider is a swappable dependency if you design it that way; self-hosting swaps vendor lock-in for hardware and expertise lock-in), and "it'll be cheaper" (usually false — see below).

The cost question, honestly

Where dedicated capacity overtakes per-token pricingPer-token API pricing stays cheaper at low and medium volume because dedicated GPU capacity is billed whether or not requests arrive; the crossover sits at sustained production throughput.10k/mo100k/mo1M/mo10M/moCrossover sits high
Figure. Per-token API pricing wins at low and medium volume; dedicated capacity only overtakes it at high, steady throughput.

The cost comparison people make is API tokens versus GPU rental. That comparison flatters self-hosting because it omits most of the cost.

API pricing is linear and has no floor. Ten requests a day costs cents. Utilisation is someone else's problem.

Self-hosted pricing is a floor plus a ceiling. You pay for the GPU whether or not a request arrives. A serious open-weight model needs meaningful accelerator memory to serve at usable latency, and that instance runs continuously. At low volume you are paying thousands per month for capacity you aren't using. At high volume the same instance absorbs enormous throughput and the per-request cost collapses.

That's the crossover, and it sits higher than most people assume — typically well into sustained production volume, not pilot volume.

The self-hosting costs the GPU quote omitsA GPU-versus-tokens comparison leaves out the continuously billed capacity floor, redundancy, idle utilisation, specialist engineering time and model upgrade churn.GPU floorRedundancyIdle capacityEngineeringUpgradesBUDGET
Figure. A GPU-versus-tokens comparison omits the continuously billed capacity floor, redundancy, idle utilisation, specialist engineering time and model upgrade churn.

What the GPU-vs-tokens comparison leaves out:

  • Engineering time. Serving stack, batching, quantisation choices, autoscaling, failover, upgrades. This is a specialist skill set and it is not cheap to hire or to borrow from your existing team.
  • Redundancy. One instance is a single point of failure. Two is double the floor cost.
  • Idle capacity. Traffic is bursty; GPUs are billed continuously. Utilisation of 20% means your effective per-token cost is five times the naive calculation.
  • Model upgrades. A new open-weight release doesn't arrive as a version bump. It's a re-evaluation, a re-tune of serving parameters, and a redeploy.
  • The quality gap. Frontier proprietary models still lead on the hardest reasoning tasks. If a smaller open-weight model needs more retries, more scaffolding or more human correction to reach the same outcome, that's a real cost that never appears in the infrastructure line.

To be fair to the other side: open-weight quality has improved enormously, and for a large class of tasks — classification, extraction, summarisation, routing, structured output over a known domain — a mid-sized open model is entirely sufficient. The gap that matters is narrower than it was, and it's narrowest exactly where high-volume workloads live.

The hybrid split most teams land on

Most teams with a genuine residency constraint don't end up all-in on either option. They split by data class.

WorkloadWhere it runsWhy
Bulk classification, extraction, routing, embeddingsSelf-hosted or managed open-weightHigh volume, narrow task, sensitive data, quality gap irrelevant
Complex reasoning over non-sensitive dataProvider APIBest quality where it matters, low volume
Anything touching regulated recordsSelf-hosted / in-region managedConstraint is binding
Internal developer toolingProvider APINo sensitive data, convenience wins

This works only if the model is a swappable dependency. Which brings us to the architectural point that matters more than the hosting decision itself.

Design so the decision stays reversible

The expensive mistake isn't picking the wrong option. It's building so that changing your mind means a rewrite.

Put every model call behind one internal interface. A single module in your codebase that owns "talk to a model". Everything else calls that. Swapping a provider, adding a self-hosted endpoint for one workload, or routing by data class then becomes a change in one place. We've written about this as the AI gateway pattern — the same seam that gives you budgets, redaction and audit trails is the seam that makes hosting reversible.

Keep prompts and evaluation sets provider-neutral. Prompts tuned to one model's quirks are a migration cost. Your evaluation set is what makes a swap a measurement instead of a leap of faith.

Classify data at the boundary, not in the prompt. Decide what class a request is before choosing where it goes. Routing by data sensitivity is a policy decision that belongs in code you can audit.

Prove the constraint before you build for it. Get the actual clause. Read it. Ask the person who owns the obligation whether an in-region managed endpoint under your existing cloud agreement satisfies it. That conversation takes a day and has repeatedly saved clients a quarter of infrastructure work.

TL;DR

Three options, not two — provider APIs, managed open-weight hosting in your region, and true self-hosting. The middle option resolves most residency objections without the operational burden, and it's the one people forget.

Self-hosting is genuinely forced by: regulation naming a jurisdiction or processor, a customer contract restricting sub-processors (the most commonly missed one, because it lives in sales, not engineering), air-gapped environments, and sustained high volume.

It's usually not forced by instinct, lock-in worry, or cost — the cost crossover sits higher than expected once idle GPU capacity, redundancy, engineering time and upgrade churn are counted.

Most teams with a real constraint split by data class: open-weight for high-volume narrow tasks and regulated data, provider APIs for hard reasoning on non-sensitive data.

Whatever you choose, put every model call behind one internal interface and keep prompts and evaluation sets provider-neutral. That's what makes the decision reversible — and at the rate models change, reversibility is worth more than getting it right the first time.

Working through a residency constraint and want a second opinion before committing to infrastructure? Let's talk.

Back to all insights