⚡The verdict up front
This is not a religious question, whatever the internet says. It is a utilisation question plus a compliance question, and both have concrete answers once you measure your own workload.
Hosted APIs win by default: zero operations, frontier quality on day one, and a cost that scales to zero when usage does. Self-hosting wins under specific, measurable conditions — sustained volume that keeps expensive GPUs occupied, or data that legally cannot leave your infrastructure. Most production systems we build end up hybrid, and the architecture that matters is the one that lets you move workloads between the two without a rewrite.
| Dimension | Hosted API | Self-hosted open weights |
|---|---|---|
| Cost structure | Per token, pure variable, scales to zero | Mostly fixed GPU cost + platform-engineering time |
| Break-even condition | Bursty or moderate volume | Sustained high utilisation |
| Model quality | Frontier models available immediately | Strong open models; behind the frontier on the hardest tasks |
| Data privacy | Enterprise terms: no training, defined retention | Data never leaves your infrastructure |
| Ops burden | None | Real — serving, scaling, upgrades, monitoring |
| Time to production | Days | Weeks to months |
| Best for | Most teams, most workloads, most of the time | Hard residency requirements or sustained high volume |
🧮The cost break-even math: GPU hours vs per-token pricing
The economics are simpler than vendor marketing on both sides suggests. An API charges per token — pure variable cost, nothing when idle. A GPU charges per hour — pure fixed cost, running whether or not it serves a single token. Break-even is the point where the tokens you would have bought per hour cost more than the GPU-hour that produces them.
On-demand H100-class GPUs run roughly $2 to $8 per GPU-hour depending on provider and commitment. A quantised 7–8B open model serves comfortably from a single mid-range card; a 70B-class model at useful precision needs roughly two H100s; the largest open models need a multi-GPU node. Against that, published API pricing for capable hosted models runs from well under a dollar to around $25 per million tokens depending on tier and direction — output tokens cost several times input tokens on every provider's rate card. Both sets of numbers move, so check current pricing before budgeting.
What we will not do is hand you a single break-even token count, because it depends on throughput — and throughput depends on model size, quantisation, context length, batching and latency targets, which vary by an order of magnitude across real deployments. What we can say from building these systems: at steady, high occupancy, a well-run GPU fleet frequently undercuts API pricing on the same workload, sometimes substantially. The same GPUs at low or bursty utilisation cost more than the API would have, because you pay for the hour, not the token.
So the first step is not procurement — it is measurement. Run your workload on an API for a quarter, log real token volumes by hour, and look at the shape. Flat and heavy favours self-hosting. Spiky and moderate favours staying on the API. Most workloads are spikier than their owners assume.
Full LLM integration cost breakdown: build vs runWhat AI agents actually cost per task
The pattern we recommend: start on a hosted API behind an abstraction layer, measure your real token shape for a quarter, then self-host only the workload slices where the utilisation math clearly wins — if any do.
🧠Open-weight model quality: where Llama, Qwen and Mistral actually stand
Open-weight models are genuinely good in 2026, and the "open models are toys" take is years out of date. Meta's Llama family, Alibaba's Qwen family and Mistral's models cover the useful size range from single-GPU 7–8B models up to large mixture-of-experts systems, and they are strong at the tasks that dominate production workloads: classification, extraction, summarisation, drafting with clear structure, translation, and code assistance in mainstream languages.
The gap that remains is at the frontier: the hardest multi-step reasoning, long-horizon agentic tool use, and ambiguous instruction-following under messy real-world context. On those tasks the frontier hosted models still lead, and for some products — a coding agent, a complex research assistant — that lead is the product. For a large share of enterprise features, the frontier is overkill and a well-prompted mid-size open model is indistinguishable in production.
One quality lever self-hosting gives you is fine-tuning the actual weights. On a narrow, repeated task, a fine-tuned open model can match or beat a prompted frontier model at a fraction of the size — which changes the cost math again. Hosted providers offer fine-tuning too, but only on their models, at their prices, with your data shaping weights you cannot take with you.
The honest summary: benchmark leaderboards shift every few months, so do not pick a model — pick an evaluation set of your real inputs and score candidates against it. That is the only quality number that matters, and it is portable across every future model.
🔒Privacy and compliance: the driver that is not about cost
For a meaningful slice of organisations, this section settles the decision before economics enter the room. Healthcare data under HIPAA, financial data under various regimes, government-adjacent work, or contractual obligations to customers can all require that prompts and documents never touch a third party's infrastructure. When that requirement is hard, self-hosting inside your own boundary is the answer and the cost discussion becomes about doing it efficiently.
But check the requirement before accepting it. The major API providers offer enterprise terms — zero data retention options, no training on your data, regional processing, BAAs for healthcare, SOC 2 attestations — that satisfy a large share of real compliance frameworks. Legal and security teams sometimes default to "self-host" as the conservative answer when a reviewed enterprise agreement would meet the actual obligation. The review is worth the weeks it takes, because it can save a permanent platform-engineering commitment.
The middle cases are where hybrid architectures earn their keep: sensitive workloads (patient data, deal documents) routed to self-hosted models inside the boundary, everything else on APIs. Sensitivity-based routing is a data-classification exercise, not an infrastructure one, and it is far cheaper than self-hosting the entire surface.
🛠️The operations burden nobody puts in the spreadsheet
Self-hosting an LLM is not "deploy a model". It is running an inference service, and inference services have the same operational demands as any latency-sensitive production system — plus some that are specific to GPUs.
The serving stack — vLLM, SGLang or similar — handles continuous batching and KV-cache management, and it is genuinely good. Around it you still own: autoscaling GPU capacity against traffic, health checks and failover, model version upgrades (which can shift behaviour on your prompts and require eval re-runs), monitoring of latency and throughput per request, cost attribution per team or feature, and an on-call rotation that understands the stack. Budget 0.25 to 1 platform-engineer FTE depending on fleet size and availability requirements, forever.
GPU capacity itself is an operational problem. The cards you want are frequently capacity-constrained at the major clouds, reserved instances commit you to spend before you have proven utilisation, and a saturated GPU serving at its batching limit degrades latency for everyone at exactly the moment traffic peaks. Teams that have run Kubernetes for years still find GPU fleets a new discipline.
There is also a cost multiplier that only shows up in production: context. Long prompts — RAG pipelines re-reading retrieved chunks, agents carrying conversation history and tool results — push far more input tokens through the model than a toy demo suggests, and long-context serving consumes KV-cache memory that directly limits how many requests a GPU can batch. The workload you benchmark in a proof of concept is rarely the workload you operate, and the gap runs against you on both the API and the self-hosted side.
If nobody in the organisation will own this service — watch it, tune it, upgrade it, get paged for it — that fact answers the build-vs-rent question more decisively than any cost model.
🔀Hybrid patterns that work in production
RAG development services for grounded, private assistants
API front door, self-hosted high-volume slices
Everything starts on the API. As volume grows, specific workloads with stable shape — classification, extraction, a fixed-format drafting task — move to a fine-tuned open model on your GPUs. This is the most common mature architecture we see.
Sensitivity-based routing
Requests are classified by data sensitivity. Sensitive traffic goes to self-hosted models inside your compliance boundary; the rest uses APIs. The routing decision is a data-classification function, auditable and cheap.
Quality-tier routing
A small cheap model handles the request; an escalation policy sends the hardest cases to a frontier API. You pay frontier prices on a small fraction of traffic and mid-tier prices on the rest.
API as failover and burst capacity
Self-hosted capacity sized for normal load, with overflow and disaster-recovery routed to an API. This fixes the peak-latency problem of saturated GPUs without paying for idle headroom.
The abstraction layer that makes all of it possible
A thin internal model gateway — one interface, per-workload routing config, per-request tracing of tokens and cost. Every pattern above is a config change behind it and a rewrite without it. Build this first, whichever direction you go.
🎯How to decide, in order
1. Check the compliance requirement for real
Is self-hosting a hard legal obligation, or a default assumption? Review the providers' enterprise terms with counsel first — they satisfy more frameworks than teams assume.
2. Measure your token shape on an API
One quarter of real usage, logged by hour and by workload. This number drives everything downstream and costs almost nothing to collect.
3. Score open models on your eval set
Not leaderboards — your inputs, your correctness criteria. If a self-hostable model clears your bar, self-hosting is an option. If only frontier models clear it, it is not.
4. Run the utilisation math honestly
GPU-hours at realistic throughput and occupancy versus your measured API bill, plus the platform-engineering cost. If utilisation is bursty, stop here and stay on the API.
5. Decide who owns the service
An inference fleet without a named owner is a future outage. If the owner does not exist, the API is the answer regardless of the spreadsheet.
