When Does Self-Hosting an LLM Actually Make Sense?

· Updated

ORGN Team

TL;DR

  • If the model only fits during a single-request test, the sizing decision is already wrong for production. Context length, KV-cache growth, concurrency, and serving overhead can push a deployment beyond one GPU even when the model weights fit comfortably.
  • When the model exceeds one GPU or throughput requirements demand several accelerators, serving becomes an infrastructure problem rather than a simple model deployment. Sharding, replication, GPU communication, request scheduling, and failure handling determine whether additional hardware actually produces useful capacity.
  • A low per-token infrastructure cost does not make self-hosting economical if demand is bursty. Provisioning for peaks creates idle GPUs during normal periods, while provisioning for averages creates queues and latency failures when traffic spikes.
  • Keeping inference inside your own environment changes the trust boundary rather than removing it. Prompts can still reach hosts, GPU memory, logs, traces, backups, and administrators, so self-hosting improves the security posture only when the organization can control those layers.
  • Self-hosting becomes economically or operationally defensible when utilization is sustained, infrastructure locality is mandated, or fine-tuning requires access to model weights. Low or unpredictable demand, frequent model changes, and teams without dedicated inference ownership push the decision toward managed infrastructure instead.
  • If the requirement is confidential execution rather than GPU ownership, running the entire inference stack isn't the only path. Hardware-isolated inference can provide protected execution and verifiable attestation while leaving GPU operations and scaling with a provider, separating the need to control infrastructure from the need to prove how inference handled sensitive data.

A self-hosted LLM can look inexpensive when the model weights are free, and the serving software is open source. The calculation changes when the model needs enough GPU memory for its weights, KV cache, context window, and concurrent requests, because the infrastructure required to keep that model available can cost considerably more than the model itself.

CNCF's 2025 Annual Survey found that 27% of surveyed organizations self-host generative AI in the cloud and another 25% self-host it on-premises, while 37% use managed generative AI services. The survey's inference-hosting breakdown shows that self-hosting is already a real deployment pattern, but it is not the default for most organizations.

Open-weight models are becoming capable enough for production workloads, and running one inside infrastructure you control can provide stronger control over data, model versions, networking, and execution. The harder question is whether your workload justifies the GPUs, memory, serving infrastructure, scaling, monitoring, and engineering ownership that come with it.

A recent discussion from a practitioner running a multi-GPU local LLM server illustrates the other side of the calculation: the hardware bill is only one part of the total cost, and utilization and depreciation can materially change the economics.

The decision therefore starts with the workload, not the model. A self-hosted LLM can make sense when control, predictable demand, model ownership, or data locality justify the operational burden. When those conditions do not exist, managed or confidential inference can provide many of the required controls without requiring the organization to operate the entire inference stack.

What "Self-Hosted" Commits You To

Self-hosting means your organization operates the infrastructure the model runs on, instead of sending inference requests to a provider that manages the serving environment. That can mean owned servers, rented cloud GPUs, or a dedicated tenant configuration. What it always means is that the boundary between "model" and "production service" becomes your team's responsibility to maintain.

The distinction worth drawing early: owning the model weights isn't the same as owning the inference system. An open-weight model eliminates licensing dependency but doesn't eliminate the infrastructure required to turn those weights into a production-grade API. A production inference path looks roughly like this:

Article illustration

The serving team has to configure, version, monitor, patch, and recover every layer in that stack when it fails. A model that runs in a developer environment hasn't been deployed. It's been loaded. Production introduces concurrency, latency SLAs, failure recovery, and capacity planning that a single local inference test won't surface.

Open-weight models are changing the decision

The capability gap between open-weight and proprietary models is no longer a sufficient reason to dismiss self-hosting.

Different models now target different workloads, including code generation, reasoning, extraction, long-context applications, embeddings, and agentic workflows. The practical question is therefore

  • less about whether an open model is universally equivalent to a frontier model, and
  • more about whether a particular open model is good enough for the workload being deployed. That changes the economics.

If an organization can achieve the required quality with a model it can download, inspect, customize, and serve itself, the model API bill is no longer the only variable. Infrastructure utilization becomes the next question, and that's where the calculation gets harder.

GPU Memory Is the Constraint Most Cost Estimates Miss

The capability gap between open-weight and closed frontier models has mostly closed for standard enterprise workloads. As of mid-2026, Qwen 3 72B at INT4 on a single H100 benchmarks comparably to GPT-4-class models on document summarization, code generation, and structured extraction. Llama 4 Maverick, DeepSeek V4-Flash, and the Mistral family all have production deployments at companies that couldn't justify frontier API costs at their token volume.

The closing quality gap makes self-hosting genuinely worth evaluating; what it doesn't change is GPU memory arithmetic.

A model doesn't consume GPU memory only because its weights exist. The serving system also needs memory for the KV cache, intermediate computation, and concurrent requests running in parallel. The relevant question isn't whether a model fits on a GPU; it's whether it fits while maintaining the context length, concurrency, and throughput the application needs under actual production load.

The KV Cache Is Where Production Deployments Break

Every active request has a context the inference engine must hold in GPU memory while generating tokens. Longer contexts consume more memory, and concurrent requests multiply that demand. A model that fits comfortably during a single-user benchmark can become memory-constrained when dozens of users submit long prompts at the same time.

For coding assistants and agentic workloads, this matters more than for chatbots. A short conversational query might sit at 1K tokens. A coding agent working across a repository, reading source files, receiving tool results, and accumulating previous actions and system instructions, can push well past 32K context per request. Fifty concurrent requests at that context length create a KV cache pressure that changes the hardware requirement entirely.

Bash
Single-user benchmark
1K context × 1 request → low KV-cache pressure

Production coding agent
32K context × 50 concurrent requests → KV-cache pressure that
                                         may exceed a single-GPU
                                         deployment

"It fits on one GPU" is a benchmark result, not a production architecture. It doesn't tell you whether the deployment can serve required concurrency, maintain time-to-first-token targets, survive a GPU failure, or handle the traffic variance that shows up in actual usage.

Multi-GPU Inference Turns Model Serving Into Infrastructure

Distributing an LLM across GPUs solves a memory or throughput problem, but it also introduces communication, scheduling, and failure considerations that do not exist in the same form on a single accelerator.

The model must be partitioned or replicated, requests must be scheduled intelligently, and the GPUs must communicate efficiently enough that the added hardware actually improves deployment.

The model can be distributed across GPUs

Inference parallelism provides several approaches for distributing computation, including tensor, pipeline, expert, and data parallelism. CNCF notes that inference parallelism lets models that exceed a single device's memory capacity run across multiple processing units.

A simplified architecture looks like this:

Article illustration

The exact topology matters; GPU-to-GPU communication can become a performance constraint, especially when the model requires frequent synchronization between accelerators. That means adding another GPU doesn't automatically add another equivalent unit of useful capacity.

Replication and sharding solve different problems

Data parallelism can replicate a model across multiple GPUs so separate requests can be handled independently. This is useful when the model already fits on a single device, but you need more throughput.

Model parallelism distributes a model across multiple GPUs when it can't fit in a single device's memory. This distinction matters for cost planning.

If one copy of the model fits on one GPU, adding more GPUs can increase request capacity by running more copies. If the model does not fit on one GPU, additional accelerators become part of the minimum infrastructure required to run the model.

Production inference also needs intelligent scheduling

The GPU is not the only resource that needs managing. A production system must decide which request goes to which model instance, when to batch requests, when to start new capacity, and when to reject or queue requests.

CNCF's work on cloud-native inference highlights why conventional load balancing is insufficient for LLM workloads: request duration, queue depth, model state, and GPU utilization can all influence the correct routing decision. At that point, the organization is no longer simply "running an LLM." It is operating an inference platform.

The Hidden Cost of Running Your Own LLM

The hardware purchase is often the easiest number to calculate; the harder costs are the capacity and engineering decisions required to keep the system useful after deployment. A self-hosted LLM therefore has at least four cost layers: compute, infrastructure, operations, and unused capacity.

GPU capacity has to account for demand

A production service cannot provision only for its average traffic. Consider an application that normally receives a moderate number of requests but experiences a large spike during business hours:

Bash
Request volume

Peak            ████████████████████████
                     ████████████████████████
Average      ████████████
                     ████████████
Low             █████
                     █████
                    ─────────────────────
                    Time →
  1. Provision for the average, and the service can queue requests or violate latency targets during peaks.
  2. Provision for the peak, and the GPUs may sit underutilized for much of the day.

That utilization gap is one of the central economic problems in self-hosted inference.

The infrastructure bill extends beyond GPUs

The production environment can also require:

  • high-bandwidth networking;
  • persistent storage for model artifacts;
  • monitoring and logging;
  • orchestration;
  • load balancing;
  • autoscaling;
  • GPU scheduling;
  • backup and recovery;
  • security patching;
  • redundancy;
  • power and cooling for owned hardware;
  • and engineers to operate the system.

For cloud deployments, the equivalent cost shows up as rented GPU time, storage, networking, managed Kubernetes or other orchestration, and supporting services. The underlying problem doesn't disappear just because you rent the GPUs. It simply moves from capital expenditure to operational expenditure.

Demand changes can make an apparently cheap deployment expensive

A workload that is cheap at 20% GPU utilization can become difficult to operate when demand doubles. More traffic may require additional replicas. Longer contexts may require more memory. A larger model may require multiple GPUs per replica. Higher availability requirements may require spare capacity; the infrastructure can therefore grow faster than requested volume. This is why "the model is open source" is a poor proxy for "the workload will be cheap to self-host."

Self-Hosting Is Not Automatically More Secure

Keeping inference inside infrastructure you control can reduce some forms of third-party exposure, but it does not eliminate the security problem. It changes who enforces the boundary.

The security question becomes less about whether another provider can access the data and more about which administrators, services, hosts, logs, and infrastructure components can access it inside your own environment.

Logs can become another copy of the data

The inference service may not intentionally retain prompts, but surrounding infrastructure can. Request logs, debugging output, traces, error messages, application telemetry, and monitoring systems can all contain fragments of sensitive context. That matters particularly for coding workloads where the prompt may contain proprietary source code, terminal output, credentials, or internal architecture details.

The deployment needs an explicit answer to a simple question:

Where does the prompt exist while the system is processing it?

The answer should include inference memory, caches, logs, traces, backups, and operational tooling.

When Does Self-Hosting an LLM Actually Make Sense?

Self-hosting becomes attractive when the organization has a reason to own the inference environment that justifies the operational cost.

That reason can be security, economics, control, availability, or some combination of them.

Self-hosting makes more sense with predictable, high utilization

HIPAA Business Associate Agreements, GDPR data residency clauses, CMMC Level 2 requirements, and air-gap procurement specifications all define acceptable inference locations at the contract or regulatory level. When a mandate specifies that inference must occur on infrastructure the organization controls, meaning physical or logical ownership rather than account-level tenancy on a shared cloud, managed APIs are out of scope. Self-hosting isn't an optimization in these cases; it's the only compliant path.

Self-hosting makes more sense when model control matters

A single H100 serving a capable 70B-class model at 60% utilization or higher starts to undercut managed inference for the same model at around 2.5 billion tokens per month. At that volume and utilization, the per-token cost advantage is real, and the ops investment pays off. Most teams build toward that threshold as a milestone rather than deploying self-hosted infrastructure on day one.

Self-hosting makes more sense when data locality is a requirement

Fine-tuning requires access to model weights. Managed APIs don't expose weights; self-hosting does. For teams that need domain-specific behavior, such as legal document classification trained on internal contracts or code completion trained on a company codebase, self-hosting is the only path that supports it. The tradeoff is that fine-tuned models require version control, regression testing, and a rollout strategy that managed APIs handle automatically for their own model updates.

When Self-Hosting Probably Does Not Make Sense

Low or unpredictable demand is the clearest counterindication. A GPU required to remain available but receiving few requests means the organization pays for idle capacity, while a managed API would absorb that variability by pooling across customers. Teams that want to evaluate a new model every few weeks face additional friction: each model brings different memory requirements, context limits, quantization tradeoffs, and hardware profiles. The more frequently the model changes, the less stable the infrastructure becomes.

Small platform teams should count engineering time honestly. A production inference deployment needs owners for upgrades, security patches, performance tuning, incident response, GPU capacity decisions, model rollouts, and failure recovery. If nobody owns those responsibilities, the deployment hasn't solved the operational problem; it's moved it into the infrastructure without an owner.

Self-Hosted LLM vs Managed and Confidential Inference

The decision is not necessarily binary. A team can choose between operating the complete inference stack, using a conventional managed API, or using an infrastructure layer that provides stronger confidentiality without requiring the customer to operate the GPUs.

The differences are primarily about ownership, control, security boundaries, and operational burden.

Self-hosted LLMManaged APIConfidential inference
GPU operationsCustomerProviderProvider
Model controlHighProvider-dependentProvider/catalog-dependent
Infrastructure controlHighLowLower
Data localityHighProvider-dependentDepends on execution environment
Scaling responsibilityCustomerProviderProvider
Operational burdenHighLowLower than self-hosting
Hardware execution proofCustomer must implementUsually unavailableAvailable on supported TEE paths
Best fitHigh-control workloadsVariable/general workloadsSensitive workloads without full GPU ownership

The key question is which option gives the organization the required security and control without creating an infrastructure burden that it cannot operate reliably.

A Confidential Alternative to Owning the Entire Inference Stack

Zero-data retention agreements cover a provider's logging and storage behavior. They don't cover what the provider's infrastructure can access during computation. Prompt content exists in plaintext GPU memory during inference on standard hardware, and a ZDR agreement doesn't change that; it's a contractual claim about what the provider does with the data, not a technical constraint on access.

Hardware-isolated inference, using technologies like Intel TDX and NVIDIA confidential compute, encrypts inference data in memory so the host OS, hypervisor, and infrastructure operator can't read it. TEE-based inference produces a cryptographic attestation receipt for each request: a verifiable record that inference ran in an isolated hardware environment and that no third party accessed the prompt or output. For regulated workloads requiring audit evidence of inference security, the difference between a ZDR agreement and a TEE receipt is the difference between a policy document and an independently verifiable artifact.

ORGN Gateway routes requests to TEE model providers, specifically NEAR and Phala using Intel TDX and Tinfoil using AMD SEV-SNP, and logs attestation metadata to ORGN Scanner, where receipts verify against Intel and NVIDIA public PKI without trusting ORGN's claims. The Gateway security model explicitly documents threat boundaries, including what TEE paths protect against (host OS access, hypervisor inspection, infrastructure-level compromise) and what they don't (compromised client environments, data exposure before requests are sent).

Bash
curl -X POST https://api.gateway.orgn.com/v1/chat/completions \
  -H "Authorization: Bearer sk-ollm-YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "phala_deepseek_v3_1",
    "messages": [
      {
        "role": "user",
        "content": "Explain how GPU memory affects LLM inference."
      }
    ]
  }'

The same OpenAI-compatible endpoint routes the request to hardware-isolated inference; the model name determines the execution environment. The current model catalog exposes execution type alongside model information so applications can query available TEE paths programmatically rather than hardcoding model identifiers.

For teams handling regulated or sensitive data who can't satisfy the residency requirement with policy agreements alone, confidential inference shifts the deployment decision from a binary choice into a three-option comparison: self-host for full infrastructure ownership, use a managed API for operational simplicity, or use confidential inference for hardware-backed execution guarantees without owning the GPUs.

What to Measure Before Committing

A good decision starts with workload measurements, not published model benchmarks. The numbers that matter: requests per second, concurrent requests in flight, input and output token counts per request, context length distribution, latency targets (time to first token and inter-token latency separately), and GPU utilization under representative load.

Don't benchmark only short prompts; a coding assistant, retrieval system, and agentic workflow have dramatically different memory and throughput profiles even when they use the same model. Measure the actual application behavior with production-representative inputs before sizing the hardware.

Then measure utilization, not just throughput. A deployment generating 500 tokens per second with one H100 idle most of the time isn't necessarily cheaper than a smaller deployment keeping its hardware busy. The relevant number is cost per completed request at the utilization level the workload produces.

Article illustration

A conceptual comparison of average demand versus peak demand for an LLM inference fleet, showing the unused GPU capacity created when infrastructure is sized for peak traffic. Focus on the gap between average utilization and provisioned capacity. This illustrates why GPU cost cannot be calculated from the model size alone.

The Decision Comes Down to Ownership

Self-hosting makes sense when the workload justifies owning the infrastructure: sustained high volume that keeps GPUs utilized, a residency mandate that rules out managed APIs, or a fine-tuning requirement that needs weight access. Below those thresholds, the infrastructure burden, covering capacity planning, security patching, version management, and ops ownership, tends to cost more than the token savings justify.

For workloads where prompt confidentiality and audit evidence matter more than infrastructure ownership, confidential inference changes the comparison. The organization retains a managed infrastructure model while gaining hardware-backed execution guarantees and verifiable attestation receipts. It's worth distinguishing between "we need to own the GPUs" and "we need to prove what happened to the data during inference." Those are related requirements, but they don't have the same answer.

Article illustration

ORGN Gateway's inference architecture showing the request moving through Gateway into either a TEE or ZDR execution environment, with attestation generated for TEE requests. The key point to notice is that model routing and execution security are separate decisions from the application integration surface.

What the Infrastructure Decision Actually Looks Like

Self-hosting an LLM is no longer primarily a question of whether an open model can run on available hardware. Open models have made the model-access problem easier, but they have exposed the infrastructure problem more clearly: memory, concurrency, multi-GPU execution, utilization, scaling, and operational ownership determine whether the deployment works economically.

The first question should therefore be about the workload. Measure demand, context length, concurrency, latency, and GPU utilization before deciding how much infrastructure the model requires.

The second question is about the control you actually need. If you need model ownership or strict infrastructure locality, self-hosting may be justified. If the requirement is primarily confidentiality and verifiable execution, you may not need to run the entire inference stack.

The practical decision is simple to state even when the implementation is not: self-host when the workload justifies owning the infrastructure, and choose another execution model when the infrastructure burden outweighs the control it provides.

FAQs

1. Is self-hosting an LLM cheaper than using an API?

Self-hosting can be cheaper when inference volume is high, predictable, and sufficient to keep GPUs highly utilized. For low or bursty workloads, idle GPU capacity, operations, networking, and engineering time can make a managed API more economical.

2. How much GPU memory does a self-hosted LLM need?

The model's parameter count is only the starting point. Actual requirements also depend on precision, KV-cache size, context length, concurrency, runtime overhead, and serving configuration. A model that fits during a single-request test may require substantially more memory under production load.

3. Why do large LLMs need multiple GPUs?

A model may exceed the memory available on one accelerator, or the workload may require more throughput than one GPU can provide. Model and inference parallelism can distribute computation across multiple GPUs, but this adds communication, scheduling, and operational complexity.

4. Is a self-hosted LLM more secure than a managed API?

Self-hosting gives the organization more control over infrastructure and data paths, but it does not automatically provide hardware isolation. The answer depends on who can access memory, hosts, GPUs, logs, and inference infrastructure during execution, and how those boundaries are independently verified.