TL;DR
- Once prompts enter the serving stack, administrators, runtime processes, monitoring systems, GPU memory, and other infrastructure components become part of the trust boundary.
- Context length, concurrent requests, KV-cache requirements, and serving configuration can push a deployment toward additional GPUs, lower concurrency, cache offloading, or quantization.
- The organization gains control over storage and network placement, but it also becomes responsible for privileged host access, GPU isolation, logging, patching, serving software, and operational failures.
- ZDR can restrict what happens to a prompt after processing, but it doesn't establish who could access plaintext during inference or provide cryptographic evidence about the execution environment.
- When the attestation binds the model identity and request and response hashes to CPU and GPU hardware evidence, a verifier can check properties of the execution environment for that specific inference session.
- Self-hosting, ZDR, and TEE-backed inference place responsibility in different parts of the stack, so the decision should follow the required trust boundary, workload shape, and operational capacity.
What Secure LLM Deployment Actually Requires
A 70B-class model doesn't become a private system simply because its weights sit on infrastructure you own. The moment a prompt enters an inference server, it exists across tokenizer memory, GPU memory, KV cache, runtime processes, and whatever logging or monitoring surrounds the server.
A secure LLM deployment therefore needs more than a private model endpoint. It needs defined controls for who can read inference data, who can modify the execution environment, what gets retained, which identity issued a request, whether workloads remain isolated, and whether the claimed execution environment can be independently verified.
Gartner projects worldwide data-center electricity consumption to rise from 448 TWh in 2025 to 980 TWh by 2030, with AI-optimized servers accounting for a large share of that increase. A practitioner building a B2B AI product described a similar tradeoff on Reddit after deploying private models, reporting substantially higher infrastructure costs and performance compromises than commercial APIs.
The article below separates model capability from inference infrastructure, then maps the security properties that matter once a model processes sensitive data. It also compares self-hosted inference with ZDR and TEE-backed execution, including what evidence an enterprise should demand from an LLM provider.
Why Better Open Models Make Self-Hosting More Attractive
The model-selection decision has changed because open-weight models have narrowed the performance difference with closed systems. The infrastructure decision hasn't become simpler just because the model catalog has improved.
Open-weight capability changes the deployment calculation
The top closed model led the top open model by 3.3% in March 2026 on Stanford's measured performance comparison, after the difference had narrowed to 0.5% in August 2024. The same report notes that six of the top ten Arena models were closed-weight at that point, so the data doesn't support treating open and closed models as interchangeable across every workload.
For an engineering team, the more useful question is whether a selected open-weight model is good enough for a defined workload while giving the organization control over the inference environment. Code generation, retrieval-heavy applications, internal assistants, and domain-specific workloads can produce very different answers from a benchmark leaderboard.
Self-hosting becomes attractive when data control, model availability, deployment control, or isolation requirements justify taking ownership of the serving stack. The tradeoff appears after model selection: somebody still has to provision GPUs, load weights, manage memory, handle concurrency, patch the serving layer, and operate the machines.
The Model May Be Open, but Inference Still Needs Hardware
Model weights describe only one part of the memory required to serve an LLM. Request length, concurrency, cache configuration, precision, and the serving runtime all affect how much accelerator memory the deployment consumes.
Model weights aren't the entire memory budget
vLLM exposes KV-cache configuration separately from model execution memory and tracks GPU blocks, CPU blocks, cache capacity, and maximum concurrency. Its current configuration also sets gpu_memory_utilization as a per-instance limit for the model executor.
The distinction matters for security because the prompt doesn't disappear after tokenization. The request continues through runtime memory while the model generates tokens, and the KV cache holds an intermediate attention state associated with active requests.
A deployment that fits the weights into available VRAM therefore isn't automatically sized for production traffic. Increasing context length or concurrent requests changes the memory pressure, which can force cache offloading, lower concurrency, quantization, or additional GPUs.
Multi-GPU serving creates another trust boundary
ORGN's current TEE deployment documentation describes an inference environment using eight NVIDIA H100 GPUs, with Intel TDX protecting the confidential VM and NVIDIA confidential computing extending protection into GPU memory.
The hardware requirement isn't merely a capacity concern. Sensitive inference data moves through an execution stack that includes CPUs, accelerators, drivers, serving processes, and interconnects, so each layer needs an explicit trust assumption.
A self-hosted team therefore trades one provider relationship for direct responsibility over the infrastructure boundary. The security review meeting shifts from provider-side data retention to administrator, process, and runtime access to plaintext prompts and model state during inference.
The ORGN attestation shows the CPU and GPU sides of its TEE deployment, including Intel TDX and NVIDIA H100 hardware. Look how ORGN separates CPU and GPU trust components, because protecting the virtual machine alone doesn't establish protection for GPU memory during inference.

Self-Hosting Removes One Trust Relationship and Creates Several Others
Owning the server changes who controls the physical and software infrastructure, but ownership doesn't remove privileged access. The important question becomes which operators and services sit outside the model's intended trust boundary.
Infrastructure administrators still matter
A self-hosted LLM runs somewhere under an operating system, hypervisor, container runtime, orchestration layer, or some combination of those components. Administrators with sufficient privileges can interact with the host environment even when the application itself exposes only an authenticated API.
Monitoring introduces another path: metrics, debugging tools, backups, and log collectors sit outside the model's core execution logic and therefore need their own data-handling controls.
ORGN's own security material makes the same distinction for self-hosted systems: keeping code or prompts away from a provider doesn't inherently prevent execution-time access by the infrastructure hosting the model.
Security review needs to follow the request
A useful review starts with the request lifecycle rather than the model card: Client → TLS termination → API server → Tokenizer → CPU memory / GPU memory / KV cache → Model execution → Logs / telemetry.
Every stage introduces a different question; TLS protects the network connection, but it doesn't by itself determine who can inspect plaintext after termination, while retention controls don't determine who can access memory during execution.
What Does "Secure LLM" Actually Mean?
"Secure LLM" describes several independent properties rather than one technical control. Treating the phrase as a single security guarantee makes architecture reviews much harder because each deployment can satisfy one property while leaving another unresolved.
Confidentiality is only one property
A security review should separate the seven different verticals:
| Property | Question |
|---|---|
| Confidentiality | Who can read the prompt during inference? |
| Integrity | Can someone modify the execution environment or model? |
| Retention | What happens to the request after inference? |
| Identity | Can the request be tied to an authenticated caller? |
| Isolation | Can another workload access the same execution state? |
| Provenance | Can you establish where inference occurred? |
| Auditability | Can you reconstruct what happened afterward? |
The distinction is important because different controls answer different questions. API authentication addresses identity, ZDR addresses retention, and hardware attestation addresses properties of the execution environment.
NIST's Generative AI profile treats AI risk management as a lifecycle activity spanning design, development, use, and evaluation rather than a single security mechanism. The same approach fits inference security: define the property first, then identify the control that establishes it.
A private endpoint doesn't answer every question
A self-hosted endpoint can give an organization direct control over storage and network placement while leaving host-level access under the same administrative domain. A ZDR service can remove post-request retention while leaving execution under the provider's infrastructure.
TEE-backed inference approaches the problem differently. The execution environment itself becomes part of the security mechanism, and attestation provides evidence about the environment rather than relying solely on a written retention statement.
Encryption Stops Being Enough Once Inference Starts
TLS protects the connection between the client and endpoint, while storage encryption protects persisted data. Neither control describes what happens to plaintext after the inference server accepts the request.
The inference path contains plaintext states
A typical serving pipeline needs to parse the request, tokenize text, prepare tensors, execute attention layers, maintain intermediate state, and return a response. Sensitive content therefore exists inside memory during computation even when the network and storage layers use encryption.
Hardware-backed confidential computing addresses that specific boundary by isolating memory from privileged infrastructure components. ORGN describes its TEE routes as combining Intel TDX confidential virtual machines with NVIDIA H100 confidential GPU compute, extending hardware-backed isolation from the VM to GPU memory.
ORGN's verification diagram traces a TEE request from submission through secure execution, measurement, attestation, and verification. Check where the hardware-generated evidence appears relative to inference, because the security claim is tied to the execution environment rather than the API connection.

Zero Data Retention Solves a Different Problem
ZDR answers what happens to data after processing, while execution security asks who can access data during processing. Treating those properties as interchangeable produces an incomplete threat model.
Retention, runtime access, and execution integrity are separate
A provider could agree not to retain prompts while still processing those prompts inside infrastructure controlled by the provider. A customer could also delete every request after inference while privileged infrastructure operators retain access to memory during execution.
ORGN's Gateway separates these execution types explicitly; TEE models produce hardware-backed attestation, while ZDR models use provider agreements and don't produce hardware attestation receipts.
A ZDR policy isn't an execution receipt
A retention policy describes what the provider agrees to do with stored data. An attestation receipt supplies cryptographic evidence about the environment that processed a particular request. The two controls therefore sit at different points in the trust model. One governs data handling after processing; the other establishes the environment's properties during processing.
What a Secure Inference Architecture Actually Looks Like
The architecture becomes easier to evaluate once model selection, request orchestration, execution, and verification sit in separate boxes. ORGN Gateway uses that separation to expose TEE and ZDR execution through an OpenAI-compatible interface.
Model selection remains an application decision
The current Gateway architecture doesn't automatically select or substitute a model. The request's model field determines the execution environment, with near_*, phala_*, and tinfoil_* models using TEE routes and vercel_* models using the ZDR path. It's important because the API surface stays similar while the security properties change with the selected model.
The ORGN Gateway architecture separates the client application, Gateway control plane, execution environments, and attestation layer. The architecture separates the API integration surface from the security mechanism used during execution. That separation is useful when an application needs access to different model classes without treating every model as having identical security properties.

A TEE request is still an ordinary API request
ORGN's Gateway integration uses the standard /v1/chat/completions endpoint with a model identifier and an API key. The following request uses phala_deepseek_v3_1, which the current integration documentation identifies as a TEE model, so the operational reason for specifying the model explicitly is to select the execution class rather than leaving routing to an opaque policy.
curl -X POST https://api.gateway.orgn.com/v1/chat/completions \
-H "Authorization: Bearer sk-ollm-YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "phala_deepseek_v3_1",
"messages": [
{
"role": "user",
"content": "Process this confidential engineering request."
}
]
}'The API key should remain server-side, and the model identifier should come from the current catalog rather than being copied indefinitely into application configuration.
What TEE Attestation Proves, and What It Doesn't
Attestation becomes useful when a security claim needs evidence rather than a provider statement. The evidence has a defined scope, so the verification process needs to stay within that scope.
The receipt ties hardware evidence to inference
ORGN's current attestation format combines an Intel TDX quote, NVIDIA GPU evidence, and a message signature binding model identity to request and response hashes. The TDX quote and GPU evidence also share a session nonce, linking the CPU and GPU evidence to the same execution session.
The resulting evidence addresses questions such as whether the inference ran inside the expected Trust Domain and whether the environment matched the expected measurements. It doesn't certify the model's behavior, the client's machine, or anything that happened before the request entered the protected execution boundary.
ORGN's attestation reference breaks the receipt into the Intel TDX quote, per-GPU NVIDIA evidence, and message signature. The important detail is the cryptographic binding between model identity, request and response hashes, and the hardware evidence, which gives the verifier something more specific than a provider status page.

Evidence still has an explicit boundary
ORGN's verification lists compromised clients, leaked API keys, unsafe model behavior, exposure before request submission, exposure after response delivery, and vulnerabilities in model weights outside the TEE guarantee. That limitation is a strength of the security model rather than a weakness in the explanation. A security control becomes useful when engineers know exactly what it establishes and where its authority ends.
The Cost of Owning the Inference Stack
Self-hosting changes the cost model from per-request service consumption toward infrastructure ownership. GPU capacity is only the visible part of that calculation.
GPU capacity is the first line item
Model weights consume accelerator memory, while KV cache and concurrent requests consume additional memory. vLLM exposes separate configuration for KV-cache capacity and GPU memory utilization, illustrating why the serving workload affects the hardware requirement beyond the model's raw parameter count.
The infrastructure bill then expands into serving software, networking, patching, model updates, capacity planning, observability, and operational support. A Reddit discussion from a B2B AI team illustrates the same calculation from the operator side, where the team compared multi-GPU instances with commercial APIs and reported substantial cost differences.
Security becomes part of the operating burden
A self-hosted team owns the controls around the model as well as the model itself. That includes deciding who has privileged host access, where logs go, how secrets enter the process, how GPU access gets isolated, and what happens when an operator needs to debug a production failure.
The economics therefore depend on workload shape; a steady workload with strict locality requirements has a different calculation from an intermittent workload that needs large models only during occasional bursts.
The same distinction applies to security architecture: the right deployment depends on which property needs proof and which operational responsibilities the organization is prepared to own.
A Practical Security Model for Self-Hosted and Managed Inference
The useful comparison isn't "private versus public." The useful comparison asks which security property each deployment establishes and who carries the operational responsibility.
| Property | Self-hosted inference | ZDR managed inference | TEE-backed inference |
|---|---|---|---|
| Data location | Customer-controlled | Provider infrastructure | Provider infrastructure |
| Retention | Customer-controlled | Provider agreement | Execution-dependent |
| Host access | Customer responsibility | Provider trust boundary | Hardware-isolated execution |
| Execution proof | Custom implementation | No hardware attestation | Per-request attestation |
| Model choice | Models the team deploys | Provider catalog | Supported TEE catalog |
| GPU operations | Customer responsibility | Provider responsibility | Provider responsibility |
| Audit evidence | Custom | Policy records | Cryptographic evidence |
The comparison shows why a single word such as "secure" isn't enough to describe an LLM deployment. Self-hosting gives direct infrastructure control but also puts host security, serving operations, and hardware management on the customer.
ZDR reduces retention concerns without supplying hardware evidence about execution. TEE-backed inference adds a hardware security boundary and per-request evidence, while still leaving client-side security and application-level handling outside the TEE's scope.
Secure LLM Is a Property of the Deployment
A secure LLM isn't defined by the model weights alone; rather the security boundary depends on where inference runs, who can access plaintext during execution, whether data is retained, how workloads are isolated, and what evidence exists to verify the environment.
Self-hosting gives organizations direct control over infrastructure, but it also makes the host, GPU memory, administrators, monitoring systems, patching process, and capacity planning part of the security model. ZDR removes retention at the provider layer, but it doesn't prove how inference is protected while the model processes a prompt. TEE-backed inference adds hardware-based isolation and cryptographic evidence that can be independently verified.
The practical choice comes down to which trust relationships an organization can accept, operate, and verify. A secure LLM deployment is ultimately a property of the execution environment, not the model itself.
FAQs
Is self-hosting an LLM more secure than using an API?
Self-hosting gives direct control over infrastructure and retention, but the organization also owns host security, privileged access, GPU isolation, logging, patching, and serving operations. The answer therefore depends on the threat model rather than the deployment label.
Are open-source LLMs as good as frontier models?
Performance differences have narrowed considerably, but they haven't disappeared across every model and task. Stanford's March 2026 comparison put the top closed model 3.3% ahead of the top open model.
Does zero data retention make an LLM secure?
ZDR addresses post-inference retention. It doesn't establish who could access plaintext during execution or provide hardware evidence about the environment that processed the request.
Does a TEE protect everything around an LLM?
No, TEE protection applies to the defined execution boundary. Client compromise, leaked credentials, data exposure before submission or after response delivery, and unsafe model behavior remain outside the documented TEE guarantee.