Confidential AI: What It Means and Why It Requires Hardware, Not Just Policy
· Updated
ORGN Team
TL;DR
- Confidential AI protects data while it's actively being processed, not just while it's stored or in transit, a gap standard encryption never covered.
- It requires hardware-enforced isolation via Trusted Execution Environments (TEEs), in which memory remains encrypted even from the host OS and the cloud operator.
- A ZDR (zero-data-retention) policy and a TEE solve different problems: one is a contractual promise, the other is a cryptographic guarantee checkable against a public PKI.
- Cryptographic attestation turns "we run inside a TEE" from a claim into a proof that is verifiable by anyone, not just the vendor's own customers.
- ORGN implements this through a CDE-layer TDX sandbox isolating every workspace by default, plus a Gateway offering explicit TEE and ZDR model tiers with per-request attestation evidence for TEE-routed calls.
Why "Encrypted" Was Never the Same as "Confidential"
Encryption at rest and encryption in transit have been baseline security expectations for well over a decade. Disk encryption protects stored data. TLS protects data moving between systems. Both are mature, well-understood, and, for most workloads, sufficient.
AI workloads expose a gap neither of those controls was built to close: data in use. To generate a response, a model needs the prompt, context, and intermediate reasoning state sitting in plaintext memory. That's true regardless of how well-encrypted the surrounding infrastructure is. During that processing window, the data is visible to the host operating system, the hypervisor, privileged administrators, and any logging or debugging tooling running on that infrastructure. Confidential AI means closing that window and protecting data during execution, not just before and after.
What Confidential AI Actually Requires
Confidential AI isn't a single feature. It's a small set of specific technical properties working together, and it's worth being precise about each one, because vendors frequently use the term loosely.
Hardware-enforced isolation. The core requirement is a Trusted Execution Environment (TEE), a hardware-isolated execution boundary where memory is encrypted and inaccessible to anything outside the enclave, including the host operating system and the cloud operator. This is a categorically different guarantee than a contractual promise not to log data. A policy can be violated by misconfiguration, a compromised credential, or a legal compulsion order. Hardware isolation doesn't depend on anyone's good behavior; it's enforced at the CPU level.
Cryptographic attestation. Isolation alone isn't verifiable without a receipt. Attestation is what turns "we run inside a TEE" from a claim into a proof: a signed report identifying the exact software running, the hardware platform it ran on, and its tamper state, checkable against a public root of trust, Intel's or NVIDIA's PKI, by anyone, not just the vendor's own customers.
Scoped, non-persistent data handling. Confidential AI implementations should minimize what they retain after a session ends. This doesn't mean zero operational data exists anywhere; billing and reliability require some metadata, but it does mean prompt and response content shouldn't sit in a retrievable database after the fact.
The distinction that trips up most evaluations is that confidential compute (memory encryption, hardware isolation) and confidential policy (a provider's contractual commitment) solve different problems and produce different kinds of evidence. A vendor answering "yes" to "do you offer confidential AI" might mean either "yes" or "no."
How Hardware Isolation and Attestation Work Together
An Intel TDX Trust Domain is a confidential virtual machine in which the entire guest memory is encrypted at the hardware level and is invisible to the host OS and hypervisor. NVIDIA GPU Attestation extends that same isolation model to accelerator hardware, verifying that the GPU state used during inference matches an expected, untampered configuration. Used together, they cover both halves of a typical inference workload: the CPU-bound orchestration and the GPU-bound model execution.

The attestation report generated by this process is what makes the isolation checkable rather than assumed. It's worth being precise about what it does and doesn't prove: attestation confirms the execution environment's integrity, that specific code ran inside a specific, unmodified enclave. It says nothing about the correctness of the model's output, or about data handling once a response leaves the TEE boundary. Those remain separate concerns that confidential compute doesn't resolve on its own.
| Traditional Cloud Inference | Confidential AI (TEE-backed) |
|---|---|
| Prompt decrypted in host-visible memory | Prompt decrypted only inside the encrypted enclave |
| Isolation guaranteed by policy and access controls | Isolation enforced by hardware, independent of operator behavior |
| Trust based on vendor claims and certifications | Trust based on a cryptographic report checkable against a public PKI |
| No per-request proof of execution environment | Signed attestation available per TEE-routed request |
How ORGN Implements Confidential AI
ORGN applies confidential compute at two distinct layers, and keeping them separate matters for understanding exactly what evidence exists for any given piece of work.
The CDE layer isolates the workspace. Every cloud worktree in ORGN runs inside a TDX-encrypted sandbox by default; this is a session-level guarantee that applies regardless of which model is used for inference. Tool calls, file edits, and terminal commands all execute inside this TDX-backed sandbox, isolated from the host infrastructure and inaccessible to ORGN operators.
The Gateway layer determines inference-specific attestation. Model access routes through two distinct tiers, and the tier selected for a given request determines whether that specific inference call produces additional, per-request attestation evidence on top of the baseline CDE sandbox isolation.
curl -X POST https://api.gateway.orgn.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "phala_deepseek_v3_1",
"messages": [
{ "role": "user", "content": "Summarize this incident report." }
]
}'Selecting a TEE model string, phala_deepseek_v3_1 or near_qwen3_30b, routes that inference call to Phala or NEAR's TDX infrastructure and produces a per-request attestation receipt in Scanner, ORGN's public attestation explorer. The exact fields in that receipt vary by provider; Phala and NEAR don't return identical evidence artifacts. Selecting a ZDR model string instead, a vercel_* alias, still executes inside the same CDE-layer TDX sandbox at the workspace level, but that specific inference call produces no hardware attestation receipt. The ZDR tier's guarantee is contractual: the provider has agreed not to retain or train on the data, but that's a policy commitment, not a cryptographic one.

Model selection is always explicit at the application layer: a developer specifies the model string directly, and there's no dynamic routing deciding which tier handles a request behind the scenes. That predictability is what makes the evidence trail traceable after the fact: the model string used is a reliable indicator of what proof exists for that call.
What Data ORGN Actually Retains
Confidential AI claims are only as good as what actually happens to data after a session ends, so it's worth being specific rather than relying on general language like "zero retention."
ORGN does not store prompt or response content. This isn't a configurable default with an opt-in or opt-out; it's a fixed property of the infrastructure. What ORGN retains is limited to: the user's email address (via Privy authentication), user-agent data, token usage for billing, and TEE attestation records, including the underlying provider quotes. None of that retained data includes the actual content processed.
For requests routed through TEE models specifically, no prompt, code, or output is ever used to train a model. That guarantee is tied to the TEE inference path; it's a structural property of how that infrastructure is built, not a setting a team enables separately.
Extending Isolation Beyond a Single Inference Call
Most discussions of confidential AI focus narrowly on the model call itself: is the prompt encrypted during inference? That framing misses a real surface of exposure: everything that happens before and after that single call. Preprocessing, retrieval steps, intermediate transformations, and business logic that touches sensitive data often run outside any protected boundary, even when the final model call is secured.
ORGN addresses this by making full-workspace confidential compute a selectable option, not something bolted onto the model layer alone. When enabled for a project, the entire development environment, not just inference, runs inside the hardware-encrypted TDX boundary. That matters for workflows like retrieval-augmented generation, where the retrieved documents themselves may be as sensitive as the final prompt sent to a model, or fraud detection pipelines where filtering logic touches high-value transaction data before any model is called at all.
Scaling Confidential AI Usage
Getting started with confidential AI shouldn't require negotiating an enterprise contract upfront. ORGN's self-serve tier runs on prepaid credits starting at $25, with no subscription required, and covers both standard and TEE-backed inference through the same Gateway.
For teams anticipating heavier usage or approaching rate limits, ORGN's Enterprise tier includes reserved capacity arrangements available through direct sales conversations, along with air-gapped and private deployment options for teams in classified or highly restricted environments where standard cloud deployment isn't an option regardless of the isolation guarantees underneath it.
Conclusion
Confidential AI is a specific, checkable technical property, not a marketing label for "we take security seriously." It requires hardware-enforced isolation during execution, cryptographic attestation that makes that isolation independently verifiable, and data handling that's minimal by design rather than minimal by policy toggle. Most AI infrastructure today satisfies none of these at the technical level; it offers encrypted transport and a compliance certificate, which answers a different, easier question than the one a security review actually asks.
If your team needs to prove, not just claim, where inference ran and what happened to the data while it did, get started with ORGN and see what a per-request cryptographic attestation record looks like in practice.
FAQs
What makes AI processing "confidential" instead of just encrypted?
Standard encryption protects data at rest and in transit but not while it's being actively processed; a model still needs plaintext access to a prompt in memory to generate a response. Confidential AI specifically closes that gap by using hardware-isolated Trusted Execution Environments to keep data encrypted even during execution, with cryptographic attestation available to prove that isolation was maintained.
Is a zero-data-retention policy the same as confidential computing?
No. A zero-data-retention (ZDR) agreement is a contractual commitment: a provider has agreed not to log or train on data. Confidential computing is a hardware guarantee, enforced at the CPU and memory levels regardless of any provider's policy. Both reduce risk, but only hardware isolation provides evidence independently verifiable against a public root of trust, rather than a promise you must trust.
How does attestation prove that confidential AI infrastructure actually worked as claimed?
Attestation is a cryptographically signed report generated by the hardware itself that identifies the exact code running, the platform it runs on, and its tamper state. That report can be checked against Intel's or NVIDIA's public PKI by anyone, an internal security team, an external auditor, or a third party with no account on the platform that generated it, without needing to trust the vendor's own infrastructure logs.
Does using a confidential AI platform mean every request is hardware-isolated?
Not necessarily, and this is a common point of confusion. Platforms that offer multiple model tiers typically apply hardware isolation per request, based on which specific model was selected; a standard or ZDR-routed model call through the same platform doesn't inherit TEE-level guarantees just because the platform also offers TEE options elsewhere in its catalog. Checking which model string was actually used is the only reliable way to know what evidence exists for a specific request.