TL;DR
- When an AI coding assistant turns source files, terminal history, and session context into a prompt, traditional DLP can miss the transfer because the sensitive data leaves as unstructured API content rather than a classified file.
- If sensitive code remains on an unmanaged developer endpoint, protecting the inference provider does not address exposure while the assistant constructs context; runtime isolation must begin before the prompt is created.
- ZDR prevents providers from retaining or training on prompts, but the data still exists in provider infrastructure during inference, so ZDR alone cannot satisfy reviews that require hardware-backed proof of execution.
- When a workflow requires independently verifiable evidence, TEE inference provides cryptographic receipts for specific model calls without retaining prompt or response content, trading some model-catalog breadth for stronger assurance.
- Runtime and inference attestations solve different audit questions: one proves where the code and agent executed, while the other proves where the model call ran, so regulated workflows requiring hardware-backed evidence need both.
- Select model security per workload rather than globally: ZDR fits non-sensitive work where policy-level retention controls suffice, while sensitive or regulated context should route through a TEE when the review requires cryptographic proof.
Where Sensitive Data Enters the LLM Pipeline
The Samsung incident from 2023 still gets cited as the reference case for developer-driven LLM exposure, but it understates what broke. Engineers didn't bypass the control, they used an AI coding tool exactly as designed, and the source code was transmitted to a third-party inference endpoint. The architecture worked, and that's the real problem.
LLM data security occupies territory that traditional application security wasn't built to cover. Access controls protect files at rest. Network controls protect data moving between known endpoints. Neither intercepts what happens when a developer's AI assistant constructs a context window from a local codebase and ships it as unstructured text to an inference API. Prompts have no file-classification label, and there's no discrete transfer event. The data moves as semantic content inside a standard HTTPS call, and most DLP stacks can't see it.
Teamblind thread asking whether companies restrict LLM tool access drew thousands of responses from engineers at Walmart, Snapchat, and others describing active bans, code-paste restrictions, and internal-only model deployments, all driven by the same underlying concern: no one knows exactly what leaves the boundary when a developer types into a prompt.

For regulated teams, including defense contractors, financial institutions, and healthcare infrastructure engineers, Sixty-Two Percent of Organizations Experienced a Deepfake Attack; 32% Faced an Attack on AI Applications, according to Gartner’s Report. The governance question has already shifted from whether AI is in use to whether security teams can control what leaves the boundary when it is. Most can't answer that second question with evidence.
This article covers where LLM data security breaks down in developer workflows: the inference layer, the execution runtime before a prompt is even sent, and the audit record that needs to exist after the fact. Using ORGN's CDE and Gateway as the concrete implementation, it shows what hardware-enforced isolation looks like at each layer and what the distinction between ZDR and TEE attestation means when a procurement reviewer asks for proof.
The Code-as-Context Problem: Why Developer Workflows Break Existing Controls
Source code in an LLM context window isn't governed by the same controls that protect the file it came from. The moment a developer's AI assistant reads a working directory to build context, that code stops being a file under access control and becomes a prompt payload bound for an inference endpoint the developer's organization doesn't operate.
Source Code in the Context Window Is Structurally Different from a File at Rest
File-level DLP tools check classification labels, destination addresses, and transfer protocols. None of those checks apply to prompt construction. The assistant reads local files silently, aggregates context across open tabs and terminal history, and sends that aggregate to an external endpoint as part of a standard API call. A developer debugging a configuration error might include environment variables in the prompt. One summarizing a database schema might include field names that appear in no other transfer log. Each step looks innocuous; together, they expose the full context of a system the organization treats as sensitive, with no entry in any audit log to let a security team reconstruct what left.
ORGN CDE addresses this at the execution layer, i.e., Cloud worktrees run inside Intel TDX Trust Domains, hardware-encrypted VM memory that the host OS, hypervisor, and ORGN operators can't inspect. Code doesn't sit on an unmanaged local endpoint where a background assistant reads it without restriction. It runs in a controlled, attestable environment. That attestation report, verifiable against Intel's public PKI, confirms the sandbox is genuine TDX hardware running an untampered image before any inference call.
Prompt Construction Is an Unmonitored Exfiltration Path
OWASP's Top 10 for LLM Applications (2025) places sensitive information disclosure at LLM02, the second-highest risk category, because output filtering must work against a model that doesn't distinguish between confidential and public content during generation. The model doesn't know that the schema file a developer included to get SQL suggestions contains column names from a regulated dataset. It processes the file and returns output that may contain details the developer never intended to expose.
Multi-turn sessions make this harder to catch, i.e., in a typical AI coding session, context accumulates: an API schema in turn one, internal endpoint names in turn three, authentication logic in turn five. No single message is obviously sensitive; the aggregate is a detailed map of a system the organization hasn't approved for external exposure. Standard gateway observability logs request counts and latency. It doesn't parse content. ORGN Gateway addresses retention: prompts and responses aren't stored or logged at any point in the pipeline, but zero retention and zero proof aren't the same, and that distinction matters when an incident responder asks what left the boundary.
A developer session rarely exposes one isolated piece of information. The assistant can pull source files, terminal history, environment variables, and open tabs into the same context window, turning several local inputs into a single outbound inference payload.
Look first at where those inputs converge inside the context window. The easiest part to miss is the boundary crossing itself, because the outbound request contains semantic context rather than a recognizable file transfer.

That flow explains why conventional DLP can miss the exposure: the controls see neither the individual context sources nor a discrete file-transfer event. By the time the request reaches the inference API, the sensitive information has already been transformed into unstructured prompt content.
What "We Don't Store Prompts" Means for a Security Reviewer
When a security team asks what data left their boundary during a specific incident window, a ZDR agreement produces no answer. Zero retention by design means no log, no cryptographic artifact, no audit record tied to a specific request. For teams that need to demonstrate to auditors, procurement boards, or incident responders that sensitive data was handled correctly, a policy promise is insufficient.
TEE-backed inference closes this through ORGN Scanner: each TEE request produces a cryptographic attestation receipt, independently verifiable against Intel and NVIDIA public PKI. Scanner retains attestation status (Verified, Pending, Failed), cryptographic hashes, and signatures, but not the underlying prompt or response content. An engineering team can hand that receipt to a security reviewer and point at a specific inference run. The reviewer can verify the execution environment independently, against public infrastructure, without trusting ORGN's word.
What ZDR Agreements Cover and the Gap They Leave During Inference
Zero-retention agreements are real security controls. Understanding where they stop being sufficient is the specific architectural decision that determines what a regulated team can prove when asked.
Policy Trust vs. Hardware Proof: Two Different Assurance Mechanisms
ZDR models, backed by contractual zero-retention agreements with providers, give developers access to the broadest frontier model catalog while establishing that prompts won't be stored, logged, or used for training. The security guarantee is real and contractually binding. The assurance mechanism is a policy document, not a hardware boundary.
During ZDR inference, the prompt exists as plaintext in the provider's infrastructure memory. The provider doesn't store it, but the execution environment is standard cloud infrastructure: shared hardware, standard hypervisor, no hardware-enforced memory encryption isolating the inference computation. An infrastructure-level misconfiguration or insider access at the provider layer sits outside what the ZDR agreement covers. That's not a flaw in the ZDR model. It's the design limit the guarantee is built around, and it's the reason ORGN Gateway makes model tier selection explicit rather than abstracting it away.
ORGN Gateway exposes this distinction at the model ID level. ZDR routes use vercel_* model IDs; TEE routes use near_* or phala_*. Same API key, same wire format, different security posture per request. The following call routes through a TEE-backed model. The near_* prefix signals that ORGN routes this to a NEAR-backed confidential VM, and a cryptographic attestation receipt is produced in Scanner on completion:
from openai import OpenAI
client = OpenAI(
base_url="https://api.gateway.orgn.com/v1",
api_key="sk-ollm-your-api-key"
)
response = client.chat.completions.create(
model="near_llama3-70b", # TEE route: hardware-isolated, receipt in Scanner
messages=[{"role": "user", "content": "Review this service mesh config..."}]
)Switching to a ZDR model means changing the model ID to vercel_gpt-4o. The API surface doesn't change, but the security posture does.
Where Intel TDX Closes the Inference-Layer Exposure
TEE inference on ORGN Gateway routes prompts through Intel TDX confidential virtual machines with NVIDIA H100 GPU attestation, or AMD SEV-SNP with confidential-compute GPU attestation on Tinfoil infrastructure. Inside a TDX Trust Domain, prompts and responses are encrypted in memory. The host OS can't read them; the cloud provider's hypervisor can't read them; ORGN operators can't read them. Every request produces a cryptographic receipt tied to that specific inference run, verifiable against public vendor PKI without going through ORGN.
That verification changes the procurement conversation, which means instead of presenting a vendor's data handling agreement to a review board, an engineering team presents a cryptographic artifact: this inference call ran inside a hardware-isolated environment on this date, verified against Intel and NVIDIA public PKI. A security reviewer checks that independently. ORGN Gateway's published threat model draws this responsibility line precisely: Gateway doesn't address compromised client environments, API key leakage, malicious prompts, or data exposure before requests are sent or after responses are received. Publishing that boundary lets security teams reason clearly about where their controls need to extend, rather than assuming platform coverage where none exists.
Choosing the Right Tier Per Workload in CDE
In ORGN CDE, model tier selection happens in the model picker inside the IDE, a per-session decision, not a platform-wide setting. Origin Agent routes all inference through ORGN Gateway. A developer selects a ZDR model for general-purpose assistance on non-sensitive work and switches to a TEE model when the context includes code or data that a procurement process would classify as sensitive.
The tradeoff is catalog breadth: TEE models cover a narrower range than ZDR's frontier options because hardware-isolated inference requires specific infrastructure. For most regulated workflows, workload-specific routing is the practical answer. Non-sensitive queries run ZDR; code adjacent to classified systems or regulated data routes to TEE. CDE's separation of runtime isolation (cloud worktrees in Intel TDX sandboxes) from inference tier selection (model picker per session) lets a developer maintain both guarantees in a single working environment without switching tools.
The model decision should therefore happen at the workload level, not as a permanent platform setting. Start with context sensitivity, then determine whether the workflow also requires independently verifiable execution evidence.
The easiest branch to misread is sensitive context by itself. Sensitive data does not automatically determine the model tier; the need for independent proof pushes the workflow from ZDR to TEE.

This routing preserves ZDR's broader model access for workloads that only require retention controls while reserving TEE for workloads where the security review requires hardware-backed evidence. The developer can therefore change the inference tier for a session without changing the surrounding development environment.
Runtime Isolation: The Security Gap Before the Prompt Is Even Sent
Inference security only addresses what happens to data after it reaches a model. The execution environment where code lives before a prompt is constructed is a separate exposure surface that most LLM security discussions don't reach.
Local Development Endpoints as Persistent Exposure Surfaces
A developer's local machine running an AI coding tool is an unmanaged endpoint. The AI assistant reads the working directory to build context, sweeps terminal history for relevant commands, and accumulates detail across an entire working session. Local chat logs have no retention policy. Credentials appearing in environment variables get included in context without any classification check. The exposure isn't a single prompt. It's everything the assistant can read across the session.
ORGN CDE's cloud worktree model relocates execution away from the local endpoint. A developer opens an active worktree in CDE, attaches via SSH, and the file system, terminal, and agent tool execution all run inside an Intel TDX sandbox, encrypted VM memory on cloud infrastructure that the cloud provider can't inspect. Cloned repositories don't persist on local storage, with no retention policy. The working environment is controlled and attested from the moment development starts.
Sandbox Attestation in CDE: Proving the Runtime Before Inference Runs
ORGN CDE provides two distinct attestation mechanisms that security reviews require separately, and conflating them produces an incomplete audit package.
Sandbox attestation, available from the TDX shield in CDE's status bar or via Show TDX Sandbox Attestation in the command palette, produces a signed TDX report proving the cloud worktree is running on genuine Intel TDX hardware with an untampered measured image. The attestation report includes the sandbox ID, worktree and project context binding, a TDX quote (hardware-signed evidence from Intel), measured launch environment digests, and the report timestamp. This answers one question: is this execution environment what it claims to be?
Inference attestation, surfaced through ORGN Scanner, answers a different question: did this specific model call run inside a TEE? Both are required for a complete compliance package. A security reviewer examining a regulated development workflow needs attestation covering the runtime and the inference. One specific constraint teams building procurement documentation should note: sandbox attestation is only available after CDE attaches to a cloud worktree via SSH. Local Open Project folders don't produce TDX reports. The attestation surface exists only once execution moves to the confidential cloud environment.
Read the two attestations as separate checkpoints in the same workflow. Start with the environment where the code and agent execute, then follow the request into the inference boundary.
The easiest distinction to miss is that a verified CDE sandbox does not verify the model call itself. Runtime attestation and inference attestation produce different evidence for different execution boundaries.

The separation matters during a security review because proving the runtime does not prove the downstream inference execution. A complete hardware-backed audit therefore needs the sandbox report for the development environment and the inference receipt for the specific model request.
Parallel Worktrees and Agent Isolation Across Team Workflows
Agentic development introduces an isolation problem that goes beyond single-session security. When multiple agents run in parallel, each with file access, terminal execution, and tool use, the question isn't only whether one agent's inference is confidential. It's whether concurrent agents contaminate each other's execution context.
ORGN CDE's parallel worktrees give each agent a separate branch and an isolated sandbox. The isolation is hardware-enforced, not policy-controlled: distinct TDX Trust Domains for distinct execution contexts. For regulated teams running automated security reviews, parallel code generation across different classification levels, or multi-step analysis pipelines that shouldn't share intermediate state, that boundary matters. One agent's context doesn't bleed into another's regardless of what's running in parallel.
Building an Audit Trail That Satisfies a Security Review
The gap between what AI providers log and what regulated auditors require isn't a compliance edge case. It's the specific problem that makes most LLM pipelines unsuitable for regulated environments without architectural changes.
What Regulated Environments Require That Standard Gateways Can't Produce
When a procurement board or incident responder asks about LLM data handling, the question is specific: can you show, with verifiable evidence, what data was processed, where it was processed, and who had access to it? Standard gateway observability covers request counts, latency, model names, and error rates. That log doesn't include the execution environment, verification status, or content-level proof of correct data handling.
A complete audit record for a regulated workflow requires content-free operational metadata (model, provider, token counts, latency, timestamps) paired with environment verification at both layers: sandbox attestation proving the runtime was genuine TDX hardware, and a TEE receipt from Scanner proving the inference call ran inside that hardware. ZDR routes substitute policy documentation for the TEE receipt, but the audit package should explicitly mark that substitution as policy-grade, not hardware-grade. Without all three, a security reviewer has an incomplete picture. That gap is exactly what auditors flag.
ORGN Gateway's logging behavior covers operational metadata and attestation status without ever logging prompt or response content. Scanner surfaces attestation status (Verified, Pending, Failed), cryptographic hashes, and signatures for TEE requests. None of that lets a reviewer reconstruct the request. That separation is what makes Scanner usable as an audit artifact without creating a secondary exposure risk.
Scanner as the Audit Layer: What It Shows and What It Won't
ORGN Scanner retains cryptographic attestation receipts for TEE inference requests, independently verifiable against Intel and NVIDIA public PKI. Attestation status, hashes, and signatures are visible. Prompt and response content are not. Anyone with Scanner access can confirm that a specific inference run happened, ran on identified infrastructure, and was hardware-verified, without reading the request.
A Scanner report showing "Verified, TEE execution on NEAR infrastructure, model near_llama3-70b, timestamp verified against Intel PKI" satisfies a compliance review. It doesn't create a secondary sensitive document. That's a design choice, documented explicitly in ORGN's Gateway security model: the separation enables transparency and auditability without compromising confidentiality.
The Shared Responsibility Line ORGN Draws Explicitly
Most AI providers don't publish a formal threat model that names what they don't protect against. ORGN Gateway does. Gateway doesn't address compromised client environments, API key leakage, malicious prompts, or data exposure before requests are sent or after responses are received. Publishing that boundary isn't a disclosure of weakness. It lets security teams scope their own controls without guessing where the platform stops.
If a platform claims to solve for everything, a security team can't reason about where their responsibility begins. For regulated teams, that explicit scope matters as much as the protection itself. Knowing that Gateway covers the inference pipeline from request receipt through response delivery and attestation, and that endpoint security, API key management, and prompt hygiene remain the organization's responsibility, is the information needed to build controls that don't overlap with the platform and don't leave gaps between them.
LLM Data Security When the Stakes Require Proof, Not Policy
The architectural decision in LLM data security for regulated environments reduces to one question: does your security review require a cryptographic artifact or a contractual document? ZDR agreements fit the first tier of regulated requirements: strong retention controls, broad model access, contractually binding. TEE attestation fits the second tier, where the reviewer asks, "can you prove it" rather than "did you agree to it." Those aren't the same tier, and a security posture built for the first won't satisfy the second.
CDE addresses the fact that the two requirements aren't always separable. Runtime isolation and inference attestation need to cover different surfaces: where code ran versus where inference ran. Simply, a team that addresses one without the other hands an auditor exactly the gap they'll look for. The decision is about knowing which tier a specific workflow requires before the procurement review and not during it.
FAQs
1. What is LLM data security and how does it differ from traditional data security?
LLM data security covers the exposure surface created when language models process input: prompt content, context windows, and completions that transit inference pipelines as unstructured text. Traditional DLP controls check file classifications and transfer destinations. Neither applies to prompt construction, where sensitive data moves as an API payload rather than a labeled file, leaving standard controls with nothing to intercept.
2. Can LLMs leak sensitive data even when using zero data retention providers?
Yes. ZDR agreements prevent providers from storing or training on prompts, but during inference the data exists as plaintext in the provider's infrastructure memory. A policy agreement doesn't create a hardware boundary. Infrastructure-level misconfigurations or insider access at the provider layer sit outside what ZDR covers. TEE execution closes that gap by encrypting prompts in hardware-isolated memory the provider can't inspect.
3. How do you prevent source code from being exposed when using AI coding assistants?
Run code in a hardware-isolated execution environment rather than on a local endpoint where the AI assistant can freely read the working directory. ORGN CDE's cloud worktrees operate inside Intel TDX Trust Domains: the file system and terminal run in encrypted VM memory the cloud provider can't read. Pair that with TEE model routing in ORGN Gateway for inference-layer isolation that produces a cryptographic receipt, not just a policy commitment.
4. What does cryptographic attestation prove about how an LLM handled my data, and what does it not cover?
Attestation proves a specific inference call ran inside a hardware-isolated Trust Domain, verified against public PKI. It doesn't prove what the model generated, and it doesn't cover execution paths outside the attested environment: local development, ZDR routes, and anything that happened before the request reached the TEE. Sandbox attestation covers the runtime; inference receipts from Scanner cover the model call. A complete audit package for a regulated review needs both.