How to Secure AI Agents: Trust Boundaries, Attestation, and Runtime Control

· Updated

ORGN Team

TL;DR

  • If an agent runs inside a shared VM or an ordinary container, its credentials and application policy can still sit atop an execution environment that exposes shared state; sensitive workloads therefore require hardware-enforced isolation before identity or inference controls can be trusted.
  • When model inference happens outside the same verified boundary as the agent, runtime isolation doesn't establish inference privacy; ZDR provides a provider commitment, while TEE inference adds independently verifiable hardware evidence for each request.
  • If parallel agents share filesystem state or execution context, one agent can cross the intended data boundary without exploiting a vulnerability; separate worktrees eliminate that runtime path, but teams must automate their creation and teardown as concurrency grows.
  • An agent inheriting a developer's credentials can exercise permissions far beyond its assigned task, so least privilege requires a separate non-human identity, dedicated credentials, explicit subagent routing rules, and an independent revocation path.
  • When an authorized agent reaches a high-impact tool call, authorization alone doesn't establish operator intent; approval gates are needed for actions such as production writes or external data transfers, with the cost of added latency handled through deployment-specific policy.
  • If an audit contains only platform logs, sandbox attestation, or inference receipts, it leaves a specific part of the trust chain unproven; runtime attestation and per-request inference evidence have to be collected together when the requirement is to verify both where the agent ran and where its model calls ran.

Why Securing AI Agents Differs from Securing Conventional Software

Gartner projects that 40% of enterprise applications will embed task-specific AI agents by the end of 2026, up from under 5% in 2025, a deployment curve that the security tooling available to most teams wasn't built to handle.

The assumption underlying conventional application security is that you can enumerate what a running process does: a known binary, scoped credentials, predictable outputs. AI agents break all three; i.e., an agent reasons, plans, and invokes tools in ways its own developers can't fully enumerate in advance. The agent is simultaneously an identity, a process, and a decision-maker within a single runtime, which means the attack surface spans its reasoning layer, its memory, and every tool it can call. Applying a standard AppSec checklist to an agentic workload doesn't close those vectors. It just maps familiar controls onto a fundamentally different threat model.

The most common failure isn't a novel exploit. It's a category error: teams secure the infrastructure their agents run on (the VM, the API keys, the network perimeter) while leaving the reasoning layer ungoverned. An agent with a valid API key and a manipulated prompt can exfiltrate data, escalate privileges through chained tool calls, or forward sensitive context to an external model endpoint, all inside the perimeter the security team thought was locked. Bessemer Venture Partners documented a controlled red-team exercise in which an autonomous agent compromised McKinsey's internal AI platform in under two hours. This speed makes human response nearly irrelevant. Prompt injection, memory poisoning, and semantic privilege escalation don't trigger conventional SIEM alerts because they operate at the content layer.

This visual shows that agent security isn't one perimeter around the workload. Runtime, inference, identity, and evidence each establish a different security claim.

Article illustration

This article walks through four control categories: runtime isolation, inference privacy, non-human identity governance, and cryptographic audit. Each section names where the failure occurs, what the correct control looks like, and where ORGN's CDE, Gateway, and Scanner close the gap.

Securing the Execution Environment Before the Agent Starts

The execution environment is the precondition for every other control. Get it wrong, and all downstream governance, identity scoping, prompt filtering, and behavioral monitoring rest on a foundation that a compromised host, a shared-memory side channel, or an untrusted hypervisor can undermine.

For most teams, "sandboxing" means Docker containers or VLANs. For agentic workloads where an agent uses tools, edits files, runs terminal commands, and makes live web requests, that's not isolation; it's containment. A containerized agent on shared cloud infrastructure still has a blast radius that extends to mounted volumes, neighboring tenant processes, and every credential the underlying VM can reach. The correct control for sensitive workloads is hardware-enforced memory encryption at the VM level, in which the host OS, the hypervisor, and even the infrastructure provider can't read the memory of running VMs.

Why Runtime Isolation Is the First Control, Not an Optional One

ORGN CDE's cloud worktrees run inside Intel TDX Trust Domains. TDX is a CPU-level technology that creates hardware-isolated virtual machines with encrypted memory, in which the VM's memory contents are inaccessible to the hypervisor and the cloud operator. The practical consequence for security teams is straightforward: an agent's code, terminal state, and tool execution context can't be read from outside the sandbox, regardless of the privileges held by the infrastructure layer above it.

That's a meaningful departure from a policy-based isolation claim. Most cloud providers promise not to inspect your workloads. A TDX Trust Domain makes that promise verifiable, because the isolation is enforced by hardware, not by contract.

The boundary is worth stating precisely; TDX attestation proves the runtime environment is genuine Intel TDX hardware running an untampered measured image. It doesn't prove what the agent did within that environment, which requires inference receipts covered in the next section. Runtime confidentiality and inference confidentiality are two separate controls with two separate proof mechanisms.

Parallel Worktrees and the Blast Radius Problem in Multi-Agent Workflows

Running multiple agents on the same execution surface is a common architecture. It's also a common source of unintended data cross-contamination that doesn't require an attack to trigger; shared filesystem state between two agents is enough.

CDE's worktree model addresses this directly. Each worktree is a separate branch running in its own TDX sandbox, so parallel agents don't share file state, terminal history, or execution context. Agent A, working on a sensitive codebase, and Agent B, handling external API calls, can run concurrently without any data path between them at the runtime layer. Worktrees are created and assigned from the Projects sidebar, one per task, one branch per agent, one sandbox per worktree.

The tradeoff is real: worktree isolation adds provisioning overhead and requires deliberate lifecycle management. Teams running ten or more concurrent agents need to build worktree creation and teardown into their deployment automation rather than handling them manually. The ORGN docs on cloud worktrees cover the SSH attach flow and sandbox recovery paths in detail.

Article illustration

Kali Environments and the Offensive Validation Gap

A deployment review that doesn't include adversarial testing isn't complete. Configuration checks, IAM policy audits, and network scans don't surface prompt-injection vectors or behavior that only emerges with adversarial input. CDE provides per-worktree Kali VMs on cloud worktrees, giving security teams a dedicated attack environment on the same hardware boundary as the agent under test.

The value here isn't just having a Kali instance. It's that the red-team environment runs inside the same confidential compute boundary as the workload being tested, so test data, including simulated sensitive payloads used in injection exercises, doesn't escape the sandbox during the exercise itself. That collapses the gap between "configuration approved for deployment" and "tested adversarially before deployment."

Inference Privacy: What ZDR Doesn't Prove, and TEE Does

Governing where your agent executes is necessary. Governing where your prompts and model outputs are processed is a separate problem, and most teams conflate the two, leaving one unaddressed.

The distinction matters most in regulated environments where a data residency or confidentiality requirement applies specifically to inference, not just to stored data or code. A correctly isolated runtime indicates that the agent's tool calls occurred in a hardware-verified sandbox. It says nothing about whether the model that processed the agent's prompts did the same.

The Policy-vs-Hardware Distinction That Most Gateway Comparisons Skip

Zero Data Retention (ZDR) is the standard enterprise AI gateway offer: a contractual commitment from the inference provider not to store your prompts. It's a policy, not a mechanism. If the provider's infrastructure is compromised, or if the agreement is interpreted differently in a jurisdiction you care about, ZDR gives you no independent recourse. You trusted the vendor's word, and that was the extent of the guarantee.

Article illustration

TEE inference is different in kind: the model executes within a hardware-backed Trust Domain, and the execution environment produces a cryptographic receipt for each request. That receipt can be verified independently, without trusting the vendor as an intermediary. ORGN Gateway's TEE routes, model prefixes near_*, phala_*, and tinfoil_*, produce per-request attestation receipts inspectable in ORGN Scanner. ZDR routes (prefixed vercel_*) appear in Scanner for observability but don't produce hardware attestation receipts.

For teams in defense contracting, regulated finance, or healthcare, this distinction belongs in procurement documentation before a breach forces the conversation.

How Intel TDX and NVIDIA GPU Attestation Bind Together Per Request

The TEE inference architecture on ORGN Gateway combines two distinct isolation layers and links them cryptographically. Intel TDX provides secure virtual machines with hardware-enforced isolation from the host OS, encrypted memory, and verifiable measurements of VM state. NVIDIA GPU attestation extends that boundary to GPU compute within the same session.

The binding mechanism is a shared session nonce. The nonce appears in both the TDX REPORT_DATA field and each GPU's SPDM evidence header, proving that the CPU and GPU attestation artifacts came from the same session and not from separately captured evidence assembled after the fact. That binding means neither half of the proof is sufficient on its own: TDX alone proves the VM was genuine; NVIDIA attestation alone proves the GPU was trusted. Together, they prove that the inference ran end-to-end in a verified environment.

Reading a Scanner Receipt in a Security Review

ORGN Scanner at scanner.orgn.com is publicly accessible without sign-in. Each TEE inference request entry surfaces the model ID, token counts, latency, attestation status, signing addresses, and Intel/NVIDIA evidence artifacts. What Scanner never shows is prompt content or model outputs. The boundary is intentional, because attestation proves where and how inference ran, not what was inferred.

That's the critical framing for security reviews and procurement questionnaires. An attestation receipt doesn't prove your prompts were safe. It proves that the environment in which they were processed was hardware-verified and untampered. Those are two separate claims, and conflating them in a compliance submission creates a gap that auditors will find.

Non-Human Identity Governance for Agents at Scale

Identity is where agentic security most directly intersects with existing enterprise security practice, and where the mismatch is most consequential. Security teams with mature IAM programs can audit every human identity in their environment: OAuth scopes, role assignments, API key ownership, revocation paths. Ask the same questions about their agents, and the answer is often silence.

The silence has a structural cause. Most credential management tooling was built for human identities. It models users, roles, and service accounts. It doesn't natively model agent-specific properties such as ephemeral tool invocation scope, subagent delegation chains, or the fact that an agent's effective permissions at runtime depend on which tools it can invoke, not just on which credentials it holds.

Why Agents Fail the "Who Approved It" Test

Agents frequently inherit the credentials of the developer who configured them. The agent's effective permissions are therefore determined by the most permissive human identity on the team, regardless of what the agent was built to do. An agent designed to summarize meeting notes doesn't need write access to a production database, but if it runs under the credentials of an engineer who does, it has that access.

The fix is treating every agent as a distinct non-human identity with its own credential set, its own least-privilege scope tied to the specific tools the agent needs for its task, and its own revocation path that doesn't require revoking the human developer's credentials to shut down the agent. CyberArk's guidance on agentic identity frames this precisely: every AI agent is an identity, and it needs credentials managed on the same lifecycle as a human identity, provisioned deliberately, rotated on schedule, and revocable instantly.

An agent can inherit the developer's permissions even when its task requires only a small subset of them. This creates a direct path from a legitimate human identity to actions the agent was never meant to perform.

Article illustration

Subagent Routing and the Delegation Trust Problem

CDE's subagent model introduces a governance surface that most teams haven't mapped. When Origin Agent delegates a subtask to a named subagent, oracle for deep reasoning, explore for codebase search, librarian for document retrieval, multimodal-looker for image and PDF processing, each subagent inherits the session model by default, or can be pinned to a specific model in Settings > AI > Agents.

Article illustration

The security implication follows directly: a subagent pinned to a model outside the TEE tier routes inference outside the hardware-verified boundary, even if the parent Origin Agent session is running on a TEE model. Teams deploying agents on export-controlled codebases or sensitive data need an explicit subagent model policy, not just a gateway-level policy, that specifies which subagent types must use TEE routes and which can use ZDR routes without elevating risk. The default "Inherits session model" behavior is convenient for development. It's a governance gap in production if the session model isn't TEE.

Approval Checkpoints Before High-Impact Tool Calls

Least-privilege identity governance addresses what an agent is authorized to access. It doesn't prevent an authorized agent from taking an action the operator didn't intend at a particular moment. That's where approval checkpoints sit: a human-in-the-loop gate before an agent executes an action whose effects can't be trivially reversed.

The category of actions that warrant a gate is deployment-specific. Still, the structural candidates are consistent: file writes to production paths, external API calls that transfer data outside the organization, terminal commands that modify system state in ways that affect other users, and any action that triggers a downstream workflow the agent didn't initiate. ORGN Studio's mission configuration supports approval gates at the task level, allowing high-impact tool calls to be flagged for explicit human confirmation before execution. The operational tradeoff is real: approval gates add latency and create a supervision bottleneck in high-throughput pipelines. Which actions require gates and which can run autonomously are policy decisions, not default settings, and they belong in your agent deployment runbook before the agent is live.

Cryptographic Audit Trails: What Attestation Proves and What It Doesn't

Most agentic platforms offer logs. Logs are produced by the same system they document, which means a compromised or misconfigured system can produce misleading logs without any obvious indicator. Attestation operates on a different trust model: the signed TDX report for a CDE cloud worktree is produced by Intel hardware and is independently verifiable against Intel's PKI, without relying on ORGN's infrastructure as an intermediary for the proof.

That distinction matters for regulated workloads, where the auditor checks whether the operator could have tampered with the audit record. A log entry saying "agent executed in a secure environment" is an operator's claim. A TDX attestation report signed by hardware is verifiable evidence that doesn't require trusting the operator.

The Difference Between Observability and Verifiability

Fetching a sandbox attestation report in CDE takes two steps: click the TDX shield in the status bar, or run Show TDX Sandbox Attestation from the command palette. The report opens in an editor tab and surfaces the sandbox ID, worktree, and project context; the TDX hardware-signed quote; cryptographic measurements of the launch environment; and the issued timestamp. Each field maps to a specific verifiable claim: the TDX quote verifies against Intel PKI; the measurements detect image tampering; the sandbox ID binds the report to the specific worktree.

Attestation is available only when CDE is attached to a cloud worktree over SSH. Local Open Project folders don't produce a TDX sandbox report because there's no hardware Trust Domain to attest. If you request attestation outside a cloud worktree, CDE returns: "TDX attestation is only available inside a CDE remote worktree." That's a hard boundary, not a configuration option.

Sandbox Attestation vs. Inference Receipts: Two Different Audit Questions

These two attestation systems answer different questions, and both belong in a complete audit package. Sandbox attestation covers the environment: was this CDE cloud worktree running on genuine Intel TDX hardware with an untampered measured image? Inference receipts cover the model call: did this specific TEE inference request execute inside a verified Trust Domain?

Article illustration

An audit that collects only sandbox attestation can prove the agent's code ran in a hardware-verified environment. It can't prove that the model calls the agent made were processed on TEE hardware. An audit that collects only inference receipts proves the model calls were TEE-verified but says nothing about whether the execution environment was isolated. Both halves are necessary for a full evidence chain on sensitive agentic workloads. Sandbox attestation surfaces in the CDE status bar and in Scanner's Sandboxes section; inference receipts surface per request at /request/:requestId.

Article illustration

Sandbox attestation is public; no sign-in is required to view the sandbox list or inspect TDX attestation evidence for any sandbox.

Building an Audit Package for Regulated Procurement

Defense contractors, fintechs, and healthcare organizations subject to CMMC, SOC 2, or HIPAA need to demonstrate, before a contract is awarded, that AI inference happened in a verified environment. A minimal ORGN-based audit package contains three components: the TDX sandbox attestation report from CDE that proves the runtime environment, the per-request TEE inference receipts from Scanner that prove the model calls, and the Gateway security model documentation as the threat boundary statement.

Scanner is publicly accessible without sign-in. Auditors can independently verify request records, attestation status, and signing addresses without requiring access to the ORGN platform. That's the point: the verification path doesn't require trusting ORGN as an intermediary, which is precisely what a hardware attestation chain is designed to provide.

When Policy-Based Controls Stop Being Enough for Secure AI Agents

The decision teams in regulated environments face isn't whether to secure their agents. It's at what point policy-based trust, ZDR agreements, access control lists, governance checklists, and configuration reviews become insufficient for the workload they're running, and what the transition to hardware-verified trust requires operationally.

Policy controls are the floor. They're necessary, auditable, and familiar to procurement teams. They don't provide independent verification of where compute ran or what hardware executed the inference. For workloads involving export-controlled code, classified system integrations, or PII at scale, the procurement checklist is the minimum, not the ceiling. The gap between "we have a ZDR agreement" and "we can independently verify every inference request ran in hardware-isolated compute" is the gap that hardware attestation closes. It's the gap that regulators are specifically beginning to ask about.

This article covered four control layers that matter for secure AI agent deployments: hardware-isolated runtime environments via CDE's TDX cloud worktrees; inference privacy through TEE model routing on ORGN Gateway, with the ZDR vs TEE distinction made explicit; non-human identity governance including subagent model policy and approval checkpoints for high-impact tool calls; and cryptographic audit via sandbox attestation reports and per-request TEE inference receipts in ORGN Scanner. The specific operational decisions- which subagents get TEE models, which tool calls require human gates, which audit artifacts belong in a procurement package- are deployment-specific. The architecture that makes those decisions verifiable rather than asserted is not.

FAQs

1. What are the biggest security risks of deploying AI agents in enterprise environments?

Prompt injection, semantic privilege escalation, and memory poisoning are the three vectors that conventional security tooling misses. All three operate at the content layer, not the network layer, which means perimeter controls and IAM policies don't catch them. An agent with valid credentials and a manipulated input can traverse systems at machine speed before an analyst responds.

2. How do you prevent prompt injection attacks in autonomous AI agents?

Prevention requires controls outside the model's reasoning loop. Deterministic policy enforcement, permission checks, scope boundaries, and tool-call validation can't live inside the LLM's context. Credential isolation, where agents authenticate to a gateway that holds credentials rather than directly to backend systems, removes one major path for exploitation. Runtime monitoring that flags anomalous tool invocation patterns catches injection attempts that reach execution.

3. What is the difference between ZDR and TEE for AI inference privacy?

ZDR is a contractual commitment from the provider not to store prompts. TEE inference runs in hardware-isolated compute and produces a per-request cryptographic receipt that you can independently verify against Intel PKI. ZDR requires trusting the vendor's infrastructure. TEE requires trusting the hardware specification, a meaningfully different trust basis for workloads that can't accept vendor-mediated assurances.

4. How should security teams govern non-human identities for AI agents?

Each agent needs its own credential set with a least-privilege scope tied to its specific task, rather than inherited from the developer who configured it. That means a dedicated provisioning and revocation path that doesn't require revoking human credentials to shut down an agent. In multi-agent systems, subagent delegation chains require the same treatment: each subagent's model routing and tool access should be explicitly scoped, rather than inherited by default from the parent session.