Agentic AI Security: What Your Execution Environment Must Prove

· Updated

ORGN Team

TL;DR

  • Prompt injection is an identity exploit, i.e., an attacker embedding instructions in a document the agent processes assumes system-prompt authority without touching a credential; the fix is cryptographic trust at the hardware level, not a better prompt.
  • Static least-privilege fails agents because they don't hold a fixed role across a session. A read-only research task and a deployment step in the same session means write permissions exist from minute one, and OWASP's Agentic Top 10 names the exact condition where effective permissions silently expand to match the agent's provisioned set.
  • Hardware attestation and ZDR answer different questions. ZDR proves a vendor won't retain data; a TDX report proves the environment was genuine hardware running an untampered image, verifiable against Intel's PKI by anyone, without vendor involvement.
  • Parallel agents on shared infrastructure collapse audit attribution before any incident occurs. Two agents on the same VM produce one undifferentiated log; one agent per isolated TDX sandbox is an architecture decision, not a forensic fix.
  • Shipping OpenTelemetry traces to an external SaaS recreates the exact data exposure the TDX sandbox was built to prevent. Execution context in a trace carries the same sensitivity classification as the code it describes.
  • Cryptographic compliance evidence is only available if the execution environment was built to generate it. Sandbox attestation, inference receipts, and in-boundary observability can't be retrofitted; the sessions that already ran on shared infrastructure produced no verifiable artifacts.

What Is Agentic AI Security?

Agentic AI describes AI systems that do more than respond to a single prompt. They plan sequences of steps, call external tools, store results in memory, spawn subagents to handle subtasks, and execute those steps without human approval at each step. The difference from a standard LLM API call is not the model; it's the execution loop: an agent takes an action, observes the result, decides what to do next, and repeats until it judges the task complete or hits a defined stopping condition.

The security model that covers a standard LLM integration, where you send a prompt and receive text, breaks down when the system starts executing tool calls. A model that produces harmful text causes content problems. An agent that executes tool calls causes operational problems: it writes to databases, calls payment APIs, commits code to production branches, and delegates work to other agents that inherit its permission scope. The harm surface changes from what it says to what it does, and what it does is often irreversible.

Agentic AI security is the discipline of controlling what agents are permitted to do, in what environment, under what identity, with what auditability. It sits at the intersection of IAM, confidential computing, and software supply chain security, and none of those three fields alone covers it. The article that follows works through why each layer requires independent, verifiable evidence rather than a single policy promise covering all three.

When Agent Autonomy Becomes an Attack Surface

Between July 20 and July 24, 2026, Claude Code shipped four consecutive releases, each closing a different sandbox isolation boundary: symlinked working directories, git redirection via GIT_DIR, leftover worktrees from adjacent projects, and unprompted network egress to non-allowlisted hosts. Four separate isolation primitives failed independently, in five days, in one product. That's the practical shape of the agentic AI security problem.

Agentic AI systems plan multi-step tasks, call external tools, delegate to subagents, and act on results without continuous human review. The security profile differs from prompt-based AI in a concrete way: a model that generates text can produce harmful output; an agent that executes tool calls can modify a production database, push to a live repo, or exfiltrate data through an API it's legitimately authorized to use. OWASP's Top 10 for Agentic Applications names goal hijacking, tool misuse, identity abuse, supply chain compromise, and rogue agents as the leading risks, and each happens at the execution layer, not the inference layer.

The gap most teams leave unaddressed is that they secure what the model is told (system prompts, guardrails) and what it retains (ZDR agreements), but not the environment where the agent runs, and they don't produce verifiable evidence of what that environment looked like when an auditor comes asking. This article works through the identity layer, the execution-environment proof layer, and the inference-attestation layer, mapping each to a verifiable control rather than a policy claim.

The Identity Problem Agentic Systems Inherited from IAM

An AI agent calling an API looks identical to a service account calling the same API. The access control layer sees a valid token hitting an authorized endpoint. What it can't see is that the agent chose that call based on runtime reasoning over external inputs it processed earlier in the session, some of which may have been crafted to redirect its behavior.

Why Non-Human Identity Breaks Conventional Access Controls

Legacy IAM frameworks were built for human users with predictable session patterns: authenticate, operate within a defined scope, close the session. An agent running inside a cloud worktree doesn't follow that pattern. It spawns subagents that inherit its permission scope. It processes external documents that contain embedded instructions. It chains tool calls across multiple systems in an order no engineer explicitly approved. A Cloud Security Alliance survey found that 68% of organizations can't reliably distinguish AI agent activity from human activity, which means they can't apply different controls to each and can't trace an anomalous action back to the specific agent session that generated it.

The blast radius of a compromised agent scales directly with its provisioned permission set. A compromised human account is bounded by human speed. An agent operates at machine speed across multiple systems, and the damage compounds before any alerting rule fires.

Prompt Injection as an Identity Exploit

The first part shows how untrusted external content can influence an agent while conventional IAM still sees a valid identity and authorizes the resulting tool call. Next, it shows the safer flow, where the agent checks actions against attested identity and policy before reaching the tool, rather than trusting instructions embedded in external content.

Article illustration

Prompt injection doesn't target the model's output filters. It targets the agent's trust model. When an agent processes external content and treats embedded instructions as trusted input, the attacker assumes the identity of whoever wrote the original system prompt without touching a credential. The EchoLeak exploit (CVE-2025-32711) against Microsoft Copilot demonstrated this at production scale: a crafted email triggered data exfiltration without user interaction, no authentication bypass, no code execution. The agent did exactly what it was built to do, following instructions it was built to follow, from a source it wasn't built to distrust.

More defensible system prompting isn't the fix; trust needs to move out of the message layer entirely. An agent that claims elevated permissions inside a chat message should receive no elevation; trust should be established cryptographically, verified against attested hardware identity, not inferred from string content. That architecture closes the multi-agent injection surface where a compromised subagent passes an override instruction to a downstream agent in the orchestration chain.

Least Privilege at Runtime vs. Least Privilege at Provisioning Time

Article illustration

A permission set applied at agent configuration time and left fixed for the session is a static control on a system that changes what it's doing every few seconds. An agent handling read-only research and then a deployment step in the same session holds write permissions it doesn't need for the first task from the moment it starts. The OWASP Agentic Top 10 identifies the specific condition where the user's effective permissions expand to match the agent's provisioned permissions, not because anyone granted elevated access, but because the agent traverses paths no human would have been allowed to combine in one session.

Runtime-scoped access means evaluating permissions per tool call, not per session. That requires an execution environment where each call is observable at the right granularity, checkable against a current policy, and logged in a way that records what was evaluated, not just what was called.

What the Execution Environment Must Prove

Before hardening anything, a team needs to answer a question most haven't formally asked: what does the execution environment need to demonstrate, to whom, and in what form? The answer varies by regulatory context, but three claims appear in every regulated agentic deployment. The environment was genuine hardware. The runtime image wasn't modified. The specific workload active at the time of the audit can be identified and attributed.

Hardware Attestation vs. Policy Trust: The Difference for Regulated Teams

A ZDR agreement says a vendor won't retain your data. It says nothing about the execution environment where your agent ran, and it produces nothing a procurement team can check independently. Hardware attestation produces a different kind of artifact: a signed report from the sandbox that anyone with Intel's public keys can verify without involving the vendor.

Intel TDX (Trust Domain Extensions) runs code inside an isolated VM with encrypted memory. The sandbox generates a cryptographic quote containing a hardware signature, the TDX module version, and measurement digests of the launch environment. Anyone holding the report can verify the hardware is genuine, and the image is untampered; no vendor login required, no trust-me-on-this. ORGN CDE's sandbox attestation produces exactly this report for cloud worktrees: a TDX attestation document bound to a specific worktree's sandbox ID, fetchable from the IDE status bar, verifiable at scanner.orgn.com by anyone with the request ID.

Article illustration

The TDX Sandbox indicator is the first thing to look for here. It shows the execution boundary before the attestation report provides the cryptographic evidence behind that boundary.

The attestation report fields a security team reads directly:

Bash
Sandbox ID:       wt-4f9a2b
Worktree context: team/project/branch
TDX quote:        [hardware-signed Intel TDX evidence]
Measurements:     [SHA-384 digests of measured launch environment]
Issued:           2026-08-12T09:14:32Z

The Measurements field is what makes this actionable for procurement reviews. A digest mismatch means the image was modified between a known-good reference and the time the report was generated, a cryptographic failure verifiable by anyone with the report, not a policy violation you're taking the vendor's word on.

Two Confidentiality Layers That Must Stay Separate

Runtime confidentiality and inference confidentiality protect different things, and the failure mode is treating them as the same control. Runtime covers code, terminal execution, and agent tool calls inside the sandbox. Inference covers prompts and model outputs in transit to and from the LLM provider. A team that secures inference via ZDR and assumes that extends to the execution environment is wrong: what the agent does with tool call results after a model response comes back runs entirely in the runtime layer, which ZDR doesn't touch.

ORGN CDE makes this distinction explicit in its documentation. Cloud worktrees run inside TDX Trust Domains, the runtime layer, with encrypted VM memory inaccessible to the cloud operator, the hypervisor, and ORGN. Inference routes through ORGN Gateway separately, where teams choose between ZDR (policy-based, no hardware proof) and TEE models (hardware-isolated inference with a per-request attestation receipt). Each layer needs independent configuration and independent evidence.

Article illustration

Bash commands Origin Agent runs no longer have to block the chat, send a command to the background and keep working while it finishes. Commands still running in the foreground show the full inline orchestration card in chat, so you can follow command and subagent progress without switching to a separate terminal tab.

One terminology note that matters in security reviews: TDX uses Trust Domains, not SGX enclaves. The two technologies have different threat models, different attestation structures, and different PKI chains. Confusing them in a procurement questionnaire is an error that auditors with confidential computing experience will catch immediately.

Parallel Agents and the Sandbox Isolation Requirement

Two agents working on the same codebase inside the same VM produce write conflicts, may leak secrets between workstreams through shared memory or temporary files, and generate an audit trail that can't be attributed to individual sessions after the fact. The problem compounds in multi-agent orchestration: a coordinator spawning four implementor agents on shared infrastructure has created four attribution problems and at least as many isolation failures waiting to happen.

The correct architecture is one agent per sandbox. In ORGN CDE, each worktree is both a separate branch and a separate TDX sandbox. Agents on different worktrees can't read each other's memory because the isolation is enforced at the hardware level, not the application level. When a worktree is torn down, the TDX-encrypted sandbox is deleted; no data persists across the worktree boundary in a recoverable form. For teams running parallel agents on sensitive repositories, sandbox isolation should be an input to the agent architecture from the start, not a property verified after incidents surface.

Inference Attestation: Proving Where the Model Call Ran

The runtime layer answers where the agent executes and the inference layer answers where the model call runs. These are separate systems with separate controls and separate audit artifacts. Most teams have a defensible answer for runtime; very few have a verifiable answer for inference.

TEE Inference vs. ZDR: What Each Controls and Where Each Fails

ZDR stops data retention at the provider after processing completes. It doesn't prevent a compromised provider endpoint from reading a prompt during inference; it doesn't prove the model ran in an isolated environment, and it produces no artifact a compliance team can independently check. For workloads where prompt contents are themselves sensitive, export-controlled source code, classified operational planning, financial models with material non-public information, ZDR's assurance boundary ends at "we won't keep it," which is a weaker claim than the workload's risk profile demands.

TEE inference runs the model call inside hardware-isolated compute and generates a per-request attestation receipt signed by the hardware. The receipt proves a specific model call ran inside a verified TEE at a specific time on hardware matching the attested configuration. It deliberately excludes prompt and completion content, because attestation proves where and how inference ran, not what was inferred, which is the claim an auditor actually needs to evaluate.

Article illustration

ORGN Gateway gives teams a choice between TEE models (prefixed near_*, phala_*, tinfoil_* in the model catalog) and ZDR models (prefixed vercel_*), selected explicitly per request. Each TEE call produces a receipt in ORGN Scanner at /request/:requestId:

Bash
Request ID:       req_7c3f...
Model:            vercel_claude_sonnet_4_6
Provider:         TEE
Token count:      4,218
Attestation:      VERIFIED
Signing address:  0x4a2b...
Intel evidence:   [provider attestation artifact]
Latency:          1,240ms

You don’t need any login to read it. An auditor with the request ID confirms the inference ran in a verified TEE without asking ORGN for anything.

Audit Trails That Survive a Compliance Review

Three categories of audit artifacts exist in a well-instrumented agentic deployment, and they have different evidential weight. Agent-generated logs record what the agent reported about its own actions, useful for debugging, low value for compliance because a compromised agent could have produced different logs. Infrastructure-level telemetry records what the execution environment observed, stronger because it doesn't depend on agent-reported data. Cryptographic receipts record what the hardware signed, strongest because no software layer can forge them after the fact.

Article illustration

Shows how audit evidence becomes stronger as it moves from agent-reported logs to infrastructure telemetry and finally to hardware-signed receipts that auditors can independently verify.

ORGN Scanner operates at the infrastructure level: it exposes request metadata derived from the Gateway's infrastructure, not from agent-reported data, and never displays prompt contents or completions. A compliance review asking, "Can you prove these inference calls ran in a verified environment?" gets answered with a stack of Scanner receipts with attestation status VERIFIED, or it gets answered with a policy promise. Those two aren't equivalent as evidence, and an experienced auditor won't treat them as such.

Observability Inside the Trust Boundary

Attestation and inference receipts answer where execution happened. Observability answers what happened during execution. For regulated workloads, both questions need answers, and both sets of answers need to stay inside the same trust boundary governing the agent.

Why Shipping Telemetry Out Breaks the Security Model

A team running agents inside a TDX sandbox and shipping OpenTelemetry traces to an external SaaS observability platform has created an intentional exfiltration channel through their own security perimeter. Every trace contains execution context: function names, error messages, timing data, call paths, database query shapes. For export-controlled or classified workloads, that telemetry carries the same sensitivity as the code itself. Shipping it to an external platform replicates the data exposure problem the TDX sandbox was designed to prevent.

Routing telemetry through an enterprise SIEM that lives outside the trust boundary doesn't fix the problem when the workload's sensitivity level determines where the trust boundary sits. A SIEM receiving traces from an agent handling classified material now holds classified material, regardless of the SIEM's security posture.

In-IDE Observability Without Leaving the Secure Perimeter

Confidential Observe (beta) in ORGN CDE gives each project its own OpenTelemetry stack, traces, logs, and metrics, without routing telemetry outside the IDE. Setup is agent-driven: an Origin Agent skill instruments the code and opens a pull request. Telemetry is ingested to intake.observe.orgn.com through a write-only API key that can't retrieve data back out. The data stays inside the Confidential Observe sidebar, viewable within the same IDE session where the agent ran, inside the same TDX boundary that governed the execution.

The setup prompt Origin Agent handles via the oxyz-official/skills package instruments every service with OpenTelemetry in a single run. No manual SDK wiring, no external telemetry pipeline, no third-party observability platform receiving execution data from a regulated workload. The compliance workflow that results: the team runs an agent on a cloud worktree, the worktree attestation report proves the sandbox was genuine TDX hardware, the Scanner receipt proves each model call ran in a verified TEE, and the Confidential Observe dashboard shows traces and errors from inside the same perimeter. Three independent evidence artifacts for three separate auditor questions, none of them requiring data to leave the trust boundary.

Agentic AI Security Requires Architecture, Not Controls Added After the Fact

The three layers this article covered, identity and access, execution environment proof, and inference attestation, share one structural property: none of them can be added to an agentic deployment after it's built. A permission model bolted on after an agent architecture is running inherits whatever access the agent was provisioned with. An attestation layer added to a deployment already routing telemetry through external SaaS doesn't un-ship the data already sent. Sandboxing added to a multi-agent system after the agents are deployed doesn't retroactively isolate the sessions that already ran in shared infrastructure.

The question to put to any regulated team evaluating agentic AI: can you produce cryptographic evidence of where each agent ran, what model handled each inference call, and what the execution environment looked like at the time of the audit? If any of those three answers requires trusting a vendor's assurance rather than verifying an artifact, the compliance exposure is structural. This article walked through what verifiable evidence looks like at each layer: sandbox attestation from a TDX Trust Domain checkable against Intel's PKI, inference receipts from TEE model calls retrievable in ORGN Scanner by request ID, and observability telemetry that stays inside the trust boundary via Confidential Observe. The decisions that determine whether a team can produce those artifacts are made at architecture time, not audit time.

FAQs

1. What makes agentic AI security different from traditional AI security?

Traditional AI security focuses on model output: harmful content, leaked data, biased responses. Agentic AI security focuses on what an agent executes: tool calls, API requests, file writes, and subagent delegation. The distinction matters because an agent takes irreversible actions at machine speed, and the blast radius of a compromised session scales with its permissions, not with the content it generates.

2. How do I prevent prompt injection attacks in AI agents?

The defensible fix is structural, not instructional. Separate untrusted external content from trusted instructions at the architecture level, require a policy check before any tool call executes, and bind agent credentials to attested identities so a message claiming elevated permissions receives no elevation regardless of its content. System prompt hardening alone can't close the surface because injection targets the agent's trust model, not the model's output filters.

3. What is the difference between ZDR and TEE for AI workloads?

ZDR is a provider agreement not to retain prompts or completions after processing, a policy control with no hardware proof. TEE inference runs the model call inside hardware-isolated compute and generates a per-request cryptographic receipt verifiable independently. For workloads where prompt contents are themselves sensitive, ZDR's assurance ends at the provider's word; TEE inference gives you an artifact you can check yourself.

4. Can an AI agent's sandbox attestation be used as compliance evidence in a FedRAMP or CMMC audit?

A TDX sandbox attestation report is a cryptographic artifact verifiable against Intel's public PKI, not a vendor's self-attestation. Whether it satisfies a specific control depends on the control's requirements and your authorizing official's interpretation, but it produces hardware-backed evidence that those frameworks are designed to evaluate. Pair it with inference attestation receipts and an audit trail that stays inside the trust boundary, and the evidence package addresses the questions auditors ask about AI workloads.