What Happens When an AI Coding Agent Gets a Terminal?

· Updated

ORGN Team

TL;DR

  • Once a coding agent gets shell access, its security boundary becomes the environment behind the shell, so reviewing the model's tool permissions alone won't reveal whether it can reach credentials, local databases, SSH configuration, or cloud sessions.
  • A single approved command can cross several authority boundaries when scripts inherit environment variables, package credentials, network access, and writable files; the failure becomes materially worse when the developer workstation exposes production identities or sensitive data.
  • Prompt injection becomes an execution problem when untrusted repository content influences an agent that can call tools, because a malicious instruction inside documentation or another project artifact can turn a model decision into a command with external side effects.
  • Human approval reduces unattended execution, but it doesn't shrink the authority of an approved command; isolation has to enforce the boundary when users approve actions without inspecting every downstream process.
  • Containers, VMs, and confidential execution solve different security questions, so choosing an execution boundary requires matching the isolation property to the threat model rather than treating every sandbox as equivalent.
  • A defensible coding-agent environment separates filesystem access, network egress, credentials, and execution from the developer host, while attestation provides evidence of where execution or inference occurred without proving the agent chose a safe action.

Secure Coding Assistants Stop Being Code Generators Once They Get Execution Access

A secure coding assistant needs controls around both model context and execution authority because modern coding agents don't stop at producing text. Nearly half of respondents in a 2026 CNCF survey of 133 contributors reported using AI assistants directly inside IDEs or command-line interfaces, putting model decisions much closer to shells, repositories, and developer credentials than a browser chatbot ever gets.

The security difference becomes concrete when an assistant receives terminal authority. One practitioner on Reddit described a coding agent running in a local development environment that discovered MongoDB-related access through existing configuration and surfaced production data, even though the requested task had nothing to do with production.

Reddit thread: "Anyone else worried about coding agents discovering access they were never meant to use?"
Reddit thread: "Anyone else worried about coding agents discovering access they were never meant to use?"

The useful boundary, then, isn't "AI-generated code versus human-generated code." Platform teams need to ask which files, commands, credentials, network destinations, and external identities become reachable after the agent decides to act.

The Terminal Changes the Security Boundary

Autocomplete places a human between generated text and execution. An agent with shell authority collapses part of that separation because a model response becomes input to an operating system interface.

From suggestion to execution

Consider a conventional completion workflow where the model proposes a function, the developer reads it, inserts it, runs the test suite, and decides whether to keep the change.

A terminal-capable agent operates through a longer feedback loop: inspect files, choose a command, execute it, read stdout or stderr, modify files, then choose another action. NIST describes agent systems in similar operational terms: agents plan and take autonomous actions that affect external systems, while security problems arise when model outputs connect to software functionality.

The shell is therefore an authority boundary: a command inherits whatever the execution environment exposes.

ORGN's workspace makes the execution location explicit: each worktree has a terminal inside the sandbox, and commands entered there execute against that worktree.

ORGN Studio workspace: chat interface and repository file tree inside an isolated worktree
ORGN Studio workspace: chat interface and repository file tree inside an isolated worktree

What the Agent Reaches Through a Terminal

A repository is only one object visible from a development shell, but the effective boundary includes every resource the process inherits, mounts into its environment, exposes through a credential helper, or reaches over the network.

The repository is the beginning of the access graph

Source files are obvious, but less obvious paths include .env files, Git history, package-manager configuration, local logs, SSH configuration, cloud CLI sessions, build output, cached credentials, local databases, and documentation stored beside the project.

Diagram: the agent's access graph — a shell process reaching into filesystem, credentials, local data, and network
Diagram: the agent's access graph — a shell process reaching into filesystem, credentials, local data, and network

Start by following the SHELL PROCESS node downward, then trace each resource it inherits or reaches. The easiest path to misread is CREDENTIALS, because the agent doesn't need a dedicated credential tool when the shell already exposes an authenticated identity.

The distinction matters because the model doesn't need a dedicated "read production credentials" tool when a shell process already has a path to them. NIST's August 2026 discussion of agent identity warns specifically about static API keys and tokens appearing in configuration files, Markdown files, and logs, and argues for agent-specific identities and tightly scoped credentials instead of sharing human credentials with an agent.

The same distinction applies to routine development commands. ORGN's public Studio uses pnpm test for running a project's test suite inside a worktree, so the operational point is where the command executes, not what pnpm itself does.

ORGN Studio changes panel listing every file modification made during a session
ORGN Studio changes panel listing every file modification made during a session

In the changes panel, one can see all file modifications made during the session or worktree. Inside an unrestricted local shell, the test process inherits the local environment subject to OS and shell configuration. Inside the documented ORGN worktree model, the command executes inside the project's TDX sandbox.

Model Authority Determines the Blast Radius

A coding agent doesn't have a single binary permission called "terminal access." Authority emerges from the combination of filesystem visibility, shell execution, credentials, network routes, Git permissions, external tools, and infrastructure identities.

Suppose an agent receives permission to execute a build script. The script invokes another executable, reads an environment variable, contacts a package registry, and writes generated files; the original approval has now crossed several boundaries without another model decision. A useful review therefore follows capability transitions:

Agent actionAuthority introducedFailure consequence
Read workspace filesData accessSource, configuration, or secrets enter agent context
Execute shell commandsProcess executionAvailable binaries and scripts become callable
Use credentialsExternal identityAgent actions inherit account permissions
Reach external hostsNetwork egressData or credentials leave the execution boundary
Write through GitRepository mutationAgent-generated changes enter shared development state
Use cloud credentialsRemote control-plane accessLocal agent decisions affect hosted resources

Installing a dependency illustrates how quickly authority expands. The command needs package-registry access and modifies dependency state. A security review of agent execution therefore has to cover the package manager's network path, configuration, credentials, lifecycle scripts, and writable files, not merely the agent's original natural-language instruction.

The Model Doesn't Need Malicious Intent

Indirect prompt injection turns data the agent reads into instructions the model could follow. Once the same agent also has execution authority, an input-handling failure can lead to shell commands and external side effects.

Untrusted context becomes part of the control path

A repository might contain instructions in documentation, issue text, generated artifacts, test fixtures, or dependency content. An agent asked to investigate the project reads that material because reading project data is part of its job.

NIST's 2026 agent-hijacking work examined more than 250,000 attack attempts from over 400 participants against 13 frontier models. Researchers found at least one successful hijacking attack against every target model, including scenarios involving coding and tool-using agents.

The failure sequence doesn't require the model to become an attacker. Untrusted content influences a model decision, the agent translates the decision into a tool call, and the environment determines how far the resulting action reaches.

CISA recorded a concrete version of the same class of problem in 2025. CVE-2025-54135 described an indirect prompt-injection path in older Cursor versions where creation of a sensitive MCP configuration file could contribute to remote code execution without user approval.

Tool Permissions Don't Define the Whole Security Boundary

An agent policy might permit "run shell commands" while the shell itself holds access the policy never names. Approval dialogs govern requested actions; they don't remove ambient authority already present inside the process environment.

A shell with an authenticated AWS CLI, an unlocked SSH agent, Git credentials, and database environment variables carries those identities regardless of whether the agent UI exposes separate AWS, SSH, Git, or database tools. Human approval isn't a complete substitute for isolation either. NIST points to consent fatigue as a weakness in human-in-the-loop controls: repeated approval requests condition users to approve actions reflexively, weakening the accountability the prompt was meant to create.

The safer architectural question is narrower: what authority remains available after the user approves the shell action?

Isolation Puts a Ceiling on Agent Authority

A process running on the developer host starts in the host's security context unless engineers explicitly remove access. Moving execution into a separate workload boundary changes the maximum set of resources available even after the agent receives permission to run a command.

Containers, VMs, and confidential execution answer different questions

A container separates filesystems, namespaces, and process resources while sharing the host kernel. A VM places the workload behind a guest OS and virtual hardware boundary; microVM designs narrow that machinery for isolated workloads.

Confidential VMs add hardware-backed protection for workload memory against parts of the hosting stack. The extra property matters when the threat model includes infrastructure operators or compromised host software, but it doesn't decide whether an agent should receive a production credential.

ORGN uses a TDX-secured sandbox for project execution, with worktrees inside that project runtime. Terminal commands, file edits, and agent tool calls run inside the sandbox rather than on the developer's laptop.

ORGN Studio workspace model: a project-level TDX sandbox around worktrees and sessions
ORGN Studio workspace model: a project-level TDX sandbox around worktrees and sessions

ORGN's workspace model places a project-level TDX sandbox around worktrees and sessions. Worktrees separate Git branches while sharing the project's runtime boundary; branch isolation and compute isolation therefore solve different problems.

Starting a development server is another ordinary command whose risk depends on its execution boundary. A development server might bind a port, read application configuration, load environment variables, and communicate with dependent services. Running it inside an isolated workspace restricts its starting environment, while network policy and credential scope still determine what sits beyond that boundary.

Filesystem, Network, Credentials, and Execution Need Separate Controls

Sandboxing answers where code executes, not which resources the workload should receive. A useful agent environment treats workspace data, egress, identity, and command execution as separate control decisions.

Filesystem scope

Mount only the project data required for the task, not the developer's home directory. Keep personal SSH configuration, unrelated repositories, cloud profiles, and host credential stores outside the agent runtime.

Network scope

Outbound access deserves its own policy because shell execution plus unrestricted egress creates a direct path from readable data to external services. Package registries, source hosts, internal APIs, and arbitrary internet destinations don't carry equivalent risk.

Credential scope

Agent credentials should describe the agent's job, not mirror the developer's authority. Short-lived identities with narrow audiences reduce the damage a stolen token can cause and make agent actions easier to distinguish from human actions.

ORGN stores project secrets as environment variables or API keys for injection into the execution environment, and its settings documentation says it encrypts and masks saved values after storage. The remaining design decision belongs to the team: a securely stored credential with excessive permissions still grants excessive authority.

Sandboxing and Confidential Inference Protect Different Boundaries

Protecting agent execution doesn't establish where model inference ran. Conversely, protecting inference doesn't restrict which files or credentials an agent reaches after a response returns.

Runtime proof versus inference proof

ORGN separates those boundaries in its current architecture. Cloud worktrees execute code, terminal commands, and agent tool calls inside an Intel TDX Trust Domain, while Origin Agent inference travels through Gateway; the selected model tier determines whether inference relies on a zero-data-retention policy or a hardware-backed TEE path.

ORGN Studio project settings: AI Models panel with Allow ZDR Models and Allow TEE Models toggles
ORGN Studio project settings: AI Models panel with Allow ZDR Models and Allow TEE Models toggles

Follow the SANDBOX EXECUTION path first and notice where its attestation is generated. The easiest state to misread is the relationship between TDX ATTESTATION and INFERENCE RECEIPT: each proves a different boundary, so one artifact doesn't establish the other.

Consider the two proof paths: sandbox attestation shows where code and tools ran, while a TEE inference receipt shows where a specific model request ran; one artifact doesn't prove the other boundary.

Attestation narrows the trust claim further; ORGN Scanner distinguishes sandbox TDX attestation, which checks the compute environment, from per-request inference receipts, which bind a specific TEE model call to hardware evidence.

Diagram: two security boundaries — ORGN Studio sandbox execution vs. ORGN Gateway TEE inference, with separate attestations
Diagram: two security boundaries — ORGN Studio sandbox execution vs. ORGN Gateway TEE inference, with separate attestations
Sandbox TDX attestationInference attestation receipt
Proves: environment is genuineProves: specific call ran in TEE
Scope: per sandboxScope: per request
Source: Daytona attest-gatewaySource: NEAR AI / Phala providers
Page: /sandboxes/:sandboxIdPage: /request/:requestId

Attestation doesn't prove that a model chose a safe command, nor does it repair an overprivileged credential. It establishes properties of the execution location, leaving agent policy and resource authority as separate security decisions.

A Secure Coding Agent Architecture Starts With the Failure Case

Designing from the happy path encourages teams to grant whatever access makes the demo work. Designing from a failed model decision forces every downstream authority to justify its presence.

A bounded architecture places the coding interface and agent controller above an isolated execution environment. The workspace contains the required repository state, credentials carry task-specific permissions, network routes expose approved destinations, and shell commands execute inside that environment rather than inheriting the developer's workstation. ORGN Studio offers one concrete implementation of the execution side: projects run in TDX sandboxes, worktrees isolate Git branches inside the project, sessions record prompts and shell activity, and project settings control injected secrets.

The security review should still assume the model eventually chooses a bad action. If one incorrect model decision inherits production-grade authority, the agent environment is too permissive for the task.

What Platform Teams Should Carry Forward

Terminal access turns a coding assistant into a principal that acts through the authority of its execution environment. Review the agent by tracing filesystem reach, shell authority, network egress, credential scope, external identities, and the boundary separating agent execution from the developer host.

We covered how terminal access expands the blast radius, why indirect prompt injection becomes more serious once model output reaches execution tools, where permission prompts fall short, how isolation limits host exposure, and why sandbox attestation and inference attestation prove different properties. If one incorrect model decision inherits production-grade authority, the agent environment is too permissive for the task.

FAQs

Should an AI coding agent have terminal access?

Only when the task requires command execution and the shell runs with bounded filesystem, credential, and network access. Terminal access on a normal developer workstation inherits authority unrelated to the requested coding task.

Are containers enough for AI coding agents?

A container fits threat models where kernel sharing and configured host integrations are acceptable. Stronger isolation becomes relevant when the workload must sit behind a separate guest boundary or when the hosting infrastructure itself belongs inside the threat model.

Should coding agents receive production credentials?

A development agent shouldn't inherit production credentials merely because the developer already has them. Agent-specific, short-lived credentials with narrow permissions reduce both accidental reach and the consequences of prompt injection or credential theft.

Does human approval make terminal-capable agents safe?

Approval reduces unattended actions but doesn't change what an approved command can reach. Repeated prompts also introduce consent fatigue, so resource isolation and narrow authority remain necessary even when a human approves tool calls.