Guide

AI coding agent security: a checklist for your own machine

A coding agent runs commands on your machine, next to your files and the tokens in your environment. This checklist covers what it can reach, what it writes down, how to check what it did, and how to shrink what it could do next. Claude Code is the worked example.

Checked against the Claude Code, GitHub, Stripe, AWS and MCP documentation, and the advisories and research under Sources. Published .

01 / Reach

Know what the agent can reach.

Most of the exposure sits in what the agent can do without asking you, so start with the mode a session starts in.

Starting mode With Claude Code v2.1.283 or later, an interactive session in a terminal or VS Code starts in auto mode, where a classifier model reviews actions instead of you. On earlier versions that is the starting mode only on Pro, Max and Team plans. The --permission-mode flag or permissions.defaultMode in a settings file starts it in another mode.
Manual mode The mode you can choose instead, config value default. Claude Code starts with read-only permissions and asks before it edits a file or runs a command that can change your system. It can write only inside the folder it was started in and its subfolders, and asks before Read, Grep or Glob go outside it.
No prompt A built-in set of read-only commands runs without a prompt in every mode: ls, cat, grep, find, head, tail and others. The set is not configurable. To make one of them ask, add an ask or deny rule for it.
Environment Bash commands run with the environment Claude Code was started in. Even sandboxed commands inherit it by default, credentials included, so a token exported in that shell is in reach of every command the agent runs.
Network Your prompts and the model's outputs go to the model provider, encrypted in transit with TLS 1.2 or later. WebFetch is a second way out. It runs in Claude Code's own process and follows the permission rules, not the sandbox's network allowlist.
MCP servers Anthropic reviews connectors against its listing criteria before adding them to its directory, but does not security-audit or manage any MCP server. The MCP specification's security guidance says clients should warn that a local MCP server runs with the same privileges as the client.
cd ~/code/myapp && claude          start in the project, not in ~
claude --permission-mode default   Manual mode: ask before edits and commands
/permissions                       in the session: every rule, and the file it came from

Started in your home directory, the folder boundary is your whole home directory, and Claude Code does not save trust for it, so the trust prompt comes back on every launch. To make the file tools refuse anything outside the working directories in every permission mode, set permissions.blockReadsOutsideWorkingDirectories. It covers Read, Grep, Glob and LSP, needs Claude Code v2.1.257 or later, and does not refuse shell commands the same way.

~/.claude/settings.json
{
  "permissions": {
    "blockReadsOutsideWorkingDirectories": true
  }
}

WebFetch's preapproved domains

In Manual and acceptEdits modes, WebFetch fetches a built-in set of documentation domains without asking. Until 2.1.163 that set held huggingface.co as a bare hostname, so an attacker's own model repository went through without a prompt too (CVE-2026-54316). Hugging Face counts each file request as a download, and Novee read a secret back one character at a time from which of 64 repositories' counters moved.

Since 2.1.162, an explicit WebFetch(domain:...) rule overrides the preapproved set, so a deny rule such as WebFetch(domain:huggingface.co) removes a domain you do not use. Where nothing needs the web, set CLAUDE_CODE_DISABLE_WEB_FETCH=1 (2.1.285 or later).

02 / Boundary

Put a boundary around it.

Permission rules decide what Claude Code's own tools may touch, and Claude Code enforces them, not the model: a line in CLAUDE.md changes nothing. The sandbox decides what every process may touch. Use both.

.claude/settings.json
{
  "permissions": {
    "deny": [
      "Read(./.env)",
      "Read(./.env.*)",
      "Read(./secrets/**)"
    ]
  }
}

In the project's .claude/settings.json the rules travel with the repository; in ~/.claude/settings.json they apply to every session you start. A ./ path is relative to the directory you start in. A single leading slash is not: in user settings, Read(/secrets/**) means ~/.claude/secrets/**. For a user-level rule that reaches into your projects, write // for an absolute path or ~/ for your home: Read(//**/.env) matches any .env on the filesystem.

Covers Claude's file tools, the Bash file commands Claude Code recognises (cat, head, tail, sed, tee), and redirection targets such as < .env.
Misses grep -r pattern . run from the folder that holds the file, and a Python or Node script that opens files itself.
Bash rules They match the command as written. As a deny rule, Bash(rm *) stops rm -rf build/, not /bin/rm -rf build/ or bash -c 'rm -rf build/'. They are not a security boundary around the program.

The sandbox

The sandbox is the OS-level boundary: the operating system enforces which files and domains a sandboxed command and its child processes can reach, whatever the command text says. Run /sandbox in a session to turn it on. macOS uses the built-in Seatbelt framework, with nothing to install. Linux and WSL2 need bubblewrap and socat. Native Windows is not supported.

By default it still lets commands read ~/.aws/credentials and ~/.ssh/, and there is no built-in credential deny list: only what you list is restricted. This is Anthropic's own example.

~/.claude/settings.json
{
  "sandbox": {
    "enabled": true,
    "credentials": {
      "files": [
        { "path": "~/.aws/credentials", "mode": "deny" },
        { "path": "~/.ssh", "mode": "deny" }
      ],
      "envVars": [
        { "name": "GITHUB_TOKEN", "mode": "deny" },
        { "name": "NPM_TOKEN", "mode": "deny" }
      ]
    }
  }
}

sandbox.credentials covers sandboxed Bash only. To strip credentials from every subprocess, sandboxed or not, set CLAUDE_CODE_SUBPROCESS_ENV_SCRUB=1. The Bash tool, hooks and MCP stdio servers then run without Anthropic and cloud provider credentials and any other variable Claude Code recognises as a credential. It removes what it recognises, so keep naming your own variables in sandbox.credentials.

Anthropic is plain about the limits. The sandbox reduces risk but is not a complete isolation boundary. Its proxy allows a connection by hostname and does not inspect TLS by default, so allowing a broad domain such as github.com can open a path for data exfiltration. And a command that fails in the sandbox can be retried outside it, through the normal permission prompt, unless you set "allowUnsandboxedCommands": false.

Run a patched build

The sandbox and the trust dialog are only as good as the build. With autoAllowBashIfSandboxed at its default, a sandboxed command runs without a prompt, so an escape meets none. Auto-update applies these fixes. If you update by hand, claude --version should print 2.1.260 or later.

2.1.64 A symlink made by a sandboxed command led Claude Code's own process to write outside the workspace (CVE-2026-39861).
2.1.84 A repository could skip the trust dialog, and run its hooks, through a spoofed git worktree commondir (CVE-2026-40068).
2.1.163 A worktree named .git, symlinks and git fsmonitor let a repository overwrite files such as ~/.zshenv outside the sandbox (CVE-2026-55607).
2.1.247 On macOS, a core.fsmonitor command planted from inside the sandbox ran outside it through Claude Code's own git calls: Accomplish AI's Beltdown, with no advisory or CVE.
2.1.260 Team and Enterprise sessions could run without the organisation's server-managed policy (CVE-2026-103012, in the audit log guide).

Organisations can set requiredMinimumVersion through MDM or a managed settings file, and an older build then refuses to start. The key needs 2.1.163 or later.

managed-settings.json
{
  "requiredMinimumVersion": "2.1.260"
}

A repository you did not write

A clone can bring its own .claude/settings.json and .mcp.json, whose hooks and MCP servers run as you once you accept the trust dialog. Under -p, trust verification is off.

It need not hold the harmful code itself. In 0DIN's demonstration of 25 June 2026, a package raised an error naming python3 -m axiom init until that had run, and Claude Code, asked only to get the project running, ran it. The script behind it piped a DNS TXT record from the attacker's domain to bash, which opened a reverse shell as the developer. The write-up names no Claude Code version or permission mode.

An install script can rewrite the configuration itself: Mitiga showed an npm postinstall script and a repository hook re-pointing an OAuth MCP server at a localhost proxy, to read its tokens (the audit log guide has the chain). In October 2025, Koi Security reported 126 PhantomRaven packages on npm, some named after packages language models invent, that ran a credential stealer at install.

cat .claude/settings*.json .mcp.json    before the first claude in a new clone
grep -n localhost ~/.claude.json         an MCP server you did not point there
npm install --ignore-scripts             no install scripts run

Run an unfamiliar project's setup yourself once you have read it, and not under claude -p. Check that a package the agent proposes is the one you meant. If an MCP server was re-pointed, remove the hook that did it and restore the URL before you rotate its token, or the new one goes the same way.

Keep Claude Code out of your .env files: the full walk-through →

03 / Disk

Know what it writes to disk.

Anything that passes through a tool is written to a transcript: file contents, command output, pasted text. In plaintext, on your disk.

PathWhat it holds
~/.claude/projects/<project>/<session>.jsonl Every message, tool call and tool result in a session. A .env value a tool read is written here.
<session>/subagents/ Subagent transcripts, removed with the parent transcript when it ages out.
<session>/tool-results/ Large tool outputs, spilled to separate files.
~/.claude/file-history/ Pre-edit snapshots of files Claude changed, kept for checkpoint restore.
~/.claude/history.jsonl Every prompt you typed, with timestamp and project path. The retention sweep does not remove it.

Claude Code's transcripts and history are not encrypted at rest. File permissions are the only protection. Claude Code deletes transcripts older than cleanupPeriodDays: 30 days by default, and 1 at the least. Lower it to hold less on disk, or raise it to keep a longer record to audit, but choose it. CLAUDE_CODE_SKIP_PROMPT_HISTORY=1 stops Claude Code writing transcripts and prompt history at all, which also leaves nothing to check afterwards.

ls -lt ~/.claude/projects | head      the projects used most recently
ls -ld ~/.claude                      who else can read it
uvx ranwhat clean --no-interactive    credentials in Claude Code transcripts; changes nothing

ranwhat clean reads session transcripts from the last 30 days by default (--days N). When the transcript names the file a credential was read out of, clean names it too. It reads subagent transcripts as well, but not tool-results/ or history.jsonl yet.

Where Claude Code keeps its history, and for how long →

04 / Audit

Check what it already ran.

The transcript is the record you already have, for as long as retention keeps it. Read it before it ages out.

$ uvx ranwhat check

One read-only pass: the risky actions in your Claude Code history, and the credentials left in its transcripts. Every ranwhat command on this page also runs as uvx ranwhat … without installing anything.

ranwhat watch --days 30    risky actions from the last 30 days
ranwhat watch --json       one record per flagged call

ranwhat watch reports nine kinds of action: credential reads, secret-shaped strings in tool calls, package publishing, cloud resource changes, payment API calls, log or history tampering, destructive git, recursive deletion, and local files uploaded with curl. It reads subagent transcripts too.

Without ranwhat, a first look is one grep away, because each transcript is one JSON event per line: grep -l 'rm -rf' ~/.claude/projects/*/*.jsonl lists the sessions that mention it, whether or not it ran.

A record from now on

A PostToolUse hook runs after each tool call that succeeds, and a command that exits non-zero fires PostToolUseFailure instead. The same entry under both, matching Bash, appends every command to a file. The log holds each command as written, so a token passed on a command line ends up in it too: keep it as private as the transcripts.

Or export events with OpenTelemetry. CLAUDE_CODE_ENABLE_TELEMETRY=1 turns it on, and Bash command strings are logged only with OTEL_LOG_TOOL_DETAILS=1, which is off by default. The audit log guide has both configurations in full.

Build an audit log for Claude Code →

If Claude Code deleted your files →

05 / Tokens

Scope the tokens it holds.

A command the agent runs can do anything a token in its environment allows. Rules and sandboxes have gaps, as section 02 shows. A token that cannot delete a repository cannot be made to.

Where In the environment or a secrets manager, not on a command line. Arguments are visible to other users on the machine through the process table, and a typed command line lands in shell history.
GitHub Fine-grained personal access tokens can be limited to specific repositories, each with specific permissions. GitHub recommends them over classic tokens wherever possible, and highly recommends an expiration date.
Stripe A restricted key with only the permissions the agent needs, tagged as an agent key when you create it. Agent-tagged keys fall under Stripe's approval rules for sensitive actions such as refunds and payouts. From 31 October 2026 the Stripe MCP server refuses full-access secret keys and restricted keys without the Agent tag.
AWS IAM last accessed information shows which services, and for some services which actions, a role or user has used, so you can remove the rest. It counts attempts, denied ones included; CloudTrail is the authoritative record.
MCP servers The MCP security guidance lists wildcard or omnibus scopes, such as *, all or full-access, as a common mistake. Grant a server the smallest set it needs.
read -rs RANWHAT_GITHUB_TOKEN    paste it: not echoed, not saved to history
export RANWHAT_GITHUB_TOKEN
ranwhat live                     ask each token's own provider, read-only
ranwhat demo                     the same report on a bundled example

ranwhat live sends each token only to the provider that issued it, and scores what it allows. A fine-grained GitHub token or an rk_ restricted key comes back with no scope list: declare its permissions in a profile and use ranwhat scan, whose --pull-usage also marks which grants were used.

AI agent token scopes: the longer guide →

06 / Leaks

When a secret has already leaked.

Rotate first. Masking the transcript stops a second leak; it does not undo the first. The value was already on disk and already in a model context you do not control.

AWS AWS's sequence: create a second access key, move every application to it, check when the old key was last used, deactivate it, then delete it.
Stripe Rotating a key in the Dashboard keeps the old key working for up to 7 days unless you set its expiration to Now. For a key that leaked, every day of overlap is a day the copy still works.
AWS CLI
aws iam create-access-key
aws iam get-access-key-last-used --access-key-id <old-key-id>
aws iam update-access-key --access-key-id <old-key-id> --status Inactive
aws iam delete-access-key --access-key-id <old-key-id>

For credentials in Claude Code transcripts, ranwhat clean opens a review session after its report. rotate lists what to rotate, grouped by provider, and mask all masks every value listed, after a backup to ~/.ranwhat/backups. The backups still hold every masked value: delete them once the transcripts look right.

ranwhat clean               report, then open the review session
ranwhat> rotate           what to rotate, grouped by provider
ranwhat> mask all         mask everything listed, backups first

What to do when Claude Code read your .env →

07 / Checklist

The checklist.

One line each, in the order worth doing them. Each points back to its section.

  1. 01 Run claude --version: 2.1.260 or later. Organisations: set requiredMinimumVersion. Boundary
  2. 02 Start Claude Code in the project directory, not in ~. Reach
  3. 03 Review /permissions, and which settings file each rule comes from. Reach
  4. 04 Deny preapproved WebFetch domains you do not use, or set CLAUDE_CODE_DISABLE_WEB_FETCH=1 where nothing needs the web. Reach
  5. 05 Deny reads of .env files and secrets directories. Boundary
  6. 06 Turn on the sandbox and list your credential files and variables in sandbox.credentials. Boundary
  7. 07 Set CLAUDE_CODE_SUBPROCESS_ENV_SCRUB=1. Boundary
  8. 08 Read a new clone's .claude/ and .mcp.json before you accept trust, and run its setup yourself. Boundary
  9. 09 Check who can read ~/.claude. Disk
  10. 10 Set cleanupPeriodDays on purpose: how long Claude Code keeps history. Disk
  11. 11 Run uvx ranwhat check now, and again after anything odd. Audit
  12. 12 Log commands from now on, with a hook or OpenTelemetry. Audit
  13. 13 Replace broad tokens with fine-grained, restricted or agent-tagged ones. Tokens
  14. 14 Rotate anything found in a transcript, then mask it. Leaks
08 / Scope

What ranwhat covers, and what it does not.

ranwhat is the after-the-fact half: what already ran, what is on disk, and what the credentials allow.

Covers watch: risky actions in Claude Code history. clean: credentials in Claude Code transcripts, masked only when you ask. scan and live: the authority a set of credentials carries.
Does not Block anything while the agent works: Claude Code's permission rules, hooks and sandbox do that. It does not scan MCP servers or skills for malicious content. It flags a WebFetch call only when a secret-shaped value is in it, and does not flag a setup command that fetches its payload at run time.
Network Reads locally. live and --pull-usage ask only the provider that issued each token.
$ uvx ranwhat check
Install