Guide
AI coding agent security: a checklist for your own machine
A coding agent runs commands on your machine, next to your files and the tokens in your environment. This checklist covers what it can reach, what it writes down, how to check what it did, and how to shrink what it could do next. Claude Code is the worked example.
Checked against the Claude Code, GitHub, Stripe, AWS and MCP documentation, and the advisories and research under Sources. Published .
Know what the agent can reach.
Most of the exposure sits in what the agent can do without asking you, so start with the mode a session starts in.
cd ~/code/myapp && claude start in the project, not in ~ claude --permission-mode default Manual mode: ask before edits and commands /permissions in the session: every rule, and the file it came from
Started in your home directory, the folder boundary is your whole home directory, and Claude Code does not save trust for it, so the trust prompt comes back on every launch. To make the file tools refuse anything outside the working directories in every permission mode, set permissions.blockReadsOutsideWorkingDirectories. It covers Read, Grep, Glob and LSP, needs Claude Code v2.1.257 or later, and does not refuse shell commands the same way.
{
"permissions": {
"blockReadsOutsideWorkingDirectories": true
}
}WebFetch's preapproved domains
In Manual and acceptEdits modes, WebFetch fetches a built-in set of documentation domains without asking. Until 2.1.163 that set held huggingface.co as a bare hostname, so an attacker's own model repository went through without a prompt too (CVE-2026-54316). Hugging Face counts each file request as a download, and Novee read a secret back one character at a time from which of 64 repositories' counters moved.
Since 2.1.162, an explicit WebFetch(domain:...) rule overrides the preapproved set, so a deny rule such as WebFetch(domain:huggingface.co) removes a domain you do not use. Where nothing needs the web, set CLAUDE_CODE_DISABLE_WEB_FETCH=1 (2.1.285 or later).
Put a boundary around it.
Permission rules decide what Claude Code's own tools may touch, and Claude Code enforces them, not the model: a line in CLAUDE.md changes nothing. The sandbox decides what every process may touch. Use both.
{
"permissions": {
"deny": [
"Read(./.env)",
"Read(./.env.*)",
"Read(./secrets/**)"
]
}
}In the project's .claude/settings.json the rules travel with the repository; in ~/.claude/settings.json they apply to every session you start. A ./ path is relative to the directory you start in. A single leading slash is not: in user settings, Read(/secrets/**) means ~/.claude/secrets/**. For a user-level rule that reaches into your projects, write // for an absolute path or ~/ for your home: Read(//**/.env) matches any .env on the filesystem.
The sandbox
The sandbox is the OS-level boundary: the operating system enforces which files and domains a sandboxed command and its child processes can reach, whatever the command text says. Run /sandbox in a session to turn it on. macOS uses the built-in Seatbelt framework, with nothing to install. Linux and WSL2 need bubblewrap and socat. Native Windows is not supported.
By default it still lets commands read ~/.aws/credentials and ~/.ssh/, and there is no built-in credential deny list: only what you list is restricted. This is Anthropic's own example.
{
"sandbox": {
"enabled": true,
"credentials": {
"files": [
{ "path": "~/.aws/credentials", "mode": "deny" },
{ "path": "~/.ssh", "mode": "deny" }
],
"envVars": [
{ "name": "GITHUB_TOKEN", "mode": "deny" },
{ "name": "NPM_TOKEN", "mode": "deny" }
]
}
}
}sandbox.credentials covers sandboxed Bash only. To strip credentials from every subprocess, sandboxed or not, set CLAUDE_CODE_SUBPROCESS_ENV_SCRUB=1. The Bash tool, hooks and MCP stdio servers then run without Anthropic and cloud provider credentials and any other variable Claude Code recognises as a credential. It removes what it recognises, so keep naming your own variables in sandbox.credentials.
Anthropic is plain about the limits. The sandbox reduces risk but is not a complete isolation boundary. Its proxy allows a connection by hostname and does not inspect TLS by default, so allowing a broad domain such as github.com can open a path for data exfiltration. And a command that fails in the sandbox can be retried outside it, through the normal permission prompt, unless you set "allowUnsandboxedCommands": false.
Run a patched build
The sandbox and the trust dialog are only as good as the build. With autoAllowBashIfSandboxed at its default, a sandboxed command runs without a prompt, so an escape meets none. Auto-update applies these fixes. If you update by hand, claude --version should print 2.1.260 or later.
Organisations can set requiredMinimumVersion through MDM or a managed settings file, and an older build then refuses to start. The key needs 2.1.163 or later.
{
"requiredMinimumVersion": "2.1.260"
}A repository you did not write
A clone can bring its own .claude/settings.json and .mcp.json, whose hooks and MCP servers run as you once you accept the trust dialog. Under -p, trust verification is off.
It need not hold the harmful code itself. In 0DIN's demonstration of 25 June 2026, a package raised an error naming python3 -m axiom init until that had run, and Claude Code, asked only to get the project running, ran it. The script behind it piped a DNS TXT record from the attacker's domain to bash, which opened a reverse shell as the developer. The write-up names no Claude Code version or permission mode.
An install script can rewrite the configuration itself: Mitiga showed an npm postinstall script and a repository hook re-pointing an OAuth MCP server at a localhost proxy, to read its tokens (the audit log guide has the chain). In October 2025, Koi Security reported 126 PhantomRaven packages on npm, some named after packages language models invent, that ran a credential stealer at install.
cat .claude/settings*.json .mcp.json before the first claude in a new clone grep -n localhost ~/.claude.json an MCP server you did not point there npm install --ignore-scripts no install scripts run
Run an unfamiliar project's setup yourself once you have read it, and not under claude -p. Check that a package the agent proposes is the one you meant. If an MCP server was re-pointed, remove the hook that did it and restore the URL before you rotate its token, or the new one goes the same way.
Keep Claude Code out of your .env files: the full walk-through →
Know what it writes to disk.
Anything that passes through a tool is written to a transcript: file contents, command output, pasted text. In plaintext, on your disk.
| Path | What it holds |
|---|---|
~/.claude/ |
Every message, tool call and tool result in a session. A .env value a tool read is written here. |
<session>/ |
Subagent transcripts, removed with the parent transcript when it ages out. |
<session>/ |
Large tool outputs, spilled to separate files. |
~/.claude/ |
Pre-edit snapshots of files Claude changed, kept for checkpoint restore. |
~/.claude/ |
Every prompt you typed, with timestamp and project path. The retention sweep does not remove it. |
Claude Code's transcripts and history are not encrypted at rest. File permissions are the only protection. Claude Code deletes transcripts older than cleanupPeriodDays: 30 days by default, and 1 at the least. Lower it to hold less on disk, or raise it to keep a longer record to audit, but choose it. CLAUDE_CODE_SKIP_PROMPT_HISTORY=1 stops Claude Code writing transcripts and prompt history at all, which also leaves nothing to check afterwards.
ls -lt ~/.claude/projects | head the projects used most recently ls -ld ~/.claude who else can read it uvx ranwhat clean --no-interactive credentials in Claude Code transcripts; changes nothing
ranwhat clean reads session transcripts from the last 30 days by default (--days N). When the transcript names the file a credential was read out of, clean names it too. It reads subagent transcripts as well, but not tool-results/ or history.jsonl yet.
Check what it already ran.
The transcript is the record you already have, for as long as retention keeps it. Read it before it ages out.
One read-only pass: the risky actions in your Claude Code history, and the credentials left in its transcripts. Every ranwhat command on this page also runs as uvx ranwhat … without installing anything.
ranwhat watch --days 30 risky actions from the last 30 days ranwhat watch --json one record per flagged call
ranwhat watch reports nine kinds of action: credential reads, secret-shaped strings in tool calls, package publishing, cloud resource changes, payment API calls, log or history tampering, destructive git, recursive deletion, and local files uploaded with curl. It reads subagent transcripts too.
Without ranwhat, a first look is one grep away, because each transcript is one JSON event per line: grep -l 'rm -rf' ~/.claude/projects/*/*.jsonl lists the sessions that mention it, whether or not it ran.
A record from now on
A PostToolUse hook runs after each tool call that succeeds, and a command that exits non-zero fires PostToolUseFailure instead. The same entry under both, matching Bash, appends every command to a file. The log holds each command as written, so a token passed on a command line ends up in it too: keep it as private as the transcripts.
Or export events with OpenTelemetry. CLAUDE_CODE_ENABLE_TELEMETRY=1 turns it on, and Bash command strings are logged only with OTEL_LOG_TOOL_DETAILS=1, which is off by default. The audit log guide has both configurations in full.
Scope the tokens it holds.
A command the agent runs can do anything a token in its environment allows. Rules and sandboxes have gaps, as section 02 shows. A token that cannot delete a repository cannot be made to.
read -rs RANWHAT_GITHUB_TOKEN paste it: not echoed, not saved to history export RANWHAT_GITHUB_TOKEN ranwhat live ask each token's own provider, read-only ranwhat demo the same report on a bundled example
ranwhat live sends each token only to the provider that issued it, and scores what it allows. A fine-grained GitHub token or an rk_ restricted key comes back with no scope list: declare its permissions in a profile and use ranwhat scan, whose --pull-usage also marks which grants were used.
When a secret has already leaked.
Rotate first. Masking the transcript stops a second leak; it does not undo the first. The value was already on disk and already in a model context you do not control.
aws iam create-access-key aws iam get-access-key-last-used --access-key-id <old-key-id> aws iam update-access-key --access-key-id <old-key-id> --status Inactive aws iam delete-access-key --access-key-id <old-key-id>
For credentials in Claude Code transcripts, ranwhat clean opens a review session after its report. rotate lists what to rotate, grouped by provider, and mask all masks every value listed, after a backup to ~/.ranwhat/backups. The backups still hold every masked value: delete them once the transcripts look right.
ranwhat clean report, then open the review session ranwhat> rotate what to rotate, grouped by provider ranwhat> mask all mask everything listed, backups first
The checklist.
One line each, in the order worth doing them. Each points back to its section.
- 01 Run claude --version: 2.1.260 or later. Organisations: set requiredMinimumVersion. Boundary
- 02 Start Claude Code in the project directory, not in ~. Reach
- 03 Review /permissions, and which settings file each rule comes from. Reach
- 04 Deny preapproved WebFetch domains you do not use, or set CLAUDE_CODE_DISABLE_WEB_FETCH=1 where nothing needs the web. Reach
- 05 Deny reads of .env files and secrets directories. Boundary
- 06 Turn on the sandbox and list your credential files and variables in sandbox.credentials. Boundary
- 07 Set CLAUDE_CODE_SUBPROCESS_ENV_SCRUB=1. Boundary
- 08 Read a new clone's .claude/ and .mcp.json before you accept trust, and run its setup yourself. Boundary
- 09 Check who can read ~/.claude. Disk
- 10 Set cleanupPeriodDays on purpose: how long Claude Code keeps history. Disk
- 11 Run uvx ranwhat check now, and again after anything odd. Audit
- 12 Log commands from now on, with a hook or OpenTelemetry. Audit
- 13 Replace broad tokens with fine-grained, restricted or agent-tagged ones. Tokens
- 14 Rotate anything found in a transcript, then mask it. Leaks
What ranwhat covers, and what it does not.
ranwhat is the after-the-fact half: what already ran, what is on disk, and what the credentials allow.
Where each fact comes from.
- Claude Code: Securitycode.claude.com
- Claude Code: Choose a permission modecode.claude.com
- Claude Code: Configure permissionscode.claude.com
- Claude Code: Settings referencecode.claude.com
- Claude Code: Sandboxed Bash toolcode.claude.com
- Claude Code: Environment variablescode.claude.com
- Claude Code: Data usagecode.claude.com
- Claude Code: The .claude directorycode.claude.com
- Claude Code: Hooks guidecode.claude.com
- Claude Code: Hooks referencecode.claude.com
- Claude Code: Monitoringcode.claude.com
- Claude Code: Tools referenceWebFetch and its preapproved domains
- Claude Code: Changelog2.1.162, 2.1.285
- GHSA-fg94-h982-f3mmCVE-2026-54316, WebFetch, fixed in 2.1.163
- Novee: the Hugging Face channelnovee.security
- GHSA-vp62-r36r-9xqpCVE-2026-39861, fixed in 2.1.64
- GHSA-q5hj-mxqh-vv77CVE-2026-40068, fixed in 2.1.84
- GHSA-7835-87q9-rgvvCVE-2026-55607, fixed in 2.1.163
- Accomplish AI: Beltdownfixed in 2.1.247
- GHSA-gfvf-j8jh-jxxwCVE-2026-103012, fixed in 2.1.260
- 0DIN: Clone this repo0din.ai
- Mitiga: MCP token theftmitiga.io
- Koi Security: PhantomRavenarchived copy
- npm: ignore-scriptsdocs.npmjs.com
- MCP: Security best practicesmodelcontextprotocol.io
- GitHub: Managing personal access tokensdocs.github.com
- Stripe: Model Context Protocoldocs.stripe.com
- Stripe: API keysdocs.stripe.com
- AWS: Refine permissions using last accessed informationdocs.aws.amazon.com
- AWS: Update access keysdocs.aws.amazon.com
- ranwhat: READMEgithub.com