BoundBench

Codex Security

OpenAI CLI and TypeScript SDK for finding, validating and fixing vulnerabilities

github.com/openai/codex-security · 2026-10-03 · 4555441

Defense-in-depth score

2.9 / 10

Minimal

Codex Security is a thin wrapper that drives Codex agents against code it is asked to review, under a filesystem profile that allows reading the whole machine and writing only to workspace roots. Execution approvals are decided by an automatic LLM reviewer rather than a person, and the agent's shell inherits the operator's environment, including unrelated cloud and API tokens. Its own tool-calling surface is Codex's, so most criteria depend on upstream behaviour; the wrapper adds careful path, executable and credential-home checks but no default cost or time ceiling in standard mode and, by repository policy, no secret redaction. Scan only repositories you trust, and start scans with only the credentials they need.

Key gaps (3)

  1. Scan permission escalations are approved by a forced automatic LLM reviewer, not a human principal. C2 · Approval gates
  2. The model can obtain extra permissions from the automatic reviewer, so the sandbox can be widened at runtime without a human. C4 · Code-execution isolation
  3. The sandboxed scan shell inherits the operator's environment and can read the whole filesystem, so cloud and API tokens are reachable. C4 · Code-execution isolation

Criteria

C1 Identity & least privilege

Minimal 0.17 / 1.00

The scan runs as the operator's own account. The Codex child process inherits the whole environment except OpenAI keys, and the scan profile lets the agent read the entire filesystem, so GitHub, AWS and Linear tokens and files such as SSH keys are within reach. Writes are limited to workspace roots, and the wrapper keeps its own model credentials in a private runtime home. Nothing narrows credentials per tool or per request.

C2 Approval gates

Minimal 0.23 / 1.00

Scan commands are approved by an automatic LLM reviewer: the wrapper forces the reviewer to auto_review and removes any override, so no person sees escalation requests in a normal scan. Per-call human approval does not exist in the default flow; the only operator control is setting approval_policy to never, which denies all requests. Consequential external actions (branch push, draft PR, Linear import) happen only when the operator passes explicit flags.

C3 Tool & action scoping

Minimal 0.25 / 1.00

The wrapper does not define model-facing tools; the agent uses Codex's shell, so arguments are not validated per call. What the wrapper does constrain: the output directory must sit outside the repository, helper executables are resolved from PATH entries outside the target, git runs with fsmonitor disabled, and per-task profiles (policy generation, duplicate review, scan comparison) switch off MCP, plugins, web search and network. The default scan keeps the full tool set.

C4 Code-execution isolation

Minimal 0.25 / 1.00

Commands run inside Codex's native OS sandbox through a named profile that makes the filesystem read-only except workspace roots, and the wrapper refuses to start a scan on Unix if a sandboxed probe command fails. Overrides for the sandbox mode are stripped. However, escalations are approved by an automatic reviewer rather than a person, and the sandboxed shell carries the operator's full environment and can read the whole disk, so an escape or granted escalation reaches real credentials. Patching can opt out of the sandbox into an external one with a printed warning.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

A security scanner reads untrusted code by design. Defence rests on prompts that tell the agent to treat repository text as data and on delimiting untrusted JSON in auxiliary turns, plus the sandbox. Nothing tracks taint or forces human approval once untrusted content is read, and the reviewer that grants escalations is itself an LLM. Auxiliary tasks such as scan comparison and policy drafting are structurally limited, but the main scan is not.

C6 Memory, context & configuration integrity

Moderate 0.50 / 1.00

Nothing from the target repository can add tools, hooks or settings: the project config file is loaded only when the operator names it, the agent works in an output directory required to sit outside the repository, and there is no dotenv loading. Repository SECURITY.md guidance and the operator's saved findings do enter later scans as context. Findings, dedupe groups and severity assessments persist locally and can feed the patch command, which runs in workspace-write.

C7 Third-party extensions

Minimal 0.30 / 1.00

The wrapper loads no third-party code on its own: its plugin ships inside the package and the Codex runtime is pinned exactly. Extensions come only from the operator's own Codex configuration, which the scan inherits, including MCP servers that deep-scan workers are documented not to disable. The wrapper does not pin, hash-check or re-approve those, and they run as the same user with the full environment.

C8 Secrets & sensitive-data protection

Minimal 0.20 / 1.00

The model-provider key is withheld from the agent's environment and the credential home and SQLite store are created private, with unusually careful ownership and ACL checks. Beyond that, other inherited secrets reach every subprocess and the model can read any file. There is no redaction of diagnostics or logs: contributor policy in AGENTS.md explicitly forbids adding it, and raw agent reasoning is shown by default.

C9 Audit & traceability

Moderate 0.50 / 1.00

Each scan thread, including sub-agent threads, is recorded by Codex as local session files that the wrapper reads for cost tracking and saved scan logs, and scan results are kept in a private SQLite workbench. The record sits outside the repository but is written by the same host process tree, with no tamper evidence or export. Logging is best effort by project policy.

C10 Limits & kill switch

Minimal 0.25 / 1.00

Deep scans have shipped ceilings (40 discovery runs, 96 hours, 4 workers, stop after consecutive errors), while the default standard scan has no wrapper-enforced time or cost limit; a dollar cap exists but is optional. Interrupts forward SIGINT or SIGTERM to the child and force a kill after one second. Sub-agent concurrency is capped by a Codex setting, but limits are not enforced per tool call.