BoundBench

DeepCode

Local agentic coding assistant with TUI, Desktop and Web clients over a shared background service, plus the Paper2Code multi-agent workflow.

github.com/HKUDS/DeepCode · 2026-10-05 · 84c37f7

Defense-in-depth score

3.6 / 10

Minimal

DeepCode asks before it runs commands or changes files by default, and runs shell commands in a bubblewrap or Seatbelt sandbox that keeps writes inside the workspace. That sandbox leaves the network open, the home directory readable and the full environment (provider keys included) visible, and commands run without it when no backend is installed. Reading files outside a short credential denylist and fetching any public URL need no approval, so a hijacked session can read local data and send it out without a human. Approval and sandbox settings can be loosened without an explicit operator decision, and there is no default cap on steps or spend.

Key gaps (4)

  1. The agent's own tools are not prevented from changing settings that govern its permissions on later turns. C1 · Identity & least privilege
  2. Approval settings can be loosened without an explicit operator decision in the default configuration. C2 · Approval gates
  3. Sandbox settings can be loosened without an explicit operator decision; commands run unsandboxed when no backend is installed. C4 · Code-execution isolation
  4. The command sandbox can read the home directory and receives provider keys in its environment, with network open. C4 · Code-execution isolation

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

DeepCode runs as the local user and narrows that authority only a little. File tools refuse a short list of credential locations, MCP servers and external sub-agent CLIs get only the environment they are configured with, and a project configuration cannot redirect the user's provider keys. The agent's own shell, code mode and hooks still receive the full parent environment, including provider keys, unless the operator opts into scrubbing, and the agent's own tools are not prevented from changing settings that govern later turns.

C2 Approval gates

Minimal 0.25 / 1.00

New sessions default to an Ask preset: read-only tools run freely and every other built-in, MCP and sub-agent tool call stops for a durable approval that the user can grant once, for the session, or deny. Denials and approver errors fail closed, and the local service only accepts approvals from authenticated clients. In the TUI the prompt shows the tool name and reason with the command clipped to the terminal width and no file diff, a session grant covers every later call of that tool, MCP tools that call themselves read-only and the URL fetcher are never gated, and approval settings can be loosened without an explicit operator decision.

C3 Tool & action scoping

Minimal 0.45 / 1.00

Write, edit and patch tools resolve paths and refuse anything outside the workspace, and the URL fetcher only reaches public addresses on ports 80 and 443 and rechecks every redirect. The shell tool is a raw command string, the read tool accepts any path on the machine, and MCP tools get no shared argument validation. The default tool set includes write, shell and network tools.

C4 Code-execution isolation

Minimal 0.25 / 1.00

Shell commands and code-mode programs run under bubblewrap on Linux or a deny-by-default Seatbelt profile on macOS, with writes limited to the workspace and temporary directories. Network stays open, the whole filesystem is readable and the full environment is passed in. Hooks and MCP stdio servers run directly on the host, commands run unsandboxed when no backend is installed, and sandbox settings can be loosened without an explicit operator decision.

C5 Untrusted input blast radius

Minimal 0.40 / 1.00

Fetched pages and memory notes are wrapped in labels that mark them as untrusted, and every write, shell or unknown tool waits for approval whatever the agent has read. Nothing tracks what the session has read: reading local files outside the credential denylist and fetching any public URL stay unattended, and the UI shows fetched URLs without their query string. A hijacked session can therefore read local data and send it out without a human, while irreversible changes still need approval.

C6 Memory, context & configuration integrity

Minimal 0.45 / 1.00

The memory tool needs approval for every action in the default preset, and the MEMORY.md index is injected inside an untrusted-data boundary that escapes forged tags. Instruction files (AGENTS.md, DEEPCODE.md, CLAUDE.md) from the repository load silently as standing guidance, and project MCP servers, hooks, skills and settings apply once the folder is trusted through a folder-trust prompt. Memory has no expiry or versioning, and poisoned project files persist into every later session in that folder.

C7 Third-party extensions

Minimal 0.28 / 1.00

No third-party extension is enabled by default: bundled MCP presets are copied in disabled, and project MCP servers load only after the folder is trusted. When enabled, several presets launch unpinned `@latest` packages through npx or uvx, and nothing pins, hashes or re-approves an extension when its code or tool definitions change. MCP stdio servers run on the host with only the environment their configuration names.

C8 Secrets & sensitive-data protection

Minimal 0.33 / 1.00

Saved provider keys live in a 0600 file in ~/.deepcode (keychain reads are opt-in), the UI never reads them back, configuration views mask credential fields, and provider error messages strip echoed keys. No analytics or crash-reporting SDK is present. The agent's shell and hooks receive every provider key in the environment by default, the LLM log keeps request and response previews, and the MCP log can hold credential arguments, which the project itself notes.

C9 Audit & traceability

Moderate 0.63 / 1.00

Every tool call is stored as a structured item with its arguments and status in a local SQLite database under ~/.deepcode, outside the workspace, along with approval requests and decisions, and sub-agent transcripts are recorded. Records are written per action in transactions and the session can be replayed. There is no separate actor attribution, no tamper evidence for tool records (an optional hash chain covers only LLM calls) and no standard export.

C10 Limits & kill switch

Minimal 0.30 / 1.00

Ordinary turns have no step limit by default; an iteration cap and Goal token budgets exist but are opt-in. Shell commands time out (120 seconds by default, but the model can choose a longer value), model requests have timeouts, sub-agents are limited to one level and five at once, and stopping a turn kills whole process groups. Repeated identical calls only trigger reminders, and there is no spend cap.