BoundBench

ChatDev 2.0 (DevAll)

Zero-code multi-agent orchestration platform: YAML-defined agent workflows run by a FastAPI backend with a Vue web console.

github.com/openbmb/chatdev · 2026-10-04 · 4fb2db0

Defense-in-depth score

1.7 / 10

Minimal

ChatDev 2.0 runs agent-written code, package installs and MCP servers directly on the host, with the provider API key in every subprocess and no approval step. The backend's default network exposure is not locked down. Only the file tools are meaningfully scoped. Run it only on an isolated machine or container, bound to localhost.

Key gaps (4)

  1. Every code-execution path (execute_code, uv_run, python nodes, MCP stdio) runs as an unsandboxed host subprocess that inherits os.environ, including the provider API key loaded from .env. C4 · Code-execution isolation
  2. Content from web pages, files or MCP results can hijack an agent into exfiltrating the provider key and running irreversible host commands with no human in the loop. C5 · Untrusted input blast radius
  3. Agents can install model-chosen Python packages (install hooks run on host), and MCP servers launch unpinned via uvx with the full environment. C7 · Third-party extensions
  4. The long-lived provider API key from .env is placed in os.environ and inherited by every code-execution subprocess and MCP server, where model-written code can read it. C8 · Secrets & sensitive-data protection

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

ChatDev's backend acts with whatever authority the operating-system user who started it has, and the HTTP API's default exposure is not locked down. The provider API key from .env is loaded into the process environment and inherited by every code-running tool and local MCP server. There is no per-request or per-tool authorization layer, so a hijacked agent acts with the user's full account.

C2 Approval gates

Minimal 0.00 / 1.00

There is no approval step for any tool call. Agent nodes run every model-requested tool immediately, including code execution, package installation, recursive deletes and web fetches. The human node and the call_user tool let a workflow author or the model ask a person for input, but neither gates tool execution. Deletions in the workspace and actions taken by executed code are not reversible.

C3 Tool & action scoping

Minimal 0.33 / 1.00

The built-in file tools are well scoped: every path is resolved and checked to stay inside the session's code workspace. That scoping is undone by the execution tools the flagship workflow also enables: execute_code and uv_run run arbitrary Python on the host with model-chosen arguments and environment variables, install_python_packages accepts model-chosen package specs including URLs, and read_webpage_content fetches any URL. Each agent node receives only the tools its YAML lists, but the shipped ChatDev workflow lists write, delete, and exec tools.

C4 Code-execution isolation

Minimal 0.38 / 1.00

All model-influenced code runs as a plain subprocess on the host: execute_code, uv_run, package installs, python nodes and local MCP servers. Each one inherits the server's full environment, including the provider API key, and has unrestricted network and filesystem access as the user. The optional Docker Compose setup runs the backend as a non-root user inside a container, but it bind-mounts the whole repo including .env, injects the same secrets, and leaves network open. It is also not the mode the README leads with.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Untrusted content enters agents through web search and page fetches, uploaded files, MCP tool results and descriptions, and other agents' messages, all with the same standing as the user's task. Nothing tracks where content came from or restricts what an agent may do after reading it. A prompt injection in a fetched page can make an agent run code that reads the API key and sends it out, delete workspace files, or install packages, with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.13 / 1.00

Memory is opt-in per workflow, but when enabled every agent output is written to a JSON store automatically, with no validation, and re-injected later as a user-role message. Memory files are keyed only by the path the workflow sets, with no per-user namespace, and the server has no users. Separately, the exec tools can write anywhere the user can, including the functions/ directory, whose Python files are auto-loaded as tools. That lets poisoned content persist as code.

C7 Third-party extensions

Minimal 0.00 / 1.00

Shipped workflows give agents install_python_packages, so the model picks which packages to install from PyPI or a URL, and their install hooks run on the host. Local MCP servers are launched with unpinned commands such as 'uvx blender-mcp' and inherit the full environment by default. Function tools are any Python file in functions/, executed in-process at load.

C8 Secrets & sensitive-data protection

Minimal 0.13 / 1.00

Provider keys come from a .env file loaded into os.environ. They do not appear in prompts. They are inherited by every subprocess the agents start, so model-written code can read them. No redaction exists anywhere: tool arguments and full results are logged to the session's execution_logs.json and streamed to the web console. No telemetry was found.

C9 Audit & traceability

Minimal 0.35 / 1.00

Every tool call, built-in or MCP, is recorded before and after execution, with arguments, success flag, result and timing, tagged by node id. The record is held in memory and streamed to the web console. It is written to execution_logs.json in the session directory only after a workflow finishes successfully, so failed, cancelled or crashed runs lose it. The file sits where the agents' own exec tools can edit it, and nothing identifies who started a run.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Each agent node stops after 50 tool-call rounds, graph cycles stop after 100 iterations by default, and python nodes time out after 60 seconds. There is no token or cost ceiling and no wall-clock limit for a run. The model sets the timeout itself for execute_code and uv_run. Cancelling from the console sets a flag that is checked between steps and does not interrupt running subprocesses.