BoundBench

MetaGPT

Multi-agent framework that simulates a software company (team leader, product manager, architect, engineer, data analyst) to turn a one-line requirement into code.

github.com/foundationagents/metagpt · 2026-10-04 · 11cdf46

Defense-in-depth score

0.8 / 10

Minimal

MetaGPT runs model-written shell commands and notebook code directly on your machine, as you, with your full environment and with no approval step. Its Engineer can write any file, browse any URL, install packages it picks, and open pull requests using your ~/.git-credentials. Anything injected through a web page or file can therefore steal keys and act irreversibly. Run it only inside a disposable VM or container with no credentials you care about.

Key gaps (6)

  1. The default shell tool runs every model command on the host with no approval gate. C2 · Approval gates
  2. The Engineer's bash shell inherits the full process environment and the PR tool reads ~/.git-credentials, giving a hijacked agent the user's whole authority. C1 · Identity & least privilege
  3. Model-generated code executes in a same-user host shell and local Jupyter kernel with credentials in the environment; no sandbox exists. C4 · Code-execution isolation
  4. Tool outputs (including web pages) enter memory as user messages, and a hijacked agent can exfiltrate and act irreversibly unattended. C5 · Untrusted input blast radius
  5. A config/config2.yaml in the current working directory is merged silently on PyPI installs and can enable tools and persistent memory. C6 · Memory, context & configuration integrity
  6. The analysis prompt directs the model to pip-install packages it chooses through the host shell, running unverified install code with full credentials. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

MetaGPT has no identity of its own: every tool runs with the full authority of the OS user who launched it. The persistent shell given to the Engineer receives a copy of the entire process environment (including any LLM, cloud or Git tokens), and the pull-request tool reads the user's ~/.git-credentials file directly. There is no authorization layer between a model-chosen command and these credentials, so a hijacked agent holds everything the user holds.

C2 Approval gates

Minimal 0.00 / 1.00

There is no approval step anywhere in the default flow. The RoleZero dispatcher executes whatever command name the model emits if it appears in the tool map, including raw shell commands, notebook code execution, file writes anywhere on disk, and opening pull requests. An ask_human tool exists, but only the model decides to call it, so it is not a gate. A plan-review prompt exists for the older DataInterpreter, but RoleZero roles force it off.

C3 Tool & action scoping

Minimal 0.00 / 1.00

The default Engineer gets a raw bash shell, a notebook code executor via the DataAnalyst, an editor that writes any absolute path, a browser that opens any URL, and a pull-request tool. Arguments are passed through unvalidated: the editor only joins relative paths onto its working directory with no containment check, and the only shell filter blocks the strings 'run dev' and 'serve ' to steer the model towards the deployer, not for safety. Every tool is enabled by default for its role.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Model-written code runs directly on the host. The Engineer's shell is a same-user bash subprocess started with a full copy of the environment, and the DataAnalyst's notebook executor starts a local Jupyter kernel in the same user account. The package contains no container, sandbox, or seccomp backend on these paths. The shipped Dockerfile only packages MetaGPT itself and is not used as a per-execution sandbox.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Agents browse arbitrary URLs and read files, and every tool result is appended to memory as a user-role message, the same standing as the human's request. Nothing marks content as untrusted or restricts later actions after untrusted content is read. A hijacked Engineer can then read secrets from its environment, send them anywhere through the shell or browser, and take irreversible actions, all without a human in the loop.

C6 Memory, context & configuration integrity

Minimal 0.10 / 1.00

Default memory is in-process and session-scoped, and long-term memory and the experience pool are off by default. However, when MetaGPT is installed from PyPI its project root falls back to the current working directory, and it merges config/config2.yaml from that directory into its settings without asking. A file in the directory where the user runs MetaGPT (or one the agent itself writes there via its unrestricted shell) can enable extra tools such as search, switch on persistent long-term memory or the experience pool, or move the workspace, and that persists across sessions.

C7 Third-party extensions

Minimal 0.00 / 1.00

MetaGPT has no plugin or MCP system, but its code-writing prompt tells the DataAnalyst to install any missing package through the Terminal tool, so the model chooses and installs packages from PyPI on its own. Those installs run package install scripts on the host, unpinned and unverified, with the user's full environment and credentials. No consent is asked.

C8 Secrets & sensitive-data protection

Minimal 0.00 / 1.00

API keys live in plaintext YAML or environment variables with no masking type or redaction anywhere in the package. The default log file captures DEBUG output, which includes every full message sent to the LLM, and those prompts carry whatever tool output the agent saw (for example an 'env' command). Every shell command receives the whole environment, so long-lived keys are one 'printenv' away from the model and its provider. No third-party telemetry was found.

C9 Audit & traceability

Minimal 0.33 / 1.00

MetaGPT writes a loguru log to stderr and to a logs/ directory under the project root. Each RoleZero step logs the parsed command list and its outputs at INFO level, and full LLM prompts at DEBUG, which gives a rough reconstruction of what each role did. The record is free text, carries no notion of who approved what (there are no approvals), and sits next to the workspace where the agent's own shell can edit or delete it.

C10 Limits & kill switch

Minimal 0.38 / 1.00

There are a few limits: each RoleZero role stops after 40-50 actions, the team runs at most 5 rounds by default, the $3 'investment' budget is checked between rounds, and notebook cells time out after 600 seconds. These are coarse: the budget is only checked between rounds and ignores models missing from the price table, a role that hits its action limit asks the human whether to reset it, and the Terminal tool has no timeout at all. Background ('daemon') shell commands keep running after the agent stops.