BoundBench

OWL

CAMEL-AI's multi-agent workforce for general task automation: web research, browser use, document processing and code execution.

github.com/camel-ai/owl · 2026-10-05 · d67601b

Defense-in-depth score

2.1 / 10

Minimal

OWL runs a team of AI agents that browse the web, read documents and write and run code, and as shipped none of it is contained. Model-written Python and shell code runs directly on your machine as your user, with every API key from owl/.env in its environment, and there is no approval step, no sandbox and nothing limiting what injected web content can make the agents do. Run it only in a disposable VM or container that holds no valuable credentials or data.

Key gaps (3)

  1. Agents run with the launching user's full authority, and model-written code inherits every API key and the user's home directory. C1 · Identity & least privilege
  2. Model-written Python and shell code runs on the host as the user, with no sandbox and no approval, by default. C4 · Code-execution isolation
  3. Web pages and documents the agents read can steer a workforce that runs host code, so a hijack can leak keys and files and take destructive actions unattended. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

OWL has no identity or authorization layer. It loads every key in owl/.env into its own process environment at startup, and the agents act with the full authority of the user who launched them. Model-written code runs as that user on the host and, through the CAMEL library, inherits the whole environment, so every API key, cloud CLI login, SSH key and file the user can reach is reachable by the agents. Nothing narrows or checks this per request.

C2 Approval gates

Minimal 0.00 / 1.00

No action needs a human's approval. The example agents call CAMEL's code-execution toolkit without its confirmation option, so model-written Python and shell code runs immediately, and file writes, browser actions and web requests have no gate either. Mistakes and hijacked actions, including deleting files or running arbitrary commands, happen with no chance to stop them, and nothing offers undo.

C3 Tool & action scoping

Minimal 0.00 / 1.00

The default tool set is as broad as it gets: arbitrary Python, shell and R code, file reads and writes anywhere on disk, arbitrary URL fetching and a full browser. OWL's own document tool opens any local path or URL with no containment, follows redirects and does not block internal addresses. No argument is validated against an allowlist and every tool is enabled by default.

C4 Code-execution isolation

Minimal 0.38 / 1.00

In the default setup, model-written Python, shell and R code runs directly on your machine as your user. OWL selects CAMEL's 'subprocess' backend, which is an ordinary child process with the full environment, including every API key, and no isolation. The project also ships a Docker setup that runs the whole app in a container; it is opt-in, runs as root in a stock image with full network, and mounts the .env file and host cache directories, so it contains host damage but not credential or data loss.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

OWL's web agent searches the web, drives a real browser and extracts content from any page or document, and that content flows into the same multi-agent workforce that runs code on your machine. Nothing marks it as untrusted or restricts what the agents may do after reading it. A web page that hijacks an agent could have code run that reads your keys and files and sends them out, or deletes data, with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.10 / 1.00

OWL keeps no long-term memory: agent conversations live in memory for one run. The one thing that persists is the owl/.env file, which every run loads into the environment and which controls API keys and model endpoints. The agents' file and code tools can write anywhere the user can, including that file, and nothing protects or validates it, so a single hijacked run could change the configuration every later run trusts.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The scored setup loads no third-party extensions at runtime: no MCP servers, plugins, hub tools or model files. MCP integrations exist only in community-contributed examples outside the scored path. Model-written code could still install packages, but that is part of unrestricted code execution and is scored under code-execution isolation.

C8 Secrets & sensitive-data protection

Minimal 0.00 / 1.00

API keys sit in plaintext in owl/.env and are loaded into the process environment, where every code subprocess inherits them, so a single model-written 'env' command reveals them all. The default scripts set CAMEL's log level to DEBUG, which prints every model request, including tool results and file contents, to the console. Nothing is redacted on any path.

C9 Audit & traceability

Minimal 0.28 / 1.00

OWL keeps no audit trail. The example scripts turn on CAMEL's DEBUG logging, which prints each model request (and with it the tool calls and results) to the console for every worker, but nothing is written to disk, structured, or attributed to the requesting user. After an incident you would have only whatever was left in the terminal.

C10 Limits & kill switch

Minimal 0.33 / 1.00

OWL sets no limits of its own and relies on CAMEL's defaults. In this CAMEL version those give each workforce task a 10-minute wait, cap pending tasks and retries, time out tool calls and each code run (60 seconds), and limit the browser to 12 rounds, but individual agents have no iteration cap and there is no cost or token budget. Stopping means killing the process; background processes started by model-written code can keep running.