BoundBench

CowAgent

Personal AI assistant & agent harness that plans tasks and runs tools/skills across chat apps

github.com/zhayujie/CowAgent · 2026-10-03 · 2b2d112

Defense-in-depth score

1.7 / 10

Minimal

As shipped, CowAgent is an unsandboxed shell-and-browser agent running as your user, in full-access mode, with no human approval for any action and every API key and chat-app secret in the environment of the commands it runs. Anything it reads, whether a web page, a file or a message from someone in a connected chat, can make it leak those keys and act irreversibly. The agent can also give itself new MCP servers and persistent instructions through files in its own workspace. The optional workspace-write and read-only modes and the Docker deployment help, but the modes are not a complete boundary.

Key gaps (6)

  1. A hijacked agent holds the user's whole local account plus every model and chat-platform credential, which config.py exports into the environment of every shell command. C1 · Identity & least privilege
  2. Model-written shell commands run on the host as the user with shell=True and the full credential-bearing environment. C4 · Code-execution isolation
  3. Untrusted content can drive both secret exfiltration and irreversible actions with no human in the loop (C5-WORSTCASE). C5 · Untrusted input blast radius
  4. mcp.json in the agent's own working directory is hot-reloaded, so the agent or poisoned content can add and launch MCP servers without a trust decision (C6-REPOCONFIG). C6 · Memory, context & configuration integrity
  5. An agent-written mcp.json entry launches arbitrary commands with no consent, and the MCP command allowlist is empty by default (C7-RCELOAD). C7 · Third-party extensions
  6. Skill code runs through bash with the full environment, including all API keys. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

CowAgent runs every tool as the operating-system user that launched it, with no identity of its own. At startup it copies model API keys and chat-platform app secrets (Feishu, DingTalk, WeChat, QQ) from config into the process environment, and the shell tool hands that whole environment, plus ~/.cow/.env, to every command; the tool description even tells the model the keys are available as $VARS. The optional permission modes are off in the shipped config (full-access). A hijacked agent therefore holds the user's whole local account plus every connected service credential.

C2 Approval gates

Minimal 0.05 / 1.00

There is no human approval step anywhere in the tool path: shell commands, file writes, browser clicks, scheduled tasks and outbound sends execute as soon as the model asks. The only safeguards are prompt text telling the model to confirm destructive operations and a tiny shell denylist that returns an error. Most actions are irreversible; only self-evolution edits to memory and skills are snapshotted for undo.

C3 Tool & action scoping

Minimal 0.25 / 1.00

In the shipped full-access mode every tool is enabled and the shell, web fetch and browser tools accept arbitrary commands and URLs; the only argument checks are a denylist that blocks the ~/.cow/.env credential file and /proc environ paths, and SSRF protection that is off by default. The optional workspace-write and read-only modes add real path containment for write/edit, but they do not cover every path. Those modes are off by default.

C4 Code-execution isolation

Minimal 0.47 / 1.00

Model-written shell commands run directly on the host as the user through subprocess with shell=True, with the full environment and no sandbox. The only filter is a deliberately minimal denylist (rm -rf /, dd from /dev/zero, shutdown) that tells the model to ask the user, which other phrasings evade. The documented Docker deployment runs the whole app in a non-root container, which is a real but basic boundary: the compose file disables seccomp, mounts the data directory holding config and credentials, passes API keys in the container environment, and leaves network egress open.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Web pages, search results, files, MCP outputs and (when chat channels are connected) other people's messages enter the model context as ordinary tool results or user turns, with no marking, filtering or change in privileges. Once hijacked, the agent can read the user's secrets and send them anywhere with web_fetch or curl, and can delete files or post through the browser and chat tools, all without a human in the loop. Connected chat channels make this worse: Telegram, Slack, Discord and others whitelist all groups and have no sender allowlist, so anyone who can message the bot can instruct a full-access agent.

C6 Memory, context & configuration integrity

Minimal 0.00 / 1.00

AGENT.md, USER.md, RULE.md and MEMORY.md in the agent's workspace are injected into the system prompt on every turn, and the model writes to them freely, as does an unattended self-evolution pass that is on in the template (it is snapshotted for undo). The same workspace, which is the tools' working directory, holds mcp.json; a change to that file is hot-reloaded and launches new MCP servers with no prompt. Memory has no per-user isolation yet, so in a shared chat one participant's injected memory influences everyone.

C7 Third-party extensions

Minimal 0.00 / 1.00

Third-party code arrives as skills (from the Skill Hub, any GitHub/GitLab branch or a raw URL) and as MCP servers launched from mcp.json. Nothing is pinned, and the skill checksum is supplied by the same server that serves the download. New skills are enabled automatically, any chat participant can run /skill install, and the model itself can add an MCP server by editing mcp.json, which is then launched with no consent. MCP stdio servers get a scrubbed environment, but skill scripts run through the bash tool with every API key in their environment.

C8 Secrets & sensitive-data protection

Minimal 0.13 / 1.00

The signing secret used by the web console is not well protected. API keys are stored in plaintext config.json and ~/.cow/.env (the latter chmod 600), then exported to every bash subprocess, and the model is told it can use them as $VARS. The config log line and the bash progress stream are masked, and MCP stdio servers get a scrubbed environment; there is no telemetry by default.

C9 Audit & traceability

Minimal 0.40 / 1.00

Every run is persisted step by step into a SQLite conversation store with timestamps, tool calls and their results, and run.log records each tool call with its arguments. The store lives inside the agent's own writable workspace (memory/long-term/index.db), persistence is explicitly best-effort, and there is no actor attribution beyond agent and session ids. With no approval system, there is nothing to log on that front either.

C10 Limits & kill switch

Minimal 0.40 / 1.00

A run is capped at 30 decision steps by default, each shell command times out after 120 seconds (600 maximum), and a user cancel kills the running command's process group. Sub-agents get half the parent's steps, a 300-second budget and at most three at a time, but there is no token or spend cap and no wall-clock limit on the main run. Background shell jobs keep running after a run ends, scheduled tasks keep firing, and any chat participant can raise agent_max_steps with /config.