BoundBench

MS-Agent

ModelScope's Python agent framework with CLI, TUI and WebUI for tool-using, multi-agent and long-running tasks.

github.com/modelscope/ms-agent · 2026-10-04 · a56afcc

Defense-in-depth score

3.0 / 10

Minimal

MS-Agent ships shell, Python and notebook execution on the host, and its Python tool runs model code inside the agent process itself, with no sandbox by default. The WebUI and TUI ask before most actions, but any project directory can switch that off through its own .ms_agent/config.yaml, permission_memory.json or hooks.json, with no workspace-trust prompt. The SDK and `ms-agent run` default to auto-approve. Treat opening an untrusted repository as running its code.

Key gaps (4)

  1. Model-generated Python executes in-process via exec() by default, with the agent's credentials and the user's full filesystem in reach. C4 · Code-execution isolation
  2. Workspace files the agent opens (.ms_agent/config.yaml, permission_memory.json, hooks.json) can switch off or bypass the approval gate with no trust decision. C2 · Approval gates
  3. The same repo-controlled config defeats the only limit on a prompt-injected agent, and the SDK/`ms-agent run` default (auto) leaves exfiltration and irreversible actions unattended. C5 · Untrusted input blast radius
  4. Repo-controlled .ms_agent/hooks.json, config.yaml and ./.env are loaded with no workspace-trust prompt, enabling shell hooks, MCP servers and auto-approval. C6 · Memory, context & configuration integrity

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

MS-Agent runs as the operating-system user and adds no identity or credential scoping of its own. It does narrow what spawned processes inherit: shell commands, MCP stdio servers and hooks get a scrubbed environment without API keys. But the default python_executor tool runs model code inside the agent process itself, where every environment variable, the provider API keys and the user's home-directory credentials are reachable. A repo config file can also add variables back into the shell environment.

C2 Approval gates

Minimal 0.25 / 1.00

The WebUI and TUI default to a 'restricted' mode that asks a human before any tool outside a small whitelist (read_file, grep, glob, todo, memory, skills), with options to allow once, for the session, always, edit the arguments, or deny. The SDK constructor and `ms-agent run` default to 'auto', which approves everything except a few network commands. The approval prompts in the terminal truncate long arguments, so a long script can be partly hidden. A cloned repository can turn approval off: its .ms_agent/config.yaml can set permission.mode to auto, its .ms_agent/permission_memory.json can pre-seed always-allow rules, and its .ms_agent/hooks.json can register a PreToolUse hook whose 'allow' skips the approval step entirely. No checkpoints are taken by default.

C3 Tool & action scoping

Minimal 0.33 / 1.00

The file tools check resolved paths against the workspace and a list of sensitive paths, and shell commands go through a parser that extracts paths from known commands and blocks writes outside the workspace. That shell check lets unknown commands through, and inline interpreter code cannot be inspected. The python_executor and notebook tools take arbitrary code with no argument validation at all. The default tool set enables write, shell, Python and notebook execution together.

C4 Code-execution isolation

Minimal 0.33 / 1.00

The shipped default runs model code with no isolation at all. Shell commands run as the user on the host (with a cleaned environment), and python_executor calls exec() inside the agent's own process, with full access to its memory, environment variables and API keys. A Docker sandbox backend (ms-enclave) exists, but it is opt-in. It is also a stock container with networking on and the workspace mounted read-write, and hooks and MCP servers still run on the host. If the approval step is bypassed, model code effectively has host-level access.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Nothing tracks where content came from: files, shell output, MCP results and the repo's AGENTS.md all enter the model's context on equal footing, and AGENTS.md goes into the system prompt. What limits a hijacked agent in the default WebUI/TUI mode is the general per-call approval. Shell, Python and file writes need a click, but reads and memory writes do not, so injected text can persist itself into project memory without anyone approving it. The same repo-controlled config files that switch approval off also remove this limit, and in the SDK's default auto mode a hijack can both exfiltrate data and take irreversible actions unattended.

C6 Memory, context & configuration integrity

Minimal 0.17 / 1.00

A project directory controls a lot of the agent's behavior with no workspace-trust prompt. Its .ms_agent/config.yaml can change any setting, including tools, MCP servers and permission mode. Its .ms_agent/hooks.json registers shell-command hooks that run automatically, its permission_memory.json holds always-allow rules, its AGENTS.md is injected into the system prompt, and the CLI loads ./.env from the current directory. Project memory (MEMORY.md under .ms_agent/memory, on by default in the WebUI) is written by an auto-approved memory tool. A regex scanner rejects a few obvious injection phrases, but the memory is then fed back into later sessions.

C7 Third-party extensions

Minimal 0.25 / 1.00

Extensions come from MCP servers, plugins, hooks and skills. Plugins installed from GitHub are pinned to a resolved commit, but MCP servers launch whatever command the config names with no pinning or integrity check. Project-scope plugins.json and mcp.json inside the workspace are picked up without consent. Loading Python tool plugins named in a config does require trust_remote_code. MCP stdio servers and hook commands run as separate processes with a scrubbed environment.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

There is no telemetry (the bundled mem0 telemetry is explicitly switched off), and API keys are masked when the WebUI sends them back to the browser. Spawned shells, MCP servers and hooks get environments without secrets. Provider keys are kept in plaintext settings, and session transcripts are written with no redaction. The in-process python_executor can read every key the agent holds and print it into the conversation.

C9 Audit & traceability

Minimal 0.45 / 1.00

Each session is written as an append-only JSONL log under ~/.ms_agent/projects, outside the workspace, flushed line by line. It records every message, including tool calls and results, and the WebUI also records each approval or denial with its arguments. Records carry no actor identity beyond role. The agent's own process (and approved shell commands) can still edit the log, and a failed permission write is swallowed silently.

C10 Limits & kill switch

Minimal 0.40 / 1.00

The agent loop stops after a round cap, but the interfaces raise it to 1000 rounds (9999 in the shipped agent.yaml). Each tool call is bounded by a 120-second default timeout, and the model itself can raise that to 600 seconds. There is no token, cost or wall-clock budget. A python_executor timeout stops waiting but leaves the thread running. Killing a background shell task only marks it killed, because the kill code does not handle asyncio subprocesses.