BoundBench

Cherry Studio

Desktop AI productivity studio with autonomous agents and MCP

github.com/CherryHQ/cherry-studio · 2026-10-05 · 50d69b6

Defense-in-depth score

2.0 / 10

Minimal

Cherry Studio ships a carefully engineered approval layer for its agents: a per-call card that shows the exact shell command, a central policy registry for its own tools, path-containment checks, and hard blocks on some destructive and global-install commands. Around that layer, the agent runs shell commands directly on the host with the user's full login environment and no sandbox, and tools run by sub-agents, scheduled tasks and chat-channel turns are allowed without asking. Web fetch, browser control (including running page scripts in a persistent browser profile) and task scheduling are auto-approved, so content the agent reads can drive data out and plant follow-up work. Workspace plugin folders and project settings files are loaded without a trust prompt.

Key gaps (4)

  1. Shell commands requested inside a sub-agent, or in scheduled and channel turns, are allowed without approval, so the most powerful action path can skip the gate. C2 · Approval gates
  2. Agent-run code executes on the host with no sandbox and the user's full login-shell environment. C4 · Code-execution isolation
  3. A hijacked agent can exfiltrate through auto-approved fetch and browser tools and run shell commands through sub-agents or scheduled tasks, with no human involved. C5 · Untrusted input blast radius
  4. Workspace .claude/plugins folders and project/local settings files are loaded without a trust decision. C6 · Memory, context & configuration integrity

Criteria

C1 Identity & least privilege

Minimal 0.05 / 1.00

The agent runs as the user with the user's whole login-shell environment: Cherry Studio starts the Claude Code process with every variable from the user's shell (cloud keys, tokens and anything else exported there), and stdio MCP servers get the same environment. A short block-list only stops an agent's own settings from overriding a few model and runtime variables. There is no per-tool credential, and no authorization check sits between a tool call and the user's ambient authority, so a hijacked agent acts with everything the user can reach from a shell, plus any sites signed into the in-app browser.

C2 Approval gates

Minimal 0.25 / 1.00

Cherry Studio has a real approval system. Shell commands that need approval appear on a card showing the exact command, Cherry's own tools are classified in one policy registry as auto, required or mode-dependent, and a guard table adds deny and ask rules, such as blocking global package installs and sqlite writes to the app's own database. But approval does not cover everything. Tool calls made inside sub-agents (the sub-agent tool is always available) and in scheduled or chat-channel turns are allowed without asking. Browser actions, including running page scripts, plus web fetch and task scheduling, are auto-approved. The built-in assistant ships in accept-edits mode, and project settings files can add their own allow rules.

C3 Tool & action scoping

Minimal 0.35 / 1.00

Some tools validate their inputs well. File tools resolve paths through symlinks and ask before touching anything outside the workspace and agent data folder, web fetch refuses addresses that resolve to private or local networks, and guard rules block global installs and writes to the app's own database. But the shell tool takes any command string, reads outside the workspace stay silent by design, and the default tool set includes shell, file writes and a scriptable browser, each of which can be switched off individually.

C4 Code-execution isolation

Minimal 0.00 / 1.00

The default Claude Code agent runtime runs shell commands directly on the host as the user, with no container or OS sandbox, and with the full login-shell environment, so any credential in that environment is visible to executed code. A separate, optional agent runtime uses third-party sandbox packages, but it is not the default and was not reviewed. If agent-run code goes wrong, it has host-level access.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing structurally limits what content the agent reads can make it do. Web pages, browser pages, files and MCP results enter the conversation like any other tool output, and fetching URLs and driving the browser are auto-approved, so data can leave without a person seeing it. A hijacked agent can also delegate shell work to a sub-agent, or schedule a task with its own prompt, and both run without approval.

C6 Memory, context & configuration integrity

Minimal 0.10 / 1.00

Agents keep persistent memory files (persona, user profile and facts) that the model can rewrite through an auto-approved memory tool or file edits, and they are loaded back into the system prompt in every later session. The workspace's instruction files are also loaded automatically. For a user-chosen workspace folder, plugins under its .claude/plugins directory and its project and local settings files are loaded without any trust prompt; plugins can carry hooks and MCP servers.

C7 Third-party extensions

Minimal 0.20 / 1.00

MCP servers are added by the user from presets or by hand, often as npx commands that fetch the latest package at each launch, with no version pin or integrity check; only managed skills are hashed into an installation baseline. The agent can also ask to install MCP servers and command-line tools, and Cherry routes those requests through per-call approval, and an MCP server the agent registers stays inactive until the user turns it on. Plugins found in the workspace's .claude/plugins folder are loaded with no consent step. Every stdio server runs as the user with the full login-shell environment.

C8 Secrets & sensitive-data protection

Minimal 0.40 / 1.00

Cherry Studio redacts secrets in its logs and crash reports. Crash reporting and analytics send nothing until the user accepts the current privacy policy, and channel replies are scrubbed of known secret patterns. Provider API keys are stored as plain JSON in the local database. The agent runs a shell with the user's full environment, so long-lived keys from that environment, and the app's own stored keys, are readable by agent-run code.

C9 Audit & traceability

Minimal 0.45 / 1.00

Agent conversations, including tool calls and their inputs and results, are stored in the app's local database, and the Claude Code runtime keeps its own session transcripts in Cherry's config directory for a retention period. Detailed tracing exists only in developer mode. Records are local and written by the same process tree the agent runs in, without tamper evidence.

C10 Limits & kill switch

Minimal 0.20 / 1.00

The agent runtime is given no turn limit and no spending cap. What exists is a breaker that blocks a shell command repeated five times with identical output, a 60-second default timeout on MCP tool calls, and a 100-call cap on plain chat assistants. Pressing stop closes the session and its Claude Code process, but scheduled tasks created by the agent keep firing.