BoundBench

PI-Desktop

Local-first Electron desktop workspace for AI coding agents, built on the pi agent harness with a Rust host core, plugins, MCP servers and subagents.

github.com/vastsa/pi-desktop · 2026-10-04 · 831b66d

Defense-in-depth score

2.7 / 10

Minimal

PI-Desktop has a real, centralized approval gate: in the default 'ask' mode every Write, Edit, Bash and MCP call stops for a per-call approval card. Behind that gate there is no sandbox. Approved commands run as you, with your full environment and network. Repository-supplied configuration is not integrity-protected, and rendered model output is not fully confined.

Key gaps (2)

  1. Shell commands and MCP servers run unsandboxed as the desktop user with the full inherited environment (credentials, home directory, network). C4 · Code-execution isolation
  2. MCP stdio servers are spawned with the full process environment and no verification. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.17 / 1.00

PI-Desktop runs every action with the desktop user's full authority. The Bash runner and the agent sidecar inherit the whole process environment, so any cloud, GitHub or SSH credential the user has is reachable from a shell command. One real narrowing exists: the sidecar can only call an allowlist of host methods and can resolve only the provider key bound to its own session, and plugin processes get a scrubbed environment. User-configured and project MCP servers, however, also receive the full environment.

C2 Approval gates

Moderate 0.50 / 1.00

Every built-in, MCP, plugin and subagent tool call is routed through one host-core permission gate, and the default mode is 'ask': Write, Edit and Bash need a per-call approval card showing the arguments, MCP tools always need approval, and denials are first-class. The gate does not cover every path and its approval card does not always show the exact call, plugins can self-declare a tool as low-risk to get it auto-approved, and 'Allow for session' approves a tool by name for all later arguments. File edits have review snapshots for rollback; shell side effects do not.

C3 Tool & action scoping

Minimal 0.45 / 1.00

File tools are well scoped: paths are resolved through symlinks and must stay inside the project or the session scratch folder, and any outside path needs an explicit approval. But the default tool set always includes a raw Bash tool that accepts any command string, plus write tools, and MCP tool arguments are passed through unvalidated. A misused shell command can reach the whole machine.

C4 Code-execution isolation

Minimal 0.00 / 1.00

There is no execution sandbox. Approved shell commands run through host-core's runner as the desktop user, in the project directory, with the full inherited environment and unrestricted network. The only protection is the approval prompt; once a command runs, it can do anything the user can.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

Nothing in the code tracks whether untrusted content (repository files, MCP results, web search results) has entered a session, and nothing changes once it has. Project instruction files and tool results go straight into context. Reading files in the workspace needs no approval, and an unattended exfiltration path exists. Destructive actions still need approval in the default mode.

C6 Memory, context & configuration integrity

Minimal 0.17 / 1.00

Opening a project silently trusts files the repository ships: AGENTS.md and CLAUDE.md are auto-loaded into the prompt, and other project-supplied capability files are not integrity-protected. Project memory is written only from the UI, and project-level subagents cannot raise their own permission mode, but neither offsets the repo-config path.

C7 Third-party extensions

Minimal 0.07 / 1.00

Plugins are handled carefully: they run in a separate process per plugin with a scrubbed environment, a permission broker, and marketplace packages are checked against a SHA-256 digest. MCP servers are not. A user-typed stdio server can name any command (typically an unpinned package runner) and is started with the user's full environment; project-supplied servers are not gated by a trust decision.

C8 Secrets & sensitive-data protection

Moderate 0.50 / 1.00

There is no telemetry and crash dumps stay local. Logs and the audit table pass through redaction for tokens, keys and bearer headers. Provider keys are stored AES-GCM encrypted, but the key sits in an owner-only file right next to them, so in practice this is plaintext protected by file permissions. Shell and MCP subprocesses inherit the full environment, and tool output containing secrets is sent to the model unredacted.

C9 Audit & traceability

Minimal 0.45 / 1.00

Host-core writes a structured, redacted audit row for every tool it executes or denies, including whether a prompt was shown and how long approval took, into a SQLite database in the app data directory. MCP, plugin and subagent calls pass through the same path. Some host-local tool paths are not audited, the audit row omits the arguments (only the session transcript has them), there is no approver identity, and write failures are silently ignored.

C10 Limits & kill switch

Minimal 0.33 / 1.00

The main agent loop has no step, token or cost limit. Bash commands default to a 60-second timeout but the model can request up to six hours per call, and subagents are capped at 10 running at once and six hours each. Stopping a turn kills the shell's whole process group. Scheduled tasks can keep starting new runs in the background.