BoundBench

Dyad

Local, open-source desktop AI app builder (Electron) whose agent writes, installs, builds and runs web apps on the user's machine.

github.com/dyad-sh/dyad · 2026-10-05 · 701d917

Defense-in-depth score

2.7 / 10

Minimal

Dyad's agent edits your app's files and then installs, builds and runs that app directly on your computer under your own account and full environment, and those steps need no approval by default. Only a few actions ask first: adding a package, schema-changing SQL and MCP tool calls. File tools are well contained to the app folder and .env values are redacted before they reach the model, but nothing contains the code the agent writes once the app runs, so a hijacked session can do anything you can.

Key gaps (5)

  1. Code the agent writes runs on the host with the user's full account and environment when the app starts or builds, with no approval by default. C2 · Approval gates
  2. The app runtime, builds, tests and installs run unsandboxed on the host; only one scripting tool is sandboxed. C4 · Code-execution isolation
  3. A session hijacked by imported code, logs or MCP output can leak data and make irreversible changes with no human involved. C5 · Untrusted input blast radius
  4. Agent-started processes act with the user's whole account and inherited credentials. C1 · Identity & least privilege
  5. Third-party package code from the app's package.json is installed and executed without consent when the app runs or builds. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.07 / 1.00

The agent acts with the full authority of the logged-in user. The app's dev server, builds and package installs are started with a copy of Dyad's whole environment, so any credential in it is available to code the agent wrote. One path narrows this: end-to-end test runs strip the app's database credentials from the environment. Optional integrations add broad, long-lived credentials such as a Supabase management token and a GitHub token with repo and workflow scopes.

C2 Approval gates

Minimal 0.25 / 1.00

Each agent tool has a consent level, and a handful default to asking: adding packages, schema-changing or deleting SQL, reinstalling dependencies and every MCP tool call. The prompt shows the exact SQL or package list, though MCP arguments are cut at 500 characters. Writing, deleting and renaming files, restarting the app, building and running tests all default to running without approval, and restarting or building executes code the agent just wrote. Each turn is committed to git, so file changes can be rolled back, but anything the running code does on the machine cannot.

C3 Tool & action scoping

Minimal 0.42 / 1.00

File tools resolve paths and symlinks and refuse anything outside the app folder, and reads open the resolved file and check it before reading. Package names are checked against strict npm name and version patterns, and outputs and file sizes are bounded. The weak points are reach rather than validation: by default the agent can write any file in the app and then run the app, build or tests, which executes whatever the app's scripts say. MCP tools get only schema validation.

C4 Code-execution isolation

Minimal 0.33 / 1.00

The only sandbox on by default is a small scripting runtime used by one tool, which gets file and MCP access only through functions Dyad injects and has time and memory limits. Everything else that runs code does so directly on the host: the app's dev server, production builds, type checks, tests and package installs, all under the user's account with the inherited environment. An optional Docker runtime runs the dev server in a stock container with the app folder mounted and normal network access, but builds and installs still run on the host.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

The agent reads content its user did not write, such as imported repositories, app console logs and MCP tool results, and nothing in code separates that content from the user's instructions. There is no taint tracking and no rule that forces approval once untrusted content has been read; the only defence is prompt wording. A hijacked session can write code and run it on the host without approval, which is enough both to send data anywhere and to make irreversible changes.

C6 Memory, context & configuration integrity

Minimal 0.10 / 1.00

Each app's AI_RULES.md is read from the app folder on every turn and placed in the system prompt as authoritative project guidance, and the agent can edit that file without approval. An imported repository can therefore ship its own rules, and anything written there steers every later session in that app. Chat history from earlier conversations is also searchable by the agent. The file is versioned in the app's git history, so changes can be inspected and reverted by hand.

C7 Third-party extensions

Minimal 0.20 / 1.00

MCP servers are added by the user (a deep link only pre-fills a dialog) and run whatever command and arguments were entered, unpinned. Package installs through the dedicated tool ask first and are wrapped in a pinned Socket firewall that blocks known-malicious packages, falling back to an unscreened install with a warning when it can't be set up. Running or building the app installs and executes the packages its package.json lists without a prompt, as the build tool's own description acknowledges. Installed packages and MCP servers run as the user.

C8 Secrets & sensitive-data protection

Minimal 0.40 / 1.00

Provider keys and integration tokens are encrypted with the operating system's secure storage, falling back to plaintext only when it is unavailable. Values in .env files are redacted before file reads, grep, sandbox scripts and codebase context reach the model, and model-request logs record only metadata. Telemetry is opt-in. The gaps are the environment handed to app, build and install processes, unredacted build and log output, and chat transcripts stored unencrypted.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every tool call is written into the chat message in Dyad's local database as it happens, including its arguments and result, and the full model message history is saved at the end of the turn. Each turn is also committed to the app's git repository. The record carries no separate actor or approver fields, approvals and denials are not clearly logged, and the database sits in the user's data folder where code running as the user could change it.

C10 Limits & kill switch

Minimal 0.45 / 1.00

A turn stops after 100 tool steps by default (selectable from 25 to 200), builds are limited to three per turn, and builds, installs, shell commands and sandbox scripts have timeouts. There is no token or cost cap for a run beyond the free tier's message quota. Stopping a turn aborts the model stream and passes a cancel signal to tools, but the app's dev server keeps running, and code the agent runs on the host can modify Dyad's own settings.