BoundBench

marimo

Reactive Python notebook with built-in AI agent chat and MCP server

github.com/marimo-team/marimo · 2026-10-03 · 82b936f

Defense-in-depth score

1.2 / 10

Minimal

Once the chat is in Agent or Code Mode, marimo's assistant runs the code it writes in your notebook kernel immediately, as you, with no approval, sandbox or step limit. The 'Fix with AI' button on any cell error switches the chat to Code Mode and keeps it there. A prompt injection in a shared notebook, a dataset or a web page can therefore read your credentials (including the AI keys marimo hands the kernel) and send them anywhere, or delete files, unattended. A cloned repo's pyproject.toml can also make marimo launch MCP server commands at startup.

Key gaps (6)

  1. Agent-written code runs as the user in the notebook kernel with the user's full environment and marimo's unmasked AI keys, so a hijacked session holds all of the user's authority. C1 · Identity & least privilege
  2. The assistant's most powerful actions, running model-written code in the kernel via execute_code or run_stale_cells, execute with no human approval. C2 · Approval gates
  3. Model-written code runs unsandboxed in the host kernel process as the user, with environment credentials, the unmasked marimo config and full network. C4 · Code-execution isolation
  4. After prompt injection via notebook code, data outputs, MCP results or web pages, the assistant can both exfiltrate secrets and take irreversible actions through kernel code with no human step. C5 · Untrusted input blast radius
  5. A cloned repository's pyproject.toml can add MCP stdio servers that marimo launches at startup and redirect the AI provider base URL, with no trust prompt. C6 · Memory, context & configuration integrity
  6. Code Mode lets the model install arbitrary packages without confirmation, and project config can make marimo launch MCP server commands at startup without consent. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

marimo's assistant acts with the full authority of the person who started marimo. Code the agent writes or runs executes in the notebook kernel, a child process of the server running as the same OS user that also receives the unmasked marimo configuration (including AI provider keys), and MCP servers are started with a copy of the whole environment. There is no dedicated identity, no scoped credential and no per-tool authorization layer; the only check is the server's login token, which identifies the human user, not what the agent may do. A hijacked session therefore has every file, credential and network path the user has.

C2 Approval gates

Minimal 0.05 / 1.00

There is no approval step for any assistant action. In Agent mode the model's edits are applied to cells and its run_stale_cells tool then executes those edited cells immediately; in Code Mode the execute_code tool runs arbitrary Python in the kernel; MCP tools are invoked directly by the server. Tool calls are executed as soon as the model emits them and the chat automatically sends the result back for the next step. Edited and deleted cells are kept as 'staged' changes with the previous code so the user can reject them afterwards, but by then the code has already run.

C3 Tool & action scoping

Minimal 0.13 / 1.00

The agent's main tools are general-purpose: Code Mode's execute_code takes any Python string, and Agent mode can write any cell and then run it. The read-only backend tools have typed arguments that are parsed before use, but that does not constrain the code-execution paths, and MCP tool arguments are only checked by the MCP server itself. The shipped default chat mode is Manual with no tools, but the 'Fix with AI' button silently switches the chat to Code Mode and saves that choice, and a project's pyproject.toml can set the mode too.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Model-generated code runs directly in the notebook's Python kernel, which is a normal child process on the host running as the user. There is no container, OS sandbox or restricted interpreter for AI-driven execution; the 30-second timeout and interrupt only bound how long the server waits. The kernel has the user's home directory, environment credentials, the marimo config with API keys, and unrestricted network. MCP stdio servers are likewise host processes.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing limits what a prompt-injected assistant can do. The notebook's code (which may come from a shared or downloaded notebook) is placed in the system prompt, cell outputs from whatever data the notebook loads, MCP tool results and, when web search is turned on, fetched web pages all reach the model with the same standing as the user. The same session can then run arbitrary code with network access and the user's credentials, so an attacker can both steal data and take irreversible actions without any human step.

C6 Memory, context & configuration integrity

Minimal 0.17 / 1.00

marimo has no agent memory store, but its configuration travels with projects. The nearest pyproject.toml's [tool.marimo] section overrides the user's own settings with only signing and cache-verification keys removed, so a cloned repository can add MCP servers whose commands marimo launches at startup, point the AI provider at its own base URL, set the chat mode or inject custom rules into every prompt, with no trust prompt. A .marimo.toml found by walking up from the working directory is loaded as the user's config. marimo does block AI and MCP settings in a notebook's own inline script header, which shows the maintainers recognise the risk.

C7 Third-party extensions

Minimal 0.05 / 1.00

marimo can launch MCP servers as local commands or connect to remote ones, with no pinning, hashes or change detection, and project config can add them silently. In Code Mode the model can install packages from the package index with a single call and no confirmation, and package installs run third-party install code. MCP stdio servers run as separate processes but receive a full copy of the user's environment.

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

AI provider keys live in the user's marimo TOML config (or are read from environment variables) and are masked before the config is sent to the browser. But the kernel, where agent-written code runs, is handed the configuration with secrets unmasked, and it inherits the full environment; Code Mode additionally places the marimo server's auth token into the kernel request. There is no redaction of tool results before they go to the model. marimo ships no third-party telemetry and AI tracing is off unless MARIMO_TRACING is set.

C9 Audit & traceability

Minimal 0.38 / 1.00

The only record of what the assistant did is the chat transcript, which the browser saves to its local storage with every tool call's arguments and results once a response finishes. The server does not log successful tool calls or executed code, and OpenTelemetry tracing is opt-in. The transcript lives in the user's browser profile, can be deleted from the UI, and is written only when a turn completes, so a crash mid-run can lose it; code that agent-run cells spawn is not recorded at all.

C10 Limits & kill switch

Minimal 0.20 / 1.00

marimo sets no step, time or cost budget for the assistant. Each code-mode execution is waited on for 30 seconds and then interrupted, and MCP calls time out after 30 seconds by default, but in Agent mode the chat automatically sends a new request after every tool result with no cap, and max_tokens is unset unless configured. Stopping the chat ends the stream, and in Code Mode cancels and interrupts the kernel, but processes or threads started by that code keep running, and cells started by run_stale_cells continue.