BoundBench

Windmill

Open-source developer platform for scripts, workflows and internal apps, with AI agent steps that call scripts, flows, MCP servers and web search as tools.

github.com/windmill-labs/windmill · 2026-10-05 · 85248ee

Defense-in-depth score

3.3 / 10

Minimal

A Windmill AI agent step runs every tool call the model makes without human review, using the full workspace permissions of whoever the flow runs as and the author's long-lived service credentials. If content it reads (a webhook payload, an email, a web page, a tool result) carries injected instructions, it can leak data and act on connected systems unattended. Authorization checks, author-pinned tool inputs, SSRF checks on MCP servers and a working cancel path are real strengths, but in the shipped compose file tool code runs as root in a privileged worker with only process-namespace isolation, and the nsjail sandbox and token scoping are opt-in.

Key gaps (4)

  1. Tool jobs carry the runner's full workspace authority and long-lived service credentials, so a hijacked agent holds everything the runner can reach. C1 · Identity & least privilege
  2. In the shipped compose file the worker runs privileged as root and jobs get only PID-namespace isolation. C4 · Code-execution isolation
  3. A prompt-injected agent can both exfiltrate data and take irreversible actions with no human involved. C5 · Untrusted input blast radius
  4. Third-party Hub scripts run as ordinary worker jobs with the runner's job token and no extra confinement. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.42 / 1.00

Every tool an agent calls runs as a Windmill job under the identity of whoever the flow runs as, with a job token that carries that user's full workspace permissions and lasts as long as the instance's maximum job duration (seven days on a self-hosted install). Authorization is real: resources, secrets and MCP servers are loaded through the job's own permissions, so a flow cannot use what its runner cannot read. Windmill also lets an author cap each step's or tool's token to named API scopes, and a tool can never get a wider token than its agent, but that cap is off unless set. A hijacked agent therefore acts with the runner's whole workspace authority plus the long-lived credentials attached to its tools.

C2 Approval gates

Minimal 0.00 / 1.00

An AI agent step runs every tool call the model chooses straight away; there is no approval step for agent tool calls. Windmill flows do have strong human-approval steps, but they sit between flow steps and cannot be attached to an individual tool call inside an agent. Whatever the attached tools can do, including writes, messages and payments through author scripts or MCP servers, happens without a person seeing the exact call.

C3 Tool & action scoping

Moderate 0.50 / 1.00

An agent only gets the tools its author attached and cannot load new ones; a run can narrow the roster but never widen it. Inputs the author pinned are removed from what the model sees and always override the model's values, MCP tools can be limited by include and exclude lists, and MCP server addresses are checked against private and metadata addresses without following redirects. Unpinned inputs are passed to the author's code as the model wrote them, and tools are general scripts that reach whatever their credentials reach.

C4 Code-execution isolation

Minimal 0.35 / 1.00

Agent tools run as Windmill jobs on workers. In the shipped Docker Compose setup the only job isolation is a separate process namespace, and the worker container itself runs privileged and as root, so a tool job can reach the worker host and its network. Windmill ships an nsjail sandbox that limits the filesystem and drops capabilities, but it is off by default and leaves network access and the job's token in place. Jobs start with a cleared environment that adds only Windmill's own variables, including the job token.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Agent steps commonly read content their author did not write: flow inputs from webhooks, email and other triggers, web search results, and the output of tools and MCP servers, which all enter the conversation like any other tool result. Nothing marks that content as untrusted or restricts what the agent may do after reading it, and no tool call needs approval. A successful prompt injection can therefore use any attached tool to send data out and take irreversible actions with the runner's authority.

C6 Memory, context & configuration integrity

Minimal 0.35 / 1.00

Agent steps keep conversation memory by default, summarising older turns when the context fills up and storing the whole conversation, including tool results, in the database for the next run of the same conversation. Memory is kept per workspace, flow and conversation id, and the model cannot choose which memory it reads. Nothing validates what is written, so injected text that lands in a tool result or summary is replayed in every later turn of that conversation. There are no workspace instruction files or repository configs that an agent loads.

C7 Third-party extensions

Minimal 0.25 / 1.00

Agents can use third-party code in two ways: Hub scripts, which are pinned to a version id and fetched from Windmill's hub, and remote MCP servers, whose tool lists are fetched fresh on every run with no pinning or change detection. Nothing third-party is enabled by default; the author adds each one. MCP servers run remotely and only receive their own token and the call arguments, but Hub scripts run on the worker like any other job, with the runner's job token and the same weak default isolation.

C8 Secrets & sensitive-data protection

Minimal 0.28 / 1.00

Secret variables are encrypted at rest with a per-workspace key, and any secret a job fetches, along with its own token, is masked in that job's logs. Jobs start from a cleared environment. Author-pinned credentials are kept out of the tool schema the model sees. Redaction covers job logs only: job results, stored conversation memory and messages sent to the model provider are not redacted, instance-level settings are stored unencrypted under the default backend, and every tool job receives a long-lived token with the runner's full permissions. Usage statistics code is closed source in this repository.

C9 Audit & traceability

Moderate 0.65 / 1.00

Each tool call is recorded as it happens in the flow's status, and Windmill tools run as separate jobs whose arguments, results, logs, timestamps, creator and run-as identity are stored with parent and root job ids, so a run can be traced across nested agents. MCP calls are recorded with their arguments, and the agent's final result carries the full message history. The audit log that would record configuration changes and secret reads is an Enterprise feature and does nothing in this code. Records live in the instance database, which tool jobs can reach through the API with the runner's permissions.

C10 Limits & kill switch

Moderate 0.50 / 1.00

An agent step stops after 10 model turns by default, an author can raise that to at most 1,000, and every model request and tool job is bounded by the job timeout, which defaults to seven days on a self-hosted install. Cancelling a run stops the loop, gives in-flight tools 30 seconds and then aborts them and cancels any tool jobs still queued or running. There is no token or cost budget, and a nested agent tool starts its own turn budget.