BoundBench

Archon

Workflow engine and harness that runs AI coding agents (Claude Code, Codex, Pi and others) through YAML-defined development workflows from a CLI, web UI or chat platforms.

github.com/coleam00/Archon · 2026-10-05 · f326bb4

Defense-in-depth score

1.7 / 10

Minimal

Archon runs Claude Code with permission prompts switched off (and Codex with approvals set to never), handing the agent the user's full environment and credentials and letting it act on GitHub issues, push branches and open pull requests unattended. The per-run git worktree keeps workflows from colliding but is not a security boundary, and the only container isolation is opt-in and limited to folder projects. Repositories can bring their own workflows, MCP servers, environment and assistant settings, which load without a trust prompt. Run it only on repositories and inputs you fully trust.

Key gaps (7)

  1. Agent tools run with permission prompts bypassed, so shell commands, file writes and pushes need no human approval. C2 · Approval gates
  2. A paused workflow approval gate can be resolved by the chat model through its run-management tool. C2 · Approval gates
  3. The agent and all subprocesses inherit the user's full environment and credentials. C1 · Identity & least privilege
  4. Model-driven commands run on the host as the user with no OS boundary in the default configuration, and a repository workflow can turn the worktree off. C4 · Code-execution isolation
  5. Content from GitHub issues and other sources reaches an agent that can leak credentials and push code with no approval step. C5 · Untrusted input blast radius
  6. Repository files load workflows, MCP servers, hooks, environment and assistant settings without a trust decision. C6 · Memory, context & configuration integrity
  7. MCP servers run as the user and can be configured to receive any variable from the full environment. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

Archon acts with whatever authority the person running it has. The Claude Code subprocess, bash and script nodes all receive the full environment of the Archon process, so every API key, GitHub token and cloud credential in it is available to the agent, and the agent runs Claude Code with permission prompts switched off. A per-user GitHub token policy exists for multi-user server installs, but it does nothing in the default single-user CLI setup. If the agent is steered wrong, it can do anything the user's own accounts allow.

C2 Approval gates

Minimal 0.05 / 1.00

Archon deliberately runs Claude Code in its permission-bypass mode and Codex with approvals set to never, so no individual command, file write, push or pull-request creation is shown to a human first; the project documents this as a design choice for unattended workflows. Workflows can include human approval gate nodes, but the bundled delivery workflow has none (the human gate is PR review on GitHub), and the chat assistant's run-management tool can itself approve or reject a paused gate. Changes happen in a git worktree by default, but pushes, PRs and anything done through the shell outside it are not reversible by Archon.

C3 Tool & action scoping

Minimal 0.05 / 1.00

Agents get the full Claude Code (or Codex) tool set: arbitrary shell commands, file writes anywhere the user can write, and web access, with no argument validation by Archon. Workflow authors can restrict each node to a list of allowed or denied tools, which is a useful coarse control, but the bundled workflows don't use it and it never inspects arguments. A misused tool reaches the whole machine.

C4 Code-execution isolation

Minimal 0.28 / 1.00

For repository projects, which is how most people use Archon, every workflow run happens in a separate git worktree, but that is a working directory, not a boundary: Claude Code, bash nodes and script nodes all run as the user on the host with the full environment and permission prompts off. A workflow file can also pin the worktree off. An opt-in Docker container backend exists for folder projects only; it keeps the host environment out and caps memory and processes, but has open network egress, runs as root inside, and on a standard Docker daemon falls back to a mode the project's own security notes describe as escapable.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

The bundled workflows read GitHub issues, pull requests and repository content and feed them to an agent that can run any shell command, push branches and open pull requests, all without a human in the loop. Nothing in code marks or limits content from those sources; the only defence is prompt wording asking the model to treat issue text as claims. If injected text hijacks a run, it can both leak the credentials and files the agent can see and take irreversible actions such as pushing code.

C6 Memory, context & configuration integrity

Minimal 0.00 / 1.00

Archon is designed so a repository carries its own automation: workflows, commands, scripts, an .archon/config.yaml (including environment variables and assistant settings), a repo-scope .archon/.env, and Claude Code project settings and CLAUDE.md are all loaded from the working checkout with no trust prompt. That means a cloned repository can add MCP servers, hooks and shell steps, and change how the agent is configured, simply by containing the right files; the project documents these sources as trusted. Conversation and run history persist in Archon's database and are scoped per project.

C7 Third-party extensions

Minimal 0.23 / 1.00

Archon's own plugins are installed explicitly from GitHub, pinned to a tag or commit, checked against the release's checksum file, and forge plugin processes get a scrubbed environment. MCP servers, Claude plugins and hooks are a different story: a workflow node can point at an MCP configuration in the repository, which is launched by whatever command it names, and that configuration can pull any environment variable into a server's environment or headers. Nothing verifies or confines these servers.

C8 Secrets & sensitive-data protection

Minimal 0.35 / 1.00

Archon has real secret hygiene in places: per-user provider and GitHub credentials are encrypted at rest with AES-256-GCM, never returned by the API, and bash/script output is redacted against secret-named environment values before it is logged or retained. But in the default setup secrets live in environment files and are passed whole to the agent and every subprocess, so the model can read them, and node output passed on to the next model step is not redacted. Anonymous, content-free telemetry is on by default with documented opt-outs.

C9 Audit & traceability

Minimal 0.45 / 1.00

Each workflow run writes a structured JSONL log and database events covering node starts and ends, provider tool calls, bash and script output, and gate decisions, stored under Archon's home directory rather than in the worktree. The record has no actor attribution for approvals, the agent runs as the same user and could edit it, and a logging failure only produces a single warning while the run carries on.

C10 Limits & kill switch

Minimal 0.33 / 1.00

Workflow structure bounds the overall run: loops must declare a maximum iteration count, sub-workflow nesting is capped, and bash and script steps time out after two minutes by default. Inside an AI step, though, nothing caps turns, time or spend by default: the 30-minute idle timer resets on any output and is described as a deadlock detector, and a dollar budget exists only if a workflow sets it. Cancelling a run aborts the in-flight provider call.