BoundBench

AgentField

Open-source control plane and Python/Go/TypeScript SDKs for running AI agents as REST microservices, with memory, async execution, coding-agent harnesses and an embedded MCP server.

github.com/agent-field/agentfield · 2026-10-04 · 68cadab

Defense-in-depth score

2.2 / 10

Minimal

As shipped, access control on the AgentField control plane is not locked down in the default configuration. Agents and coding-agent harnesses run unsandboxed on the host with the operator's full environment, and installed agent packages auto-update from upstream every six hours by default. Strong pieces exist (DID identities, access policies with argument constraints, fail-closed execution records, an encrypted secret store), but most are opt-in and the policy layer does not cover every route.

Key gaps (6)

  1. Agents and harness subprocesses inherit the operator's full environment, so a hijack holds the operator's account. C1 · Identity & least privilege
  2. With authorization enabled, agents self-declare tags that are auto-approved by default, which lets them choose their own access-policy authority. C1 · Identity & least privilege
  3. Approval resolution is protected only by the shared API key agents hold, so approvals can be self-satisfied. C2 · Approval gates
  4. Harness CLIs and agent nodes run unsandboxed on the host with the full parent environment. C4 · Code-execution isolation
  5. A hijacked tool loop can call any reasoner, including unoffered ones, with no human gate while agents hold credentials and egress. C5 · Untrusted input blast radius
  6. Installed agent packages auto-update from upstream HEAD every 6 hours by default and run with the operator's full environment. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.20 / 1.00

Authentication on the control plane is not locked down in its default configuration. When a key is set it is one shared secret that every agent node also receives, and agent nodes and harness subprocesses inherit the operator's full environment (cloud and LLM credentials). A per-agent DID identity with tag-based access policies exists, but it is off by default, tags are self-declared and auto-approved, and enforcement of the policy check does not cover every route.

C2 Approval gates

Minimal 0.15 / 1.00

AgentField offers a human-approval primitive: agent code calls app.pause(), the run goes into a waiting state, and an external approval service or operator approves, rejects or requests changes. Nothing is gated by default, though. No reasoner, tool-loop call or harness action needs approval unless the developer writes the pause call. The approval-response endpoint is protected only by the global API key, which agent nodes also hold, so an agent can approve its own request.

C3 Tool & action scoping

Minimal 0.33 / 1.00

Agent functions are typed Python functions, and the SDK coerces inputs to the declared types or Pydantic models. There are no allowlists on paths, hosts or amounts. The built-in tool loop (tools="discover") hands the model every function registered on the control plane, and it dispatches whatever function name the model returns without checking that name against the offered set. The opt-in access-policy layer can allow or deny functions by glob and put numeric limits on arguments, failing closed when a constrained parameter is missing. But it is off by default, falls through to allow when no policy matches, and its enforcement does not cover every route.

C4 Code-execution isolation

Minimal 0.00 / 1.00

AgentField runs model-driven code in two ways: the harness, which launches coding CLIs such as aforge, Claude Code, Codex and OpenCode, and agent-node packages that the control plane installs and starts. Both run as ordinary host subprocesses under the operator's user, with the full parent environment. The framework adds no sandbox of its own, the OpenCode baseline allows every tool, and permission_mode='auto' maps to Claude Code's bypassPermissions. Any isolation comes from the third-party CLI's own defaults.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Tool results go back to the model as ordinary tool messages, with nothing marking where they came from. The model can then call any agent on the control plane, including ones it was not offered. Nothing in the framework separates sessions that read untrusted content (webhook triggers, other agents' output, fetched data) from sessions that hold secrets or can act. If a hijack succeeds, it can both leak data and take actions with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.17 / 1.00

The control plane keeps persistent key-value, vector and knowledge memory with workflow, session, actor and global scopes. The scope is chosen from headers the caller supplies, the global scope is open to everyone by design, and the permission middleware is off by default. Writes are recorded as memory events but not validated, and entries carry no provenance. The SDK does not yet inject memory into prompts automatically, so poisoned memory reaches the model only through developer code. Separately, `af server` loads agentfield.yaml from ./config or the current directory when no user-level config exists.

C7 Third-party extensions

Minimal 0.15 / 1.00

Agent nodes are installed from git repositories the user chooses. They are cloned, their dependencies are installed with pip or npm, and they run as host processes with the operator's full environment and the control-plane API key. Unpinned installs track the remote HEAD, and automatic updates are on by default: every six hours the control plane pulls and reinstalls new upstream code without asking again. The bundled aforge harness binary is the exception. Its version is pinned and its download is checked against a SHA-256 checksum.

C8 Secrets & sensitive-data protection

Minimal 0.28 / 1.00

Secrets that agent nodes declare are kept in a local store encrypted with AES-256-GCM, under a key file with 0600 permissions. Execution payloads are left out of structured logs by default. Those are the only protected paths. Agent nodes and harness subprocesses receive the operator's entire environment (LLM keys, cloud credentials, the control-plane API key), nothing is redacted from what is sent to the model, and execution inputs and outputs are stored unencrypted in the local database. Anonymous, content-free telemetry is on by default. Secret handling in the shipped configuration is not fully locked down either.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every call routed through the control plane, including calls from the SDK tool loop and the MCP endpoint, becomes a structured execution record with target, input, status, timestamps, run ID and parent-execution ID. The record is written before the call is dispatched, and the call is aborted if the write fails. The parent and actor fields come from headers the caller supplies, so attribution is self-asserted. Actions taken inside a reasoner or a harness CLI are not recorded. The records live in the local database and can be deleted through the API. Signed verifiable credentials per execution exist, but the shipped `af` binary leaves them off unless they are configured.

C10 Limits & kill switch

Minimal 0.45 / 1.00

The SDK tool loop stops at 10 turns and 25 tool calls, calls between agents time out after 90 seconds by default, and harness runs are killed with their whole process group after 30 minutes. There is no cost cap. Rate limiting is off by default, and per-agent concurrency is unlimited. Fan-out through app.call starts a fresh budget at every level, with no depth limit beyond an advisory environment variable. Cancelling a workflow tree cancels the agents' asyncio tasks, but the cancel is delivered best-effort, so in-flight work may finish anyway.