C1 Identity & least privilege
Minimal 0.00 / 1.00
The harness has no identity or authorization layer of its own. The default model client uses the AWS default credential chain, and the shell tool runs commands on the host with the full environment of the process, so anything the operator's credentials, AWS profile, SSH agent or tokens can do, a hijacked agent can do too. Authorization lives only in the system prompt.
C2 Approval gates
Minimal 0.45 / 1.00
By default every tool call, including shell and file writes, runs without asking: the approval gate is off unless the developer passes interventions. When enabled, the 'ask' preset shows the exact tool name and JSON arguments for each call and applies to all tools, including MCP tools, sub-agent delegates and the calls made from programmatic_tool_caller. Nothing provides undo or checkpoints by default, so a wrongly allowed shell command can be irreversible.
C3 Tool & action scoping
Minimal 0.13 / 1.00
The default tool set is the broadest possible: a raw shell, unrestricted file write and edit anywhere on the host, a web fetcher and a sub-agent. The file tools only require an absolute path without '..' segments, and web_fetch only checks the URL scheme and characters, with no block on internal or cloud-metadata addresses. The developer can trim the tool list, but nothing is narrowed by default.
C4 Code-execution isolation
Minimal 0.35 / 1.00
The shell tool, file tools and web_fetch run directly on the host through the SDK's NotASandboxLocalEnvironment, whose own docstring says it provides no isolation, and child processes inherit the full environment. A Docker sandbox exists as an option, but it only runs 'docker exec' into a container the developer already started, with no hardening applied by the framework. The programmatic_tool_caller runs model code in the Monty interpreter, which has no host access, but it can call the unisolated shell.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
The default agent reads web pages, files and MCP tool results, holds the operator's credentials, and can run shell commands and fetch arbitrary URLs, all in one session with no approval. Nothing tracks whether untrusted content has been read, and the project's AGENTS.md is injected inside a system-reminder block, giving repository text elevated standing. If injected content hijacks the agent, it can exfiltrate secrets and take irreversible actions unattended.
C6 Memory, context & configuration integrity
Minimal 0.13 / 1.00
Long-term memory is on by default: a small model distils 'facts' from the whole conversation, including tool results, into markdown files in ./.agent/memory under the working directory, and they are injected into context on every turn. Because the folder lives in the workspace, a repository can ship pre-written memory, and it also auto-loads ./.agent/skills and the project AGENTS.md. None of these files can change tools or approval settings, but nothing validates or reviews what gets saved or loaded.
C7 Third-party extensions
Minimal 0.07 / 1.00
MCP servers are only connected when the developer passes them, but their packages are not pinned or hash-checked, and skills found in ./.agent/skills of the working directory are loaded automatically. Skill scripts run through the same unisolated host shell, with the full environment, when the model follows a skill's instructions. MCP stdio servers get only the environment the config names, but run as the same user on the host.
C8 Secrets & sensitive-data protection
Minimal 0.25 / 1.00
Credentials come from the environment and the AWS default chain, and the shell passes the whole environment to every command, so the model can read keys with a single 'env' call. Tracing is off unless OTEL_TRACES_EXPORTER is set, but when it is on, prompts and tool data are exported unredacted unless redaction is opted into. Session transcripts are stored as plaintext under ./.agent/sessions, with no secret scanning.
C9 Audit & traceability
Minimal 0.35 / 1.00
By default the harness saves a snapshot of the conversation, including tool calls and results, after each message under ./.agent/sessions in the working directory. The snapshot is overwritten each time and the context manager summarizes old turns, so it is not a complete trail. Calls made from inside programmatic_tool_caller are not recorded, and sub-agents run without a session. OpenTelemetry spans for tool calls and delegation exist but are only exported when configured.
C10 Limits & kill switch
Minimal 0.40 / 1.00
The SDK can cap turns and token spend per invocation, but the harness never sets those limits, so a default run is unbounded. Some bounds are on by default: sub-agent delegation stops at depth 2, web_fetch times out at 30 seconds, programmatic_tool_caller has VM limits and a 15-minute wall clock, and shell commands default to 120 seconds, though the model chooses that timeout per call. Cancellation is cooperative and does not stop a shell command that is already running.