BoundBench

Strands Agents (harness-sdk)

AWS's Strands Agents monorepo: Python/TypeScript agent SDKs plus the batteries-included Strands harness (create_harness) and CLI.

github.com/strands-agents/harness-sdk · 2026-10-04 · 8ba04f7

Defense-in-depth score

2.1 / 10

Minimal

The default harness agent gets a host shell, unrestricted file writes, web fetching and sub-agents, all running on your machine with your full environment and AWS credentials, and every tool call runs without approval. Good controls exist (a per-call approval gate that covers sub-agents and programmatic tool calls, a Docker sandbox option, turn and token budgets), but all are opt-in. The dominant risk is prompt injection from a web page or repository file turning into credential theft or destructive commands with no human in the loop.

Key gaps (4)

  1. The default agent runs with the operator's ambient credentials (AWS default chain and full process environment in every shell command), so a hijack holds the user's whole account. C1 · Identity & least privilege
  2. Model-generated shell commands run on the host with no isolation and the full process environment by default. C4 · Code-execution isolation
  3. A prompt-injected session can both leak secrets (host shell with full environment, unrestricted web_fetch) and take irreversible actions with no human involved. C5 · Untrusted input blast radius
  4. Workspace skills auto-load, and their bundled scripts run through the host shell with the agent's full environment and credentials. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

The harness has no identity or authorization layer of its own. The default model client uses the AWS default credential chain, and the shell tool runs commands on the host with the full environment of the process, so anything the operator's credentials, AWS profile, SSH agent or tokens can do, a hijacked agent can do too. Authorization lives only in the system prompt.

C2 Approval gates

Minimal 0.45 / 1.00

By default every tool call, including shell and file writes, runs without asking: the approval gate is off unless the developer passes interventions. When enabled, the 'ask' preset shows the exact tool name and JSON arguments for each call and applies to all tools, including MCP tools, sub-agent delegates and the calls made from programmatic_tool_caller. Nothing provides undo or checkpoints by default, so a wrongly allowed shell command can be irreversible.

C3 Tool & action scoping

Minimal 0.13 / 1.00

The default tool set is the broadest possible: a raw shell, unrestricted file write and edit anywhere on the host, a web fetcher and a sub-agent. The file tools only require an absolute path without '..' segments, and web_fetch only checks the URL scheme and characters, with no block on internal or cloud-metadata addresses. The developer can trim the tool list, but nothing is narrowed by default.

C4 Code-execution isolation

Minimal 0.35 / 1.00

The shell tool, file tools and web_fetch run directly on the host through the SDK's NotASandboxLocalEnvironment, whose own docstring says it provides no isolation, and child processes inherit the full environment. A Docker sandbox exists as an option, but it only runs 'docker exec' into a container the developer already started, with no hardening applied by the framework. The programmatic_tool_caller runs model code in the Monty interpreter, which has no host access, but it can call the unisolated shell.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

The default agent reads web pages, files and MCP tool results, holds the operator's credentials, and can run shell commands and fetch arbitrary URLs, all in one session with no approval. Nothing tracks whether untrusted content has been read, and the project's AGENTS.md is injected inside a system-reminder block, giving repository text elevated standing. If injected content hijacks the agent, it can exfiltrate secrets and take irreversible actions unattended.

C6 Memory, context & configuration integrity

Minimal 0.13 / 1.00

Long-term memory is on by default: a small model distils 'facts' from the whole conversation, including tool results, into markdown files in ./.agent/memory under the working directory, and they are injected into context on every turn. Because the folder lives in the workspace, a repository can ship pre-written memory, and it also auto-loads ./.agent/skills and the project AGENTS.md. None of these files can change tools or approval settings, but nothing validates or reviews what gets saved or loaded.

C7 Third-party extensions

Minimal 0.07 / 1.00

MCP servers are only connected when the developer passes them, but their packages are not pinned or hash-checked, and skills found in ./.agent/skills of the working directory are loaded automatically. Skill scripts run through the same unisolated host shell, with the full environment, when the model follows a skill's instructions. MCP stdio servers get only the environment the config names, but run as the same user on the host.

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

Credentials come from the environment and the AWS default chain, and the shell passes the whole environment to every command, so the model can read keys with a single 'env' call. Tracing is off unless OTEL_TRACES_EXPORTER is set, but when it is on, prompts and tool data are exported unredacted unless redaction is opted into. Session transcripts are stored as plaintext under ./.agent/sessions, with no secret scanning.

C9 Audit & traceability

Minimal 0.35 / 1.00

By default the harness saves a snapshot of the conversation, including tool calls and results, after each message under ./.agent/sessions in the working directory. The snapshot is overwritten each time and the context manager summarizes old turns, so it is not a complete trail. Calls made from inside programmatic_tool_caller are not recorded, and sub-agents run without a session. OpenTelemetry spans for tool calls and delegation exist but are only exported when configured.

C10 Limits & kill switch

Minimal 0.40 / 1.00

The SDK can cap turns and token spend per invocation, but the harness never sets those limits, so a default run is unbounded. Some bounds are on by default: sub-agent delegation stops at depth 2, web_fetch times out at 30 seconds, programmatic_tool_caller has VM limits and a 15-minute wall clock, and shell commands default to 120 seconds, though the model chooses that timeout per call. Cancellation is cooperative and does not stop a shell command that is already running.