BoundBench

Pydantic AI

Type-safe Python agent framework

github.com/pydantic/pydantic-ai · 2026-10-03 · 6695132

Defense-in-depth score

3.5 / 10

Minimal

Pydantic AI ships a minimal default (no tools) and some well-engineered primitives, such as a per-call human-approval gate, an SSRF-hardened web fetch tool and a 50-request run cap. But approval, isolation and telemetry are all opt-in. Registered tools run unattended in your process with all of its credentials, so if the agent reads attacker-controlled content it can leak data and act with no human involved. To change that, you have to mark consequential tools requires_approval and run code in a sandbox yourself.

Key gaps (2)

  1. A hijacked agent can exfiltrate data and run any registered state-changing tool unattended; tool approval is off by default and no egress control is tied to untrusted input. C5 · Untrusted input blast radius
  2. The core library's only local execution backend runs commands as host subprocesses that, per its own docs, isolate nothing; the user's whole home directory and network are reachable. C4 · Code-execution isolation

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

Pydantic AI has no identity or authorization layer of its own: registered tools run inside the developer's Python process with whatever credentials that process holds, and any per-user authorization is left to developer code. It does narrow ambient authority in two places that ship on by default: the core local workspace gives subprocesses only PATH, HOME and locale variables, and the UI adapters drop client-submitted s3:// or gs:// file URLs and uploaded-file references that the model provider would otherwise fetch with the server's IAM role. In-process function tools still hold the full process authority, so a hijacked agent can use whatever the application holds.

C2 Approval gates

Minimal 0.30 / 1.00

Pydantic AI has a well-built human-approval primitive: a tool marked requires_approval, or any toolset wrapped with approval_required(), pauses the run and hands the exact tool call (name, arguments, id) back to the application, which answers approve, approve-with-edited-arguments (re-validated), or deny. The gate is enforced on both execution paths in code. But it is off by default: every tool runs unattended unless the developer flags it, and provider-executed native tools such as code execution or web fetch never pass through it. There is no undo, checkpoint, or dry-run primitive.

C3 Tool & action scoping

Moderate 0.50 / 1.00

Every function tool's arguments are validated against a Pydantic schema generated from its type hints, and developers can add an args_validator, but that is type checking rather than an allowlist. The one bundled network tool, web_fetch, is genuinely hardened: it blocks private and cloud-metadata addresses, pins the resolved IP, and re-validates every redirect. MCP tool arguments are only checked to be a JSON object before being forwarded. The agent starts with no tools at all, so the default posture is minimal, but anything registered runs with its full reach.

C4 Code-execution isolation

Minimal 0.47 / 1.00

The core library runs nothing model-generated by default, but its only local execution backend, LocalWorkspace, runs commands as plain host subprocesses and is documented by the maintainers as isolating nothing; function tools themselves run in the application's process. When code execution is enabled this way, an injected command reaches the user's whole home directory and network. The separate pydantic-ai-harness package in the same repo offers real sandboxes (E2B remote VMs, bubblewrap); E2B is the strongest, but it is opt-in, has internet access on by default, and only covers workspace commands, not function tools or MCP servers.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Nothing in the framework limits what a hijacked agent can do once it reads attacker-controlled content: tool results, MCP results and fetched pages enter the conversation as ordinary tool returns, and no step disables egress or forces approval after untrusted content arrives. The only measures are labelling ones: tool-produced files are wrapped in provenance tags (which the code itself says 'identify rather than prove'), and the UI adapters strip client-injected system prompts. Because its docs routinely combine web fetching, private application data and outbound tools, a prompt injection can leak data and trigger any registered tool with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.30 / 1.00

The core library has no memory store and auto-loads no workspace instruction or config files; persistence happens only when the developer saves message history and passes it back, or implements the tool behind Anthropic's native memory tool. History passed back is replayed as trusted context, so injected text saved in a prior run can steer tool calls in later ones. Client-submitted history through the UI adapters is sanitized by default (system prompts, workspace references, unresolved tool calls and provider-IAM file references are stripped).

C7 Third-party extensions

Minimal 0.23 / 1.00

Third-party code comes in mainly through MCP servers, which the developer names explicitly (a URL, a command, or a JSON config passed by path); nothing is enabled by default and no workspace file adds servers. There is no version pinning, hash, or signature check, and no detection if a server's tools change between sessions. Stdio servers run as separate processes under the same user; their environment handling is delegated to the fastmcp and MCP SDK libraries.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

Model-provider keys are read from environment variables and provider reprs omit them, and workspace subprocesses receive only PATH, HOME and locale variables. There is no type-level secret masking, no redaction of secrets in tool results sent to the model, and no secret scanning. Telemetry is off by default, but when OpenTelemetry/Logfire instrumentation is enabled it records prompts, tool arguments and tool results by default. MCP JSON configs can expand any environment variable into server arguments or headers.

C9 Audit & traceability

Minimal 0.42 / 1.00

By default, the only record is the in-memory message history the run returns, which captures every tool call with arguments, IDs and timestamps but is lost unless the developer saves it. Opt-in OpenTelemetry/Logfire instrumentation emits standard gen_ai spans for each run and tool call (including MCP tools and deferrals) and can ship them off-host, but it does not record who requested or approved an action.

C10 Limits & kill switch

Minimal 0.45 / 1.00

Each run is capped at 50 model requests by default, and developers can add token, cost and tool-call limits plus per-tool timeouts, all enforced in code. There is no default wall-clock, token or cost limit, and tools have no timeout unless set. Cancelling a run cancels the driving task and in-flight async tool tasks, but synchronous tools run in threads that are not abandoned, and LocalWorkspace background jobs outlive the run. A sub-agent shares the parent's budget only if the developer passes usage through.