C1 Identity & least privilege
Minimal 0.33 / 1.00
Built-in tools act with whatever long-lived credentials the operator puts in the environment: the GitHub tools, for example, read a personal token from the environment and use it for every call, and there is no authorization layer between the agent and its tools. The toolkit does ship per-user OAuth 2.0 authentication providers that MCP and A2A clients can use to act on behalf of the signed-in user, but they have to be configured and only cover the clients that reference them. As shipped, a hijacked agent holds the operator's full authority for every configured service.
C2 Approval gates
Minimal 0.15 / 1.00
Nothing in the default workflow asks a human before a tool runs: built-in agents call GitHub write operations, code execution, search and MCP tools directly. The toolkit ships a human-in-the-loop middleware base class that can pause before or after a function, but it is abstract (a developer must subclass and register it), it is attached only to the functions it is configured for, and the prompt it shows is fixed text rather than the actual call and arguments. A workflow as shipped performs every consequential action unattended.
C3 Tool & action scoping
Moderate 0.50 / 1.00
Every toolkit function validates its input against a typed schema before it runs, and a few built-in tools go further: the GitHub commit tool resolves paths and keeps them inside the configured repository directory, and repository and branch names are checked before they are used in URLs. Other tools take free-form input, such as arbitrary Python for the code-execution tool or arbitrary queries for search tools, and there is no shared allowlist or policy layer. No tools are enabled by default; each agent receives only the tools named in its configuration.
C4 Code-execution isolation
Moderate 0.63 / 1.00
The only built-in way for the model to run code is the code-execution tool, which never runs code on the host: it sends the code to a separately deployed Piston sandbox server and returns an error if that server is unreachable or times out. The toolkit does not ship or harden the sandbox itself, so how strong the isolation is depends on how the operator deploys Piston, and the client asks for no memory limit. No shell tool or local Python interpreter is exposed to the model.
C5 Untrusted input blast radius
Minimal 0.15 / 1.00
Tool results, web search results, GitHub issue text and MCP tool output all go back to the model as ordinary tool messages, with nothing marking them as untrusted and nothing restricting what the agent does after reading them. The optional security package adds LLM-based checks (a pre-tool verifier, content-safety guard, output verifier and PII filter) and a NeMo Guardrails middleware, but these are detection layers, must be configured, and by default only log a violation rather than block it. In a workflow that combines web or issue content with GitHub write tools or search queries, a successful injection can both send data out and make changes without a human.
C6 Memory, context & configuration integrity
Minimal 0.10 / 1.00
Memory is optional, but when it is used the model can store anything through the add-memory tool, and the auto-memory wrapper saves every user and agent message by default and feeds retrieved memories back as a system message, the highest-priority context. Memory is kept per user: the user identity comes from trusted configuration or the runtime session, never from the model, and operations fail if no identity is available. There is no validation, expiry or review of stored memories, so injected text can persist across a user's sessions and steer later tool use. Workflow configuration comes only from the file the operator names; nothing in a working directory is loaded as instructions.
C7 Third-party extensions
Minimal 0.25 / 1.00
Extensions come in two forms: Python plugin packages, which the toolkit discovers through package entry points and imports into its own process at startup, and MCP servers, which the operator lists in the workflow file with the exact command or URL. Nothing is pinned, hashed or re-approved when a server's tools or a package's code change, and every installed plugin package is loaded automatically once present. Hugging Face models do default to not running remote code. A malicious plugin runs with everything the agent holds; stdio MCP servers run as separate processes on the host.
C8 Secrets & sensitive-data protection
Minimal 0.25 / 1.00
API keys are read from environment variables or the workflow file into secret-typed fields that hide their value when printed, but no redaction applies to logs by default and trace redaction is an opt-in setting. Anonymous CLI usage telemetry is content-free and asked for at first run, but the prompt defaults to yes. The README's hello-world sets verbose agent logging, which writes tool inputs and outputs to the console unredacted. Keys are long-lived and as broad as whatever the operator exported.
C9 Audit & traceability
Moderate 0.50 / 1.00
Every toolkit function call, including MCP tools and agents used as tools, emits structured start and end events with inputs, outputs, timing and parent links, and the toolkit can export them as OpenTelemetry spans (with user and session attribution) to a file or to many tracing backends. None of that is recorded by default: no exporter is configured unless the operator adds one, and the default console logging does not record tool calls. When enabled, exports are batched and best-effort.
C10 Limits & kill switch
Minimal 0.42 / 1.00
The built-in ReAct and tool-calling agents stop after 15 tool calls by default, enforced through the LangGraph recursion limit, and some tools carry their own timeouts (10 seconds for code execution, 60 seconds for MCP calls). There is no wall-clock or cost limit for a run, model requests have no timeout by default, and the timeout and circuit-breaker middleware are optional. An agent used as a tool by another agent gets its own fresh budget, so nesting multiplies the limit.