# Defense-in-depth score: NVIDIA NeMo Agent Toolkit

**Repo:** https://github.com/NVIDIA/NeMo-Agent-Toolkit · **Commit:** `c7e1162a1c7ff18bbd797e090a56cad97c281c92` · **Reviewed:** 2026-10-05
**What it is:** NVIDIA's Python toolkit for building, serving, observing, evaluating and optimizing agent workflows defined in YAML, across LangChain, LlamaIndex, CrewAI and other frameworks.
**Category:** Agent Frameworks
**Scored configuration:** Framework defaults: a YAML workflow run with `nat run` (or `nat serve` on localhost) using a built-in agent such as react_agent with the default arguments of each built-in tool, middleware and memory component, and no optional middleware, exporters or auth providers configured.
**Agent surface (default):** code execution opt-in · filesystem write no · network egress yes · external credentials yes · persistent memory opt-in · untrusted input yes · third party extensions opt-in · sub agents opt-in · external communication opt-in

## Score: 3.3 / 10.0 (Minimal)

| # | Criterion | S | C | D | B | Raw | Cap | Score | Confidence |
|---|---|---|---|---|---|---|---|---|---|
| C1 | Identity & least privilege | L2 | L1 | L0 | L2 | 0.33 | G1 | **0.33** (alt) | Medium |
| C2 | Approval gates | L1 | L1 | L0 | L0 | 0.15 | G1 | **0.15** (alt) | High |
| C3 | Tool & action scoping | L2 | L2 | L3 | L1 | 0.50 | none | **0.50** | High |
| C4 | Code-execution isolation | L2 | L3 | L3 | L2 | 0.62 | none | **0.62** | Medium |
| C5 | Untrusted input blast radius | L1 | L1 | L0 | L0 | 0.15 | G1 | **0.15** (alt) | High |
| C6 | Memory, context & configuration integrity | L0 | L0 | L1 | L1 | 0.10 | none | **0.10** | High |
| C7 | Third-party extensions | L1 | L1 | L2 | L0 | 0.25 | none | **0.25** | High |
| C8 | Secrets & sensitive-data protection | L1 | L1 | L1 | L1 | 0.25 | none | **0.25** | High |
| C9 | Audit & traceability | L3 | L3 | L0 | L1 | 0.50 | G1 | **0.50** | High |
| C10 | Limits & kill switch | L2 | L1 | L2 | L2 | 0.42 | none | **0.42** | High |


NeMo Agent Toolkit gives developers solid building blocks (typed tool inputs, explicit per-agent tool lists, a remote-only code sandbox, per-user memory namespaces and rich tracing), but almost every safeguard that limits a misbehaving agent is optional and off by default. As shipped, agents call write-capable tools such as GitHub issue, pull request and commit operations with no human approval, using whatever long-lived tokens are in the environment. The dominant risk is prompt injection through web search, GitHub issues or MCP results steering an agent into leaking data or making changes unattended; the human-in-the-loop, guardrail, timeout and trace-export features have to be added and configured by the developer.

## Critical gaps
- Untrusted content read by a default agent can drive both data egress and irreversible GitHub writes with no human approval. (ASI01, LLM01; C5). Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py:255](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py#L255); [packages/nvidia_nat_core/src/nat/tool/github_tools.py:244](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L244)
- Installed plugin packages are discovered and imported into the agent's own process automatically, with all of its credentials. (ASI04, LLM03; C7). Evidence: [packages/nvidia_nat_core/src/nat/runtime/loader.py:195](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/runtime/loader.py#L195)

## Criterion details

### C1 Identity & least privilege: 0.33 (medium confidence)

Built-in tools act with whatever long-lived credentials the operator puts in the environment: the GitHub tools, for example, read a personal token from the environment and use it for every call, and there is no authorization layer between the agent and its tools. The toolkit does ship per-user OAuth 2.0 authentication providers that MCP and A2A clients can use to act on behalf of the signed-in user, but they have to be configured and only cover the clients that reference them. As shipped, a hijacked agent holds the operator's full authority for every configured service.

- **default configuration** (default; raw 0.05 → 0.05)
  - **S L0:** The GitHub function group takes the first token found in GITHUB_TOKEN, GITHUB_PAT or GH_TOKEN and attaches it as a bearer header to one client used for reads and writes alike. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:226-236](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L226-L236) (verified)
    - *To reach the next level:* No dedicated or scoped identity: built-in tools use whatever ambient token the operator exported, with read and write sharing one credential.
  - **C L0:** Tools construct their own privileged clients from environment credentials; the generic function invocation path applies no authorization check before a tool runs. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:236](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L236); [packages/nvidia_nat_core/src/nat/builder/function.py:198-202](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/builder/function.py#L198-L202) (verified)
    - *To reach the next level:* No shared authorization layer that every tool call passes through.
  - **D L0:** Nothing narrows the default: the authority a workflow holds is exactly what the operator exported, and least privilege is left to manual token creation. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:226-227](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L226-L227) (verified)
    - *To reach the next level:* No narrower default identity or role; least privilege depends entirely on the operator minting restricted tokens.
  - **B L1:** Documented workflows combine an LLM API key with write-capable tools such as GitHub issue, pull request and commit operations, so a hijacked identity can write to several systems. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:323](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L323); [packages/nvidia_nat_core/src/nat/llm/nim_llm.py:39](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/llm/nim_llm.py#L39) (verified)
    - *To reach the next level:* Credentials are long-lived and not limited to one project or to read-only operations.
- **opt-in per-user OAuth 2.0 auth providers for MCP and A2A clients** (alt; raw 0.33, cap G1 → 0.33) ← counted
  - **S L2:** The MCP auth provider obtains OAuth 2.0 tokens per user_id through the authorization-code flow, so remote tool calls run on behalf of the requesting user, with scopes chosen by the operator (empty list by default). Evidence: [packages/nvidia_nat_mcp/src/nat/plugins/mcp/auth/auth_provider.py:530](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_mcp/src/nat/plugins/mcp/auth/auth_provider.py#L530); [packages/nvidia_nat_core/src/nat/authentication/oauth2/oauth2_auth_code_flow_provider_config.py:34](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/authentication/oauth2/oauth2_auth_code_flow_provider_config.py#L34) (verified)
    - *To reach the next level:* Scopes are operator-chosen and static per provider, not narrowed per tool or per request.
  - **C L1:** Only clients that reference an auth_provider (MCP streamable-http, A2A) use it; built-in tools such as the GitHub group keep using ambient environment tokens. Evidence: [packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py:53](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py#L53) (verified)
    - *To reach the next level:* Built-in tools and stdio MCP servers do not go through the per-user identity.
  - **D L0:** The auth provider is off unless the operator configures it on each client. Evidence: [packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py:53](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py#L53) (verified)
    - *To reach the next level:* Not enabled by default.
  - **B L2:** When used, a hijacked call is limited to what the signed-in user's OAuth grant allows on that one remote service; inferred from the per-user flow, scope depends on the operator's provider configuration. Evidence: [packages/nvidia_nat_mcp/src/nat/plugins/mcp/auth/auth_provider.py:462](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_mcp/src/nat/plugins/mcp/auth/auth_provider.py#L462) (inferred)
    - *To reach the next level:* Grants are not limited to read-only or to one project, and the agent's own policy is not intersected with the user's grant.
- **Cap:** G1: Opt-in mechanism: off in the scored default configuration.

### C2 Approval gates: 0.15 (high confidence)

Nothing in the default workflow asks a human before a tool runs: built-in agents call GitHub write operations, code execution, search and MCP tools directly. The toolkit ships a human-in-the-loop middleware base class that can pause before or after a function, but it is abstract (a developer must subclass and register it), it is attached only to the functions it is configured for, and the prompt it shows is fixed text rather than the actual call and arguments. A workflow as shipped performs every consequential action unattended.

- **default configuration** (default; raw 0.00 → 0.00)
  - **S L0:** The agent's tool node invokes the selected tool directly and the default workflow configuration contains no middleware, so no approval step exists. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py:242](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py#L242); [packages/nvidia_nat_core/src/nat/data_models/config.py:284](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/data_models/config.py#L284) (verified)
    - *To reach the next level:* No approval of any kind before consequential tool calls.
  - **C L0:** The most powerful tools (code execution, GitHub commit, MCP tools) run without crossing any gate. Evidence: [packages/nvidia_nat_core/src/nat/builder/function.py:200](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/builder/function.py#L200) (verified)
    - *To reach the next level:* The most powerful tool paths are not gated.
  - **D L0:** Approval is opt-in: no concrete approval middleware is registered or configured by default. Evidence: searched `rg -n -i 'hitl'` in `packages/nvidia_nat_core/src/nat/middleware/register.py` → 0 hits (the middleware registry imports cache, circuit breaker, dynamic, logging and timeout only; no HITL type is registered) (verified)
    - *To reach the next level:* Approval is not on by default.
  - **B L0:** Built-in actions include creating issues and pull requests and committing to a remote branch through the GitHub API, which are external and not undone by the toolkit. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:244](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L244) (verified)
    - *To reach the next level:* No checkpoints, previews or dry-runs for external actions.
- **opt-in HITL middleware base class** (alt; raw 0.15, cap G1 → 0.15) ← counted
  - **S L1:** HITLMiddleware shows a configured static prompt before (or after) each intercepted call and leaves the decision to a subclass; the approver is not shown the call's arguments. Evidence: [packages/nvidia_nat_core/src/nat/middleware/hitl/hitl_middleware.py:95-100](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/middleware/hitl/hitl_middleware.py#L95-L100) (verified)
    - *To reach the next level:* The approver does not see the exact call and arguments, and there are no risk tiers.
  - **C L1:** The middleware intercepts only the functions it is attached to; auto-registration on workflow functions defaults to off, and tools not wrapped as toolkit functions are not seen. Evidence: [packages/nvidia_nat_core/src/nat/middleware/dynamic/dynamic_middleware_config.py:121](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/middleware/dynamic/dynamic_middleware_config.py#L121) (verified)
    - *To reach the next level:* Not every tool path traverses the gate; unknown tools are not rejected.
  - **D L0:** The class is abstract and must be subclassed, registered and configured; pre_invoke_prompt defaults to None, making it a no-op. Evidence: [packages/nvidia_nat_core/src/nat/middleware/hitl/hitl_middleware_config.py:31](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/middleware/hitl/hitl_middleware_config.py#L31) (verified)
    - *To reach the next level:* Not on by default.
  - **B L0:** A wrongly approved call reaches the same irreversible external actions; nothing adds undo or rate limits on approvals. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:323](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L323) (verified)
    - *To reach the next level:* No rollback, previews or rate limits on consequential actions.
- **Cap:** G1: Opt-in mechanism: off in the scored default configuration.

### C3 Tool & action scoping: 0.50 (high confidence)

Every toolkit function validates its input against a typed schema before it runs, and a few built-in tools go further: the GitHub commit tool resolves paths and keeps them inside the configured repository directory, and repository and branch names are checked before they are used in URLs. Other tools take free-form input, such as arbitrary Python for the code-execution tool or arbitrary queries for search tools, and there is no shared allowlist or policy layer. No tools are enabled by default; each agent receives only the tools named in its configuration.

- **S L2:** Inputs are converted to typed schemas on every invocation, and the GitHub tools add resolved-path containment and URL-component checks, but general tools such as code execution accept raw code. Evidence: [packages/nvidia_nat_core/src/nat/builder/function.py:198](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/builder/function.py#L198); [packages/nvidia_nat_core/src/nat/tool/github_tools.py:340-342](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L340-L342); [packages/nvidia_nat_core/src/nat/tool/github_tools.py:44-45](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L44-L45) (verified)
  - *To reach the next level:* No allowlist validation for most tools; general-purpose inputs (code, search queries, MCP arguments) pass through unchecked beyond their schema.
- **C L2:** Schema conversion applies to every toolkit function including MCP tools, but stronger argument checks exist only in the GitHub tools. Evidence: [packages/nvidia_nat_core/src/nat/builder/function.py:290](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/builder/function.py#L290) (verified)
  - *To reach the next level:* No shared validation layer beyond type schemas that extension tools inherit.
- **D L3:** Agents start with an empty tool list and receive only the tools named in tool_names; the model cannot add tools at runtime, and function groups can be narrowed with include. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/react_agent/register.py:67-68](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/react_agent/register.py#L67-L68); [packages/nvidia_nat_core/src/nat/data_models/function.py:50](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/data_models/function.py#L50) (verified)
  - *To reach the next level:* A listed function group (an MCP server, the GitHub group) exposes all of its functions unless the operator narrows it, so tool sets are not per-task by default.
- **B L1:** Configured tools are broad: arbitrary code in the sandbox, repository-wide GitHub writes, and arbitrary search queries, with only minor limits such as output truncation. Evidence: [packages/nvidia_nat_core/src/nat/tool/code_execution/register.py:39](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/register.py#L39); [packages/nvidia_nat_core/src/nat/tool/github_tools.py:323](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L323) (verified)
  - *To reach the next level:* Tools are not quantity-bounded or scoped to a narrow resource.
- **Cap:** none

### C4 Code-execution isolation: 0.62 (medium confidence)

The only built-in way for the model to run code is the code-execution tool, which never runs code on the host: it sends the code to a separately deployed Piston sandbox server and returns an error if that server is unreachable or times out. The toolkit does not ship or harden the sandbox itself, so how strong the isolation is depends on how the operator deploys Piston, and the client asks for no memory limit. No shell tool or local Python interpreter is exposed to the model.

- **S L2:** Code runs out of process on a remote Piston server over HTTP; Piston's own isolation is not in this repository, so its strength is inferred rather than verified. Evidence: [packages/nvidia_nat_core/src/nat/tool/code_execution/register.py:35-37](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/register.py#L35-L37); [packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py:170-175](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py#L170-L175) (inferred)
  - *To reach the next level:* The toolkit does not ship or verify a hardened sandbox (non-root, dropped capabilities, seccomp, no network); isolation is whatever the operator's Piston deployment provides.
- **C L3:** The code-execution tool is the only model-reachable execution path and it always goes to the sandbox; no shell tool or in-process interpreter was found, and stdio MCP servers run commands from operator configuration, not model output. Evidence: searched `rg -n -i 'python_repl|PythonREPL|CommandLineCodeExecutor|code_interpreter|shell_tool|ShellTool|bash_tool|run_command'` in `packages` → 1 hits (the single hit is a test of the Responses API agent passing OpenAI's provider-hosted code_interpreter tool, which runs on the provider side; no shell or local interpreter tool in any shipped package); [packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py:129-133](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py#L129-L133) (verified)
  - *To reach the next level:* Processes spawned by extensions (stdio MCP servers, plugins) run on the host outside any sandbox.
- **D L3:** Piston is the only supported sandbox type and there is no option to execute on the host; the sandbox address comes from operator configuration, not from the model. Evidence: [packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py:172-173](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py#L172-L173); [packages/nvidia_nat_core/src/nat/tool/code_execution/register.py:37](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/register.py#L37) (verified)
  - *To reach the next level:* The sandbox policy itself (limits, network) lives in the operator's separate deployment, outside the toolkit's control.
- **B L2:** Only the code string is sent to the sandbox (no host credentials or files), with a 10-second run timeout, but the client disables Piston's memory limit and network reach depends on the deployment. Evidence: [packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py:166](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/code_sandbox.py#L166); [packages/nvidia_nat_core/src/nat/tool/code_execution/register.py:38](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/register.py#L38) (inferred)
  - *To reach the next level:* No memory limit is requested and egress from the sandbox is not restricted by the toolkit.
- **Cap:** none

### C5 Untrusted input blast radius: 0.15 (high confidence)

Tool results, web search results, GitHub issue text and MCP tool output all go back to the model as ordinary tool messages, with nothing marking them as untrusted and nothing restricting what the agent does after reading them. The optional security package adds LLM-based checks (a pre-tool verifier, content-safety guard, output verifier and PII filter) and a NeMo Guardrails middleware, but these are detection layers, must be configured, and by default only log a violation rather than block it. In a workflow that combines web or issue content with GitHub write tools or search queries, a successful injection can both send data out and make changes without a human.

- **default configuration** (default; raw 0.00, cap C5-WORSTCASE → 0.00)
  - **S L0:** Tool output is wrapped in a ToolMessage with the raw response as content; no taint tracking, approval-after-untrusted-read or quarantine exists in the default agents. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py:255](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py#L255) (verified)
    - *To reach the next level:* Nothing in code limits what the agent may do after it reads untrusted content.
  - **C L0:** Untrusted sources are not distinguished from the principal's instructions in the default agents. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py:242-255](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py#L242-L255) (verified)
    - *To reach the next level:* No untrusted source is handled differently from the user's input.
  - **D L0:** No containment is on by default; the default workflow configuration has no middleware. Evidence: [packages/nvidia_nat_core/src/nat/data_models/config.py:284](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/data_models/config.py#L284) (verified)
    - *To reach the next level:* Not on by default.
  - **B L0:** Documented workflows pair untrusted input (web search, GitHub issues) with credentialed write tools (issue and pull request creation, commits) and outbound channels (search queries), all unattended. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:244](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L244); [packages/nvidia_nat_langchain/src/nat/plugins/langchain/tools/tavily_internet_search.py:37](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/tools/tavily_internet_search.py#L37) (verified)
    - *To reach the next level:* Exfiltration and irreversible external actions both happen without human approval after untrusted content is read.
- **opt-in security package defense middleware** (alt; raw 0.15, cap G1 → 0.15) ← counted
  - **S L1:** The pre-tool verifier and related defenses ask an LLM to classify inputs or outputs for injection and policy violations, which is detection, not a structural limit. Evidence: [packages/nvidia_nat_security/src/nat/plugins/security/middleware/defense/defense_middleware_pre_tool_verifier.py:93-94](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_security/src/nat/plugins/security/middleware/defense/defense_middleware_pre_tool_verifier.py#L93-L94) (verified)
    - *To reach the next level:* Detection only; no Rule-of-Two enforcement or quarantined-model design.
  - **C L1:** Each defense instance is attached to a configured target function or group, so coverage depends on which tools the operator wires it to. Evidence: [packages/nvidia_nat_security/src/nat/plugins/security/middleware/defense/defense_middleware.py:98](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_security/src/nat/plugins/security/middleware/defense/defense_middleware.py#L98) (verified)
    - *To reach the next level:* Tool descriptions, MCP results and peer-agent messages are not covered unless each is targeted explicitly.
  - **D L0:** The defenses are opt-in, and even when configured the default action is partial_compliance, which logs a violation and lets the call proceed. Evidence: [packages/nvidia_nat_security/src/nat/plugins/security/middleware/defense/defense_middleware.py:87-88](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_security/src/nat/plugins/security/middleware/defense/defense_middleware.py#L87-L88) (verified)
    - *To reach the next level:* Not on by default, and the default action does not block.
  - **B L0:** A classifier miss leaves the same unattended egress and write tools available. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:323](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L323) (verified)
    - *To reach the next level:* No human approval for egress or irreversible actions after a miss.
- **Cap:** G1: Opt-in mechanism: off in the scored default configuration.

### C6 Memory, context & configuration integrity: 0.10 (high confidence)

Memory is optional, but when it is used the model can store anything through the add-memory tool, and the auto-memory wrapper saves every user and agent message by default and feeds retrieved memories back as a system message, the highest-priority context. Memory is kept per user: the user identity comes from trusted configuration or the runtime session, never from the model, and operations fail if no identity is available. There is no validation, expiry or review of stored memories, so injected text can persist across a user's sessions and steer later tool use. Workflow configuration comes only from the file the operator names; nothing in a working directory is loaded as instructions.

- **S L0:** The auto-memory wrapper re-injects retrieved memories as a SystemMessage, and the add-memory tool stores model-supplied text without validation. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py:168](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py#L168); [packages/nvidia_nat_core/src/nat/tool/memory_tools/add_memory_tool.py:64](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/memory_tools/add_memory_tool.py#L64) (verified)
  - *To reach the next level:* Memory writes are not gated or validated, and retrieved memories are presented as trusted system context rather than data.
- **C L0:** No memory or retrieval path applies write controls or provenance. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py:55-57](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py#L55-L57) (verified)
  - *To reach the next level:* No memory store has a write or load control.
- **D L1:** Memory operations are namespaced by a user identity taken from trusted configuration or the session context and fail closed when none is available; rated L1 because the write mechanism itself is L0 and D may be at most one level above S. Evidence: [packages/nvidia_nat_core/src/nat/tool/memory_tools/common.py:81](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/memory_tools/common.py#L81); [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py:78](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py#L78) (verified)
  - *To reach the next level:* Namespacing is solid, but the rating is held down by the uncontrolled write path; retention limits are not on by default.
- **B L1:** Poisoned memories persist across the same user's sessions and are injected before the agent plans, so they can trigger tool use; per-user namespacing keeps them from other users. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py:144-168](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/auto_memory_wrapper/agent.py#L144-L168) (verified)
  - *To reach the next level:* Poisoned memories are not limited to text output or gated actions, and there is no review or rollback.
- **Cap:** none
- **Notes:** Workflow YAML is loaded only from the path the operator passes, with `base:` inheritance and `file://` prompt includes restricted to text extensions. The CLI calls `load_dotenv()` at import; python-dotenv's default search starts from the calling module's directory and walks up (inferred library behaviour), so it reads a .env above the install location, not the current working directory. Not treated as repo-controlled configuration.

### C7 Third-party extensions: 0.25 (high confidence)

Extensions come in two forms: Python plugin packages, which the toolkit discovers through package entry points and imports into its own process at startup, and MCP servers, which the operator lists in the workflow file with the exact command or URL. Nothing is pinned, hashed or re-approved when a server's tools or a package's code change, and every installed plugin package is loaded automatically once present. Hugging Face models do default to not running remote code. A malicious plugin runs with everything the agent holds; stdio MCP servers run as separate processes on the host.

- **S L1:** Extensions come from sources the operator chooses (installed packages, MCP commands and URLs in YAML) with no version pinning or integrity check in the loader. Evidence: [packages/nvidia_nat_core/src/nat/runtime/loader.py:195](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/runtime/loader.py#L195); [packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py:47](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py#L47) (verified)
  - *To reach the next level:* No pinning, hash or signature verification for plugins or MCP servers.
- **C L1:** Only model loading has a safe default (trust_remote_code off for Hugging Face LLMs and embedders); plugins and MCP servers are not verified. Evidence: [packages/nvidia_nat_core/src/nat/llm/huggingface_llm.py:88](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/llm/huggingface_llm.py#L88); [packages/nvidia_nat_core/src/nat/embedder/huggingface_embedder.py:52](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/embedder/huggingface_embedder.py#L52) (verified)
  - *To reach the next level:* Plugins and MCP servers have no verification.
- **D L2:** MCP servers are added only by editing the workflow file, which states the command, but any installed package that declares a toolkit entry point is imported automatically at startup without being shown or confirmed. Evidence: [packages/nvidia_nat_core/src/nat/runtime/loader.py:133](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/runtime/loader.py#L133) (verified)
  - *To reach the next level:* Installed plugin packages are enabled on discovery rather than through an explicit step that shows what will run.
- **B L0:** Plugin packages are imported into the agent's own process, with its credentials, environment and network. Evidence: [packages/nvidia_nat_core/src/nat/runtime/loader.py:195](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/runtime/loader.py#L195) (verified)
  - *To reach the next level:* Plugins are not confined to a separate process with a scrubbed environment.
- **Cap:** none
- **Notes:** Stdio MCP servers are launched with the configured env (default None); the MCP Python SDK then supplies its default minimal environment (inferred library behaviour), so they run as separate processes as the same OS user with a reduced environment.

### C8 Secrets & sensitive-data protection: 0.25 (high confidence)

API keys are read from environment variables or the workflow file into secret-typed fields that hide their value when printed, but no redaction applies to logs by default and trace redaction is an opt-in setting. Anonymous CLI usage telemetry is content-free and asked for at first run, but the prompt defaults to yes. The README's hello-world sets verbose agent logging, which writes tool inputs and outputs to the console unredacted. Keys are long-lived and as broad as whatever the operator exported.

- **S L1:** Secrets come from environment variables into SecretStr-typed config fields, which mask their repr; there are no log filters on the main paths and trace redaction defaults to off. Evidence: [packages/nvidia_nat_core/src/nat/llm/nim_llm.py:39](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/llm/nim_llm.py#L39); [packages/nvidia_nat_core/src/nat/observability/mixin/redaction_config_mixin.py:26](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/observability/mixin/redaction_config_mixin.py#L26) (verified)
  - *To reach the next level:* No log filters or redaction on the main logging and trace paths.
- **C L1:** Masking covers object representations only; logs and trace exports are not redacted by default. Evidence: [packages/nvidia_nat_core/src/nat/observability/mixin/redaction_config_mixin.py:26-28](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/observability/mixin/redaction_config_mixin.py#L26-L28) (verified)
  - *To reach the next level:* Logs, traces and other output paths are not covered by any default redaction.
- **D L1:** Telemetry is content-free but the first-run prompt treats an empty answer as consent, and verbose agent logging (used in the README example) prints tool inputs and outputs without redaction. Evidence: [packages/nvidia_nat_core/src/nat/utils/telemetry/consent.py:263](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/utils/telemetry/consent.py#L263); [README.md:129](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/README.md#L129); [packages/nvidia_nat_core/src/nat/data_models/agent.py:31](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/data_models/agent.py#L31) (verified)
  - *To reach the next level:* Telemetry is not strictly opt-in and verbose logs are unredacted.
- **B L1:** Keys are long-lived (LLM API keys, GitHub personal tokens) and as broad as the operator made them; model-generated code runs remotely and does not see them, but every in-process plugin does. Evidence: [packages/nvidia_nat_core/src/nat/tool/github_tools.py:226-227](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/github_tools.py#L226-L227) (verified)
  - *To reach the next level:* Credentials are not scoped per task or short-lived.
- **Cap:** none

### C9 Audit & traceability: 0.50 (high confidence)

Every toolkit function call, including MCP tools and agents used as tools, emits structured start and end events with inputs, outputs, timing and parent links, and the toolkit can export them as OpenTelemetry spans (with user and session attribution) to a file or to many tracing backends. None of that is recorded by default: no exporter is configured unless the operator adds one, and the default console logging does not record tool calls. When enabled, exports are batched and best-effort.

- **S L3:** Function start and end events carry structured inputs, outputs and ancestry, and span exporters attach user.id and session.id from the runtime context. Evidence: [packages/nvidia_nat_core/src/nat/builder/context.py:268](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/builder/context.py#L268); [packages/nvidia_nat_core/src/nat/observability/exporter/span_exporter.py:232-234](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/observability/exporter/span_exporter.py#L232-L234) (verified)
  - *To reach the next level:* No tamper-evident record of its own; integrity depends on the backend the operator picks.
- **C L3:** Events are emitted in the shared function invocation path, so built-in tools, function-group tools (including MCP) and nested agents are all covered. Evidence: [packages/nvidia_nat_core/src/nat/builder/context.py:284](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/builder/context.py#L284); [packages/nvidia_nat_core/src/nat/builder/function.py:195-196](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/builder/function.py#L195-L196) (verified)
  - *To reach the next level:* Configuration changes, memory writes and credential use are not recorded as distinct audit events.
- **D L0:** The tracing section defaults to empty, so events are not persisted anywhere unless the operator configures an exporter, and per-call agent logging is off by default. Evidence: [packages/nvidia_nat_core/src/nat/data_models/config.py:157](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/data_models/config.py#L157) (verified)
  - *To reach the next level:* Recording is opt-in.
- **B L1:** Exporters batch spans and flush on an interval, so records are written after the actions they describe and can be lost on a crash. Evidence: [packages/nvidia_nat_core/src/nat/observability/mixin/batch_config_mixin.py:22-23](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/observability/mixin/batch_config_mixin.py#L22-L23) (verified)
  - *To reach the next level:* Records are not flushed durably per action.
- **Cap:** G1: Trace and audit export exists but no exporter is configured by default.

### C10 Limits & kill switch: 0.42 (high confidence)

The built-in ReAct and tool-calling agents stop after 15 tool calls by default, enforced through the LangGraph recursion limit, and some tools carry their own timeouts (10 seconds for code execution, 60 seconds for MCP calls). There is no wall-clock or cost limit for a run, model requests have no timeout by default, and the timeout and circuit-breaker middleware are optional. An agent used as a tool by another agent gets its own fresh budget, so nesting multiplies the limit.

- **S L2:** An iteration cap is enforced in code, and some tools have per-execution timeouts; there is no run-level token, cost or wall-clock cap. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/react_agent/register.py:49](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/react_agent/register.py#L49); [packages/nvidia_nat_core/src/nat/tool/code_execution/register.py:38](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/tool/code_execution/register.py#L38); [packages/nvidia_nat_core/src/nat/llm/openai_llm.py:56](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/llm/openai_llm.py#L56) (verified)
  - *To reach the next level:* No token or cost cap and no timeouts on every kind of step (model calls default to no timeout).
- **C L1:** The cap applies to each agent's own loop; tool timeouts exist only on some tools, and nested agents do not share the parent's budget. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/tool_calling_agent/register.py:169](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/tool_calling_agent/register.py#L169); [packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py:102-103](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_mcp/src/nat/plugins/mcp/client/client_config.py#L102-L103) (verified)
  - *To reach the next level:* Not every tool has a timeout, and sub-agents do not count against the parent's limit.
- **D L2:** Defaults are sensible (15 tool calls) and set by the operator; the model cannot raise them, but it can reach a fresh budget by calling an agent configured as a tool. Evidence: [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/react_agent/register.py:79](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/react_agent/register.py#L79); [packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/tool_calling_agent/register.py:87](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/tool_calling_agent/register.py#L87) (verified)
  - *To reach the next level:* Delegation to a nested agent resets the budget.
- **B L2:** A runaway is bounded at a moderate number of tool calls per agent, but there is no spend ceiling and the optional timeout middleware only stops waiting on a call (asyncio wait_for) rather than killing work. Evidence: [packages/nvidia_nat_core/src/nat/middleware/timeout/timeout_middleware.py:68-71](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/c7e1162a1c7ff18bbd797e090a56cad97c281c92/packages/nvidia_nat_core/src/nat/middleware/timeout/timeout_middleware.py#L68-L71) (verified)
  - *To reach the next level:* No per-run time or cost ceiling, and pending work is not cancelled on stop.
- **Cap:** none

## Rule-of-Two check
[A] untrusted input: Web search, GitHub issue and file content, MCP tool results (packages/nvidia_nat_langchain/src/nat/plugins/langchain/agent/base.py:255) · [B] sensitive data/systems: Environment API keys and GitHub tokens, per-user memory (packages/nvidia_nat_core/src/nat/tool/github_tools.py:226) · [C] state change / egress: GitHub issue, pull request and commit writes; outbound search queries (packages/nvidia_nat_core/src/nat/tool/github_tools.py:244) · Same default session? Yes

## Highest-impact improvements
1. Ship a concrete, registered approval middleware that shows the exact tool name and arguments and attach it by default to tools that write or send (GitHub writes, code execution, MCP tools). (C2 S L0→L3, +0.225 before caps; Playbook 5)
2. Persist function start/end events by default (for example a local file exporter outside the working directory) so every run leaves a record. (C9 D L0→L2, +0.100 before caps; Playbook 1 step 3)
3. After a tool returns untrusted content (search, GitHub issue text, MCP results), require approval for egress and write tools for the rest of the run. (C5 S L0→L3, +0.225 before caps; Playbook 1)
4. Present retrieved memories as data in a tool or user-role message instead of a system message, and require approval or validation for memory writes. (C6 S L0→L2, +0.150 before caps; Playbook 2)
5. Add a run-level wall-clock and token budget that nested agents share, and give model calls a default timeout. (C10 S L2→L3, +0.075 before caps; Playbook 3 step 3)

## Re-audit log
- No changes.

## Limitations
- Static source review of the pinned commit only; nothing was executed, installed, or probed.
- Scored as a framework from the default arguments of its first-class features; the many framework adapter packages (CrewAI, LlamaIndex, Semantic Kernel, ADK, Agno, AutoGen, Strands), the evaluation, optimizer and fine-tuning tooling, and the separate NeMo-Agent-Toolkit-UI repository were only sampled or not reviewed.
- Isolation of the code-execution tool depends on the operator's Piston deployment, which is not part of this repository; C4 strength and blast radius are inferred.
- Behaviour of third-party libraries (python-dotenv search path, MCP SDK default subprocess environment, LangGraph recursion limit) is inferred from their documented behaviour, not read at a pinned version.
- The repository's AGENTS.md and skills/ directory contain instructions for coding agents working on the repo; they were treated as data and no text aimed at reviewers was found in the files read.
