# Defense-in-Depth Score: LiteLLM

**Repo:** https://github.com/BerriAI/litellm · **Commit:** `a2bf67a03707e474be41e07806ee4fd9791ca7cd` (1.105.0) · **Reviewed:** 2026-10-04
**What it is:** Open-source AI gateway (proxy server) and Python SDK that unifies 100+ LLM providers behind an OpenAI-compatible API, with virtual keys, spend tracking, guardrails, an MCP gateway and A2A agent routing.
**Category:** Agent Frameworks
**Scored configuration:** LiteLLM proxy (AI gateway) via the shipped docker-compose or litellm --config with a master key set, admin-registered HTTP MCP servers with default fields, and callers using default request fields (require_approval unset).
**Agent surface (default):** code execution opt-in · filesystem write no · network egress yes · external credentials yes · persistent memory yes · untrusted input yes · third party extensions opt-in · sub agents opt-in · external communication yes

## Score: 4.5 / 10.0 (Minimal)

| # | Criterion | S | C | D | B | Raw | Cap | Score | Confidence |
|---|---|---|---|---|---|---|---|---|---|
| C1 | Identity & least privilege | L2 | L3 | L2 | L1 | 0.53 | — | **0.53** | High |
| C2 | Approval gates | L1 | L2 | L1 | L1 | 0.33 | — | **0.33** | High |
| C3 | Tool & action scoping | L2 | L3 | L2 | L1 | 0.53 | — | **0.53** | High |
| C4 | Code-execution isolation | L3 | L3 | L3 | L2 | 0.70 | — | **0.70** | High |
| C5 | Untrusted input blast radius | L2 | L1 | L0 | L1 | 0.28 | G1 | **0.28** (alt) | High |
| C6 | Memory, context & configuration integrity | L2 | L2 | L2 | L2 | 0.50 | — | **0.50** | Medium |
| C7 | Third-party extensions | L2 | L1 | L0 | L0 | 0.23 | G1 | **0.23** (alt) | High |
| C8 | Secrets & sensitive-data protection | L1 | L2 | L2 | L1 | 0.38 | — | **0.38** | High |
| C9 | Audit & traceability | L2 | L2 | L3 | L1 | 0.50 | — | **0.50** | High |
| C10 | Limits & kill switch | L3 | L3 | L0 | L2 | 0.55 | G1 | **0.50** (alt) | High |


LiteLLM's gateway puts strong authentication and per-key authorization in front of every model and MCP tool call: it refuses to boot without a real master key, MCP servers are deny-by-default per key, and model-written code only runs in a remote sandbox. The dominant risk is what happens once tools run server-side: the README's lead example sets require_approval to never, after which the gateway executes model-chosen MCP calls with shared upstream credentials, for up to five rounds on streaming, with no taint tracking and no default spend cap. Config-named plugins and callbacks load in-process with no integrity check, and the shipped docker-compose defaults are not locked down.

## Critical gaps
- Config-named callback, guardrail and router-plugin modules (including ones fetched from S3/GCS) are executed in the gateway process with no integrity check, giving a malicious plugin every provider key, the master/salt key and the database. (ASI04, T17, LLM03; C7) — [litellm/proxy/types_utils/utils.py:22-29](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L22-L29); [litellm/proxy/types_utils/utils.py:167](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L167)

## Criterion details

### C1 Identity & least privilege — 0.53 (high)

The gateway refuses to start without a strong master key, and every request is then authorised in code against the calling virtual key, team, user and organisation before the gateway attaches its own credentials. MCP servers are deny-by-default per key (only servers granted to the key or team, or marked allow_all_keys, are reachable), and per-key tool permissions are re-checked on every tool call. The credentials behind that check are static and shared: the operator's provider API keys and each MCP server's single configured credential serve every caller, and a key with no model list can call every model. Per-user OAuth, token exchange and passthrough modes exist for MCP but are opt-in.

- **default configuration** (default; raw 0.53 → 0.53) ← counted
  - **S L2:** A deterministic per-request authorization layer (models, MCP servers, tools) sits in front of static, shared provider and MCP credentials; a key with an empty model list may call every model. — [litellm/proxy/auth/auth_checks.py:4448](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/auth_checks.py#L4448); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2479](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2479); [litellm/proxy/auth/master_key_boot_check.py:129](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/master_key_boot_check.py#L129) (verified)
    - *To reach the next level:* Credentials are not narrowed per tool or per caller by default; per-user OAuth / token exchange for MCP is opt-in and provider keys are always shared.
  - **C L3:** Every MCP tool call passes the same pre-call authorization (server allow/deny list, per-key/team tool permission, allowed params), and server access is computed per key with allow_all_keys off by default. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5753](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5753); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2665](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2665); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5735](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5735) (verified)
    - *To reach the next level:* Entitlement lookups that fault early leave no ceiling for key auth (documented fail-open), and tool-permission enforcement does not cover every path.
  - **D L2:** Boot refuses an unset, empty or publicly known master key unless an explicit dangerously_* override is set; new keys default to all models and no MCP servers. — [litellm/proxy/auth/master_key_boot_check.py:89](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/master_key_boot_check.py#L89); [litellm/proxy/auth/master_key_boot_check.py:129](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/master_key_boot_check.py#L129); [litellm/proxy/auth/auth_checks.py:4448](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/auth_checks.py#L4448) (verified)
    - *To reach the next level:* Defaults are not near-minimal: a new key can call every configured model, and widening (allow_all_keys, model lists) is an unlogged admin config change.
  - **B L1:** A hijacked key or gateway session acts with the operator's long-lived provider keys and each MCP server's shared credential, so write access spans every registered system. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2479](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2479); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2665](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2665) (verified)
    - *To reach the next level:* Credentials would need to be per-tenant or per-user and short-lived to keep a hijack to one tenant.
- **opt-in per-user MCP auth (oauth2 per user, RFC 8693 token exchange)** (alt; raw 0.47, cap G1 → 0.47)
  - **S L3:** MCP auth types include per-user OAuth and RFC 8693 token exchange that mint a credential for the requesting user. — [litellm/types/mcp.py:50-53](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/types/mcp.py#L50-L53) (verified)
    - *To reach the next level:* Credentials are not task-scoped and revoked after use.
  - **C L2:** Applies only to MCP servers configured with those auth types; provider keys and other servers stay shared. — [litellm/types/mcp.py:50-53](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/types/mcp.py#L50-L53); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2479](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2479) (verified)
    - *To reach the next level:* Provider calls and default-auth MCP servers do not use per-principal credentials.
  - **D L0:** The default auth_type is unset (shared/no credential); per-user modes must be configured per server. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2479](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2479) (verified)
    - *To reach the next level:* Per-user credentials would need to be the default MCP auth posture.
  - **B L2:** With per-user credentials a hijack is limited to the requesting user's grants on that server. — [litellm/types/mcp.py:50-53](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/types/mcp.py#L50-L53) (verified)
    - *To reach the next level:* Tokens are long-lived OAuth grants rather than minutes-lived task tokens.
- **Cap:** none
- **Notes:** C1-PASSTHRU not applied: true_passthrough / oauth_delegate (forwarding the client's token upstream) exist but are opt-in per server.

### C2 Approval gates — 0.33 (high)

When a caller attaches gateway MCP tools to /chat/completions or /responses, LiteLLM executes the model's tool calls itself only if every MCP tool reference in the request sets require_approval to "never"; otherwise the exact tool calls are handed back to the calling application unexecuted. That check is deterministic and the model cannot set it, but it is a blanket per-request switch held by the caller, there is no human approval step or risk tiering in the gateway, and the README's lead MCP example turns auto-execution on against a GitHub MCP server. Once enabled, upstream tool actions are as irreversible as the upstream makes them.

- **S L1:** Auto-execution is one blanket per-request opt-in (require_approval == "never" on every reference); otherwise exact calls are returned to the app, and the gateway itself offers no human approval or risk tiers. — [litellm/responses/mcp/litellm_proxy_mcp_handler.py:561](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L561); [litellm/responses/mcp/litellm_proxy_mcp_handler.py:540](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L540) (verified)
  - *To reach the next level:* No per-call human approval showing the exact call, and no risk tiers deciding which tools need one.
- **C L2:** All three server-side execution paths (Responses, Responses streaming, chat completions) use the same check before calling global_mcp_server_manager.call_tool. — [litellm/responses/mcp/litellm_proxy_mcp_handler.py:561](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L561); [litellm/responses/mcp/litellm_proxy_mcp_handler.py:852](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L852); [litellm/responses/mcp/mcp_streaming_iterator.py:927](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/mcp_streaming_iterator.py#L927) (verified)
  - *To reach the next level:* Operator-enabled interception features (code interpreter, web search) run tools without this check, and unknown tools are not rejected by default.
- **D L1:** Off unless the caller opts in, but the README's flagship MCP example sets require_approval: never, so documented use routinely disables the gate (lowered one level per the examples rule). — [README.md:245](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/README.md#L245); [litellm/responses/mcp/litellm_proxy_mcp_handler.py:561](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L561) (verified)
  - *To reach the next level:* Disabling auto-execution approval would need an operator-level flag rather than a per-request field shown in the lead example.
- **B L1:** Auto-executed calls go to whatever the upstream MCP server does (issues, messages, writes) with no rollback; per-key tool permissions and allowed_params still narrow what can run. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5753](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5753); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5760](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5760) (verified)
  - *To reach the next level:* No checkpoints, previews, or quantity bounds on consequential MCP actions.
- **Cap:** none

### C3 Tool & action scoping — 0.53 (high)

Tool scoping is by name: each MCP server can carry an allowed or disallowed tool list, an optional allowed-parameter-name list per tool, and pinned tool definitions, and per-key and per-team tool permissions narrow further. All of these checks run in one pre-call function that every gateway tool call goes through. Argument values are not validated against bounds or allowlists, and by default every tool an upstream server lists is exposed to any key that has the server. User-supplied URLs fetched by the gateway (images, files) go through an SSRF guard that rechecks redirects.

- **S L2:** Name-level allow/deny lists and parameter-name allowlists are enforced in code; argument values pass through unvalidated. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5735](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5735); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5760](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5760); [litellm/litellm_core_utils/url_utils.py:92](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/litellm_core_utils/url_utils.py#L92) (verified)
  - *To reach the next level:* No value-level validation (bounds, path/URL/recipient allowlists) on MCP tool arguments.
- **C L3:** pre_call_tool_check applies the same name, pin, permission and parameter checks to every gateway MCP tool call, including auto-executed ones. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5735](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5735); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5743](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5743); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5753](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5753); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5760](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5760) (verified)
  - *To reach the next level:* Coverage is capped one level above strength; checks are name-level only.
- **D L2:** allowed_tools defaults to None, so a registered server exposes every upstream tool (including write tools) to keys granted the server. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2658](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2658); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2665](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2665) (verified)
  - *To reach the next level:* Read-only tool sets are not the default; write tools need no explicit enabling.
- **B L1:** A misused tool reaches whatever the upstream server can do with the gateway's shared credential. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2479](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2479) (verified)
  - *To reach the next level:* No quantity bounds or workspace scoping on tool operations.
- **Cap:** none

### C4 Code-execution isolation — 0.70 (high)

In the default gateway nothing executes model-written code: Jinja chat templates render in Jinja's immutable sandbox, operator-written custom-code guardrails run under RestrictedPython, and stdio MCP servers are disabled unless an environment flag is set. The one model-code path, the opt-in code-interpreter interception callback, always runs code in a remote sandbox service (E2B or OpenSandbox) and raises an error rather than falling back to local execution. Those sandboxes get internet access by default, and stdio MCP servers, when enabled, run as host subprocesses (root in the shipped image) with a scrubbed environment.

- **S L3:** Model-emitted code only runs through a remote sandbox provider API (E2B microVM or OpenSandbox containers); isolation strength depends on the provider chosen. — [litellm/integrations/code_interpreter_interception/handler.py:740-743](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/integrations/code_interpreter_interception/handler.py#L740-L743); [litellm/integrations/code_interpreter_interception/handler.py:758](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/integrations/code_interpreter_interception/handler.py#L758) (verified)
  - *To reach the next level:* OpenSandbox may be stock containers; isolation is not guaranteed kernel-level for every supported provider.
- **C L3:** The only model-code path requires a sandbox and fails closed; stdio MCP servers (operator commands) run on host only when LITELLM_ENABLE_MCP_STDIO=true. — [litellm/integrations/code_interpreter_interception/handler.py:758](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/integrations/code_interpreter_interception/handler.py#L758); [litellm/proxy/_experimental/mcp_server/stdio_gate.py:14-15](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/stdio_gate.py#L14-L15); [litellm/litellm_core_utils/prompt_templates/factory.py:14](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/litellm_core_utils/prompt_templates/factory.py#L14) (verified)
  - *To reach the next level:* Processes spawned by extensions (stdio MCP servers) are not sandboxed.
- **D L3:** The sandbox is mandatory whenever code interpretation is enabled, and the stdio host-exec path needs an explicit operator environment flag. — [litellm/proxy/common_utils/callback_utils.py:165](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/common_utils/callback_utils.py#L165); [litellm/proxy/_experimental/mcp_server/stdio_gate.py:14-15](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/stdio_gate.py#L14-L15) (verified)
  - *To reach the next level:* Sandbox policy (e.g. internet access) is set in operator config with permissive defaults rather than locked outside configuration.
- **B L2:** Sandboxes hold no gateway secrets but E2B internet access defaults to on; stdio servers run as the container user (root in the image) with a scrubbed environment. — [litellm/llms/e2b/sandbox/transformation.py:66](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/llms/e2b/sandbox/transformation.py#L66); [litellm/experimental_mcp_client/client.py:510-514](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/experimental_mcp_client/client.py#L510-L514); [Dockerfile:132](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/Dockerfile#L132) (verified)
  - *To reach the next level:* Network egress from the sandbox is not off or allowlisted by default.
- **Cap:** none
- **Notes:** The SDK's litellm.agent harness (Claude Code, Codex, OpenCode in sandbox.local or Docker) is not exposed by the proxy and was not scored.

### C5 Untrusted input blast radius — 0.28 (high)

MCP tool results and upstream tool descriptions enter the model's context like any other message, with no provenance tagging or taint tracking by default. When auto-execution is on, the streaming Responses path lets the model choose further tool calls after reading tool output, for up to five rounds, so injected tool output can drive actions with the shared MCP credentials. By default the gateway returns tool calls to the calling application instead of running them. An opt-in tool_policy guardrail blocks tools marked trusted-input once untrusted tool output is in a chat conversation, and many third-party injection-detection guardrails can be enabled.

- **default configuration** (default; raw 0.05 → 0.05)
  - **S L0:** Nothing structural limits a hijacked session by default; the streaming auto-execute loop re-runs model-chosen tools after reading tool output. — [litellm/responses/mcp/mcp_streaming_iterator.py:35](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/mcp_streaming_iterator.py#L35); [litellm/responses/mcp/mcp_streaming_iterator.py:927](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/mcp_streaming_iterator.py#L927) (verified)
    - *To reach the next level:* Once untrusted content is in context, egress and state-changing tools would need to be disabled or forced through approval.
  - **C L0:** Tool results are appended as ordinary tool messages; no source is distinguished. — [litellm/responses/mcp/chat_completions_handler.py:634](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/chat_completions_handler.py#L634) (verified)
    - *To reach the next level:* Untrusted sources (tool results, tool descriptions, fetched content) are not distinguished from principal input.
  - **D L0:** No untrusted-input control is on by default. — [litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py:167](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py#L167) (verified)
    - *To reach the next level:* A taint or Rule-of-Two control would need to be on by default.
  - **B L1:** By default the gateway only returns tool calls to the app, but the README's never-approval flow lets a hijacked session leak and act via shared MCP credentials unattended; multi-tenant shared credentials lower B one level. — [litellm/responses/mcp/litellm_proxy_mcp_handler.py:561](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L561); [README.md:245](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/README.md#L245); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2479](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L2479) (verified)
    - *To reach the next level:* Exfiltration and irreversible actions should require approval in every documented flow, with per-user credentials.
- **opt-in tool_policy trust-chain guardrail** (alt; raw 0.28, cap G1 → 0.28) ← counted
  - **S L2:** Blocks tools with input_policy=trusted when the conversation contains output from tools marked untrusted. — [litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py:208](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py#L208) (verified)
    - *To reach the next level:* Egress and state change are not generally disabled after untrusted content is read; only tools explicitly marked trusted are blocked.
  - **C L1:** Only inspects chat-completion messages with role tool; Responses input items, tool descriptions and user-supplied documents are not considered. — [litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py:202](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py#L202) (verified)
    - *To reach the next level:* All untrusted sources, including tool descriptions and Responses input, would need to be covered.
  - **D L0:** Opt-in guardrail, and it passes everything when the tool registry is not initialised. — [litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py:167](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/guardrails/guardrail_hooks/tool_policy/tool_policy_guardrail.py#L167) (verified)
    - *To reach the next level:* Would need to be on by default and fail closed.
  - **B L1:** Same worst case as the default for tools not marked trusted. — [litellm/responses/mcp/litellm_proxy_mcp_handler.py:561](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L561); [README.md:245](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/README.md#L245) (verified)
    - *To reach the next level:* Untrusted-content sessions would need no unattended egress.
- **Cap:** G1 — Opt-in mechanism: off in the scored default configuration.

### C6 Memory, context & configuration integrity — 0.50 (medium)

The model has no memory tool and the gateway auto-loads no workspace instruction files; configuration comes from the operator's config file or database. The /v1/memory store is written and read by applications and is filtered to the caller's user or team in the query, and it is never injected into prompts automatically. Responses sessions can be rebuilt from stored spend logs by previous_response_id (prompt content is only stored when store_prompts_in_spend_logs is enabled); session replay is not equally scoped. Cached or replayed context is not validated or provenance-tagged.

- **S L2:** Memory rows are app-managed data scoped per user/team and never auto-injected; replayed session history is re-injected unvalidated. — [litellm/proxy/memory/memory_endpoints.py:84-94](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/memory/memory_endpoints.py#L84-L94); [litellm/responses/litellm_completion_transformation/session_handler.py:295-305](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/litellm_completion_transformation/session_handler.py#L295-L305) (verified)
  - *To reach the next level:* Writes are not gated or expiring, and replayed session context carries no provenance.
- **C L2:** The memory store is scoped; session replay and response caching are not equally controlled. — [litellm/proxy/memory/memory_endpoints.py:84-94](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/memory/memory_endpoints.py#L84-L94) (verified)
  - *To reach the next level:* Session replay and caches would need the same controls.
- **D L2:** Memory queries are restricted to the caller's user_id/team_id by default; session replay is not equally scoped. — [litellm/proxy/memory/memory_endpoints.py:84-94](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/memory/memory_endpoints.py#L84-L94) (verified)
  - *To reach the next level:* Session replay needs the same namespace guarantees.
- **B L2:** Replayed history can steer later turns, including auto-executed tool calls when require_approval is never; content is only stored when prompt storage is enabled. — [litellm/proxy/spend_tracking/spend_tracking_utils.py:1720](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/spend_tracking/spend_tracking_utils.py#L1720); [litellm/responses/litellm_completion_transformation/session_handler.py:295-305](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/litellm_completion_transformation/session_handler.py#L295-L305) (inferred)
  - *To reach the next level:* Poisoned context would need to be session-scoped or easily purged per user.
- **Cap:** none
- **Notes:** SDK footnote: importing litellm calls load_dotenv() when LITELLM_MODE is unset (litellm/__init__.py:30), which can pick up a .env (including API base URLs) from the surroundings of an untrusted checkout; not scored because the proxy's working directory is operator-controlled.

### C7 Third-party extensions — 0.23 (high)

Third-party code reaches the gateway three ways: Python callbacks, guardrails and router plugins named in the config (including modules downloaded from S3 or GCS) are imported and executed in-process with no hash or signature check; stdio MCP servers run operator-chosen commands as subprocesses when explicitly enabled; and remote MCP servers are connected over HTTP. Only proxy admins can add servers (user submissions need admin approval and cannot be stdio), and stdio subprocesses get a minimal environment. Pinning an MCP server's tool catalog, which serves the pinned definitions and alerts on drift, is available but off by default.

- **default configuration** (default; raw 0.17 → 0.17)
  - **S L1:** Config-named modules (local or s3:// / gcs://) are exec'd as-is, and the community router-plugin catalog lists plugins without a PyPI pin. — [litellm/proxy/types_utils/utils.py:22-29](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L22-L29); [litellm/proxy/types_utils/utils.py:167](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L167); [router_plugins.json:22-24](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/router_plugins.json#L22-L24) (verified)
    - *To reach the next level:* No version pinning or integrity check for loaded modules or MCP servers.
  - **C L0:** No extension type is verified by default. — [litellm/proxy/types_utils/utils.py:50](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L50); [litellm/proxy/types_utils/utils.py:167](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L167) (verified)
    - *To reach the next level:* Integrity checks would need to cover at least one extension type by default.
  - **D L2:** Nothing third-party is enabled by default and only PROXY_ADMIN can add MCP servers; stdio submissions from users are refused (capped one level above S). — [litellm/proxy/management_endpoints/mcp_management_endpoints.py:1842](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/management_endpoints/mcp_management_endpoints.py#L1842); [litellm/proxy/management_endpoints/mcp_management_endpoints.py:1419](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/management_endpoints/mcp_management_endpoints.py#L1419); [litellm/proxy/_experimental/mcp_server/stdio_gate.py:14-15](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/stdio_gate.py#L14-L15) (verified)
    - *To reach the next level:* Adding an extension does not show a verified package/hash; capped by the weak verification.
  - **B L0:** Callback and plugin modules run inside the gateway process with every provider key, the master and salt keys, and the database. — [litellm/proxy/types_utils/utils.py:167](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L167); [litellm/proxy/common_utils/encrypt_decrypt_utils.py:29-32](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/common_utils/encrypt_decrypt_utils.py#L29-L32) (verified)
    - *To reach the next level:* Extensions would need to run out of process with only their own configuration.
- **opt-in MCP tool-catalog pinning** (alt; raw 0.23, cap G1 → 0.23) ← counted
  - **S L2:** Pinned servers serve the pinned tool descriptions and schemas and block unpinned tools; drift is detected by comparing description and input schema. — [litellm/proxy/_experimental/mcp_server/tool_catalog_guard.py:122](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/tool_catalog_guard.py#L122); [litellm/proxy/_experimental/mcp_server/tool_catalog_guard.py:140-141](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/tool_catalog_guard.py#L140-L141); [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5743](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5743) (verified)
    - *To reach the next level:* No hash or signature over server code; definitions only.
  - **C L1:** Covers remote/stdio MCP tool definitions only, not Python callbacks or plugins. — [litellm/proxy/_experimental/mcp_server/tool_catalog_guard.py:122](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/tool_catalog_guard.py#L122) (verified)
    - *To reach the next level:* Callbacks, guardrail modules and router plugins are not covered.
  - **D L0:** pinned_tools is unset unless an admin pins the server. — [litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:5743](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/_experimental/mcp_server/mcp_server_manager.py#L5743) (verified)
    - *To reach the next level:* Pinning would need to be the default for new servers.
  - **B L0:** Same confinement as the default: in-process modules hold all gateway secrets. — [litellm/proxy/types_utils/utils.py:167](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/types_utils/utils.py#L167) (verified)
    - *To reach the next level:* Extensions would need process isolation and scoped credentials.
- **Cap:** G1 — Opt-in mechanism: off in the scored default configuration.

### C8 Secrets & sensitive-data protection — 0.38 (high)

Credentials stored in the database are encrypted with a key derived from LITELLM_SALT_KEY (falling back to the master key), virtual keys are stored hashed, log records pass through a credential-redaction filter by default, prompts are not written to spend logs unless enabled, and there is no vendor telemetry. Secrets are not placed in model context, but nothing scans tool results or model-bound messages for secrets. The shipped docker-compose file's defaults are not locked down, and the redaction filter can be turned off with an environment variable.

- **S L1:** Encryption at rest and log redaction exist, but the shipped compose defaults are not locked down. — [litellm/proxy/common_utils/encrypt_decrypt_utils.py:29-32](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/common_utils/encrypt_decrypt_utils.py#L29-L32); [litellm/_logging.py:71](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/_logging.py#L71) (verified)
  - *To reach the next level:* Hardened deployment defaults; redaction before model-bound messages.
- **C L2:** Logs and spend logs are protected and stdio subprocesses get a scrubbed environment; model-bound messages and tool results are not scanned. — [litellm/_logging.py:71](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/_logging.py#L71); [litellm/proxy/spend_tracking/spend_tracking_utils.py:1720](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/spend_tracking/spend_tracking_utils.py#L1720); [litellm/experimental_mcp_client/client.py:510-514](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/experimental_mcp_client/client.py#L510-L514) (verified)
  - *To reach the next level:* Model-bound messages and third-party logging callbacks (turn_off_message_logging defaults False) are not covered.
- **D L2:** No telemetry and no prompt storage by default; secret redaction can be disabled by LITELLM_DISABLE_REDACT_SECRETS. — [litellm/_logging.py:71](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/_logging.py#L71); [litellm/__init__.py:208](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/__init__.py#L208) (verified)
  - *To reach the next level:* Redaction would need to be always on.
- **B L1:** Leaked material is long-lived provider API keys and MCP credentials; encrypted DB values fall to the master key when no salt key is set. — [litellm/proxy/common_utils/encrypt_decrypt_utils.py:29-32](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/common_utils/encrypt_decrypt_utils.py#L29-L32) (verified)
  - *To reach the next level:* Keys would need to be scoped and short-lived or rotatable automatically.
- **Cap:** none

### C9 Audit & traceability — 0.50 (high)

Every gateway MCP tool call is logged through the same logging object as model calls, and the spend-log row carries the tool name, arguments, server, key, user, team and trace id, stored in Postgres outside anything the model can touch. Denied calls are logged as failures. There is no approval record because the gateway has no approval step, and change audit logs for keys, teams and MCP servers are written only for enterprise (premium) deployments. Logging is best-effort: setup failures are swallowed at debug level and spend logs are batched asynchronously.

- **S L2:** Structured per-call records with arguments, status, timestamps and key/user/team attribution. — [litellm/proxy/spend_tracking/spend_tracking_utils.py:266](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/spend_tracking/spend_tracking_utils.py#L266); [litellm/responses/mcp/litellm_proxy_mcp_handler.py:830-833](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L830-L833) (verified)
  - *To reach the next level:* No approver or delegation-chain attribution and no tamper evidence.
- **C L2:** All gateway tool calls and denials are recorded; configuration change audit logs are premium-only. — [litellm/proxy/spend_tracking/spend_tracking_utils.py:266](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/spend_tracking/spend_tracking_utils.py#L266); [litellm/proxy/management_helpers/audit_logs.py:216](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/management_helpers/audit_logs.py#L216) (verified)
  - *To reach the next level:* Approvals and configuration changes are not recorded in the open-source build.
- **D L3:** Spend logging is on by default with a database and written by the gateway, not the model; it can be turned off by an operator setting. — [litellm/proxy/proxy_server.py:6879](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/proxy_server.py#L6879); [litellm/proxy/spend_tracking/spend_tracking_utils.py:266](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/spend_tracking/spend_tracking_utils.py#L266) (verified)
  - *To reach the next level:* Disabling logging is not itself logged.
- **B L1:** Logging initialisation and success-handler failures are caught and the tool call proceeds. — [litellm/responses/mcp/litellm_proxy_mcp_handler.py:812](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L812); [litellm/responses/mcp/litellm_proxy_mcp_handler.py:900](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/litellm_proxy_mcp_handler.py#L900) (verified)
  - *To reach the next level:* Errors would need to be surfaced and records flushed per action.
- **Cap:** none

### C10 Limits & kill switch — 0.50 (high)

Server-side tool execution is bounded: one round on the non-streaming paths and at most five rounds on streaming Responses, MCP calls time out after 60 seconds, and requests after 6000 seconds. Key, team and user budgets, RPM/TPM limits and a per-session iteration limiter are enforced in code but are all unset by default, so out of the box there is no spend ceiling. Blocking a key stops new requests but does not cancel calls already in flight.

- **default configuration** (default; raw 0.40 → 0.40)
  - **S L2:** Round cap plus per-MCP-call and per-request timeouts are enforced by default. — [litellm/responses/mcp/mcp_streaming_iterator.py:35](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/mcp_streaming_iterator.py#L35); [litellm/constants.py:202](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/constants.py#L202); [litellm/constants.py:568](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/constants.py#L568) (verified)
    - *To reach the next level:* Cost/token caps and rate limits are not on by default.
  - **C L2:** Limits apply to the gateway's tool loop and each tool call. — [litellm/responses/mcp/mcp_streaming_iterator.py:927](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/responses/mcp/mcp_streaming_iterator.py#L927); [litellm/constants.py:202](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/constants.py#L202) (verified)
    - *To reach the next level:* Spend caps across delegated/A2A calls are not on by default.
  - **D L1:** Global max_budget is 0 (disabled) and the request timeout default is 100 minutes. — [litellm/__init__.py:408](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/__init__.py#L408); [litellm/constants.py:568](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/constants.py#L568) (verified)
    - *To reach the next level:* Sensible default spend and wall-clock ceilings.
  - **B L1:** With no default budget, spend is unbounded; /key/block only refuses future requests. — [litellm/__init__.py:408](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/__init__.py#L408); [litellm/proxy/management_endpoints/key_management_endpoints.py:7308](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/management_endpoints/key_management_endpoints.py#L7308) (verified)
    - *To reach the next level:* Tight default ceilings and cancellation of in-flight work.
- **opt-in budgets, rate limits and per-session iteration limits** (alt; raw 0.55, cap G1 → 0.50) ← counted
  - **S L3:** Per-key max_budget is checked before each request, RPM/TPM limiters and a per-session iteration cap exist, on top of round caps and timeouts. — [litellm/proxy/auth/auth_checks.py:5526](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/auth_checks.py#L5526); [litellm/proxy/hooks/parallel_request_limiter_v3.py:674](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/hooks/parallel_request_limiter_v3.py#L674); [litellm/proxy/hooks/max_iterations_limiter.py:2-7](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/hooks/max_iterations_limiter.py#L2-L7) (verified)
    - *To reach the next level:* Halt does not interrupt in-flight execution.
  - **C L3:** Budgets and rate limits apply to every request a key makes, including agent and MCP-sampled calls. — [litellm/proxy/auth/auth_checks.py:5526](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/auth_checks.py#L5526); [litellm/proxy/hooks/parallel_request_limiter_v3.py:674](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/hooks/parallel_request_limiter_v3.py#L674) (verified)
    - *To reach the next level:* No cap on concurrent tasks or delegation depth.
  - **D L0:** All are unset unless configured per key/team. — [litellm/__init__.py:408](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/__init__.py#L408) (verified)
    - *To reach the next level:* Would need to be on by default.
  - **B L2:** Budget checks run before a request, so a multi-round auto-execute request can overshoot; blocking does not cancel in-flight calls. — [litellm/proxy/auth/auth_checks.py:5526](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/auth/auth_checks.py#L5526); [litellm/proxy/management_endpoints/key_management_endpoints.py:7308](https://github.com/BerriAI/litellm/blob/a2bf67a03707e474be41e07806ee4fd9791ca7cd/litellm/proxy/management_endpoints/key_management_endpoints.py#L7308) (verified)
    - *To reach the next level:* Provider-side budgets and cancellation of in-flight calls.
- **Cap:** G1 — Opt-in mechanism: off in the scored default configuration.

## Rule-of-Two check
[A] untrusted input: MCP tool results and upstream tool descriptions enter model context (litellm/responses/mcp/chat_completions_handler.py:634) · [B] sensitive data/systems: Shared provider keys and MCP server credentials held by the gateway (litellm/proxy/_experimental/mcp_server/mcp_server_manager.py:2479) · [C] state change / egress: Server-side MCP tool execution when require_approval is never (litellm/responses/mcp/litellm_proxy_mcp_handler.py:852) · Same default session? Yes

## Highest-impact improvements
1. Drop require_approval: never from the README's lead MCP example and require an operator-level allowlist of auto-executable servers/tools. — C2 D L1→L3, +0.100 before caps (Playbook 5)
2. Ship a default global and per-key spend ceiling and a shorter default request timeout. — C10 D L1→L2, +0.050 before caps (Playbook 3 step 3)
3. Harden the shipped docker-compose defaults. — C8 S L1→L2, +0.075 before caps (Playbook 4)
4. Once untrusted tool output is in a streaming auto-execute session, stop executing further tool calls server-side and return them to the caller. — C5 S L0→L2, +0.150 before caps (Playbook 1)
5. Require a pinned version and hash for config-loaded plugin modules, especially s3:// and gcs:// sources. — C7 S L1→L3, +0.150 before caps (Playbook 3)

## Re-audit log
- No changes.

## Limitations
- Static source review of the pinned commit only; nothing was executed, installed, or probed.
- Scored the proxy/AI gateway; the Python SDK (litellm.completion, experimental MCP client, litellm.agent harness with local/Docker sandboxes) is footnoted, not scored.
- The repository is very large (cloned with blobs over 1 MB filtered); enterprise/ code, the admin UI, Helm/Terraform, A2A routing, pass-through endpoints and the 50+ third-party guardrail integrations were not examined in depth.
- Remote sandbox isolation (E2B, OpenSandbox) is provider-side and was not verified beyond the API calls in this repo.
- No reviewer-steering text was found; AGENTS.md contains instructions for contributors' coding agents (e.g. to commit and push without asking), not for reviewers.
