C1 Identity & least privilege
Moderate 0.50 / 1.00
DeerFlow's tools do not hold external credentials by default: web search is keyless, page fetches go through Jina, and model provider keys stay inside the gateway process. Everything the agent touches on disk is placed under per-user, per-thread directories derived from the logged-in user, and thread routes check ownership. The fine-grained authorization layer (RBAC per tool, model, and sandbox) exists but ships disabled, so every authenticated user can use every tool. MCP servers an admin adds run with one shared credential for all users, so they are not scoped to the requesting user.
C2 Approval gates
Minimal 0.05 / 1.00
There is no human approval step for any tool call in the default configuration. File writes and web fetches run as soon as the model asks for them; the only 'approval' is the model choosing to call ask_clarification, which is a prompt instruction. The optional guardrail middleware is an automated allow/deny policy, not a human gate, and is off by default. Because the default tool set has no shell, email, or chat channel, the actions that slip through are mostly file overwrites inside the user's thread and outbound page fetches.
C3 Tool & action scoping
Moderate 0.50 / 1.00
The file tools are well bounded: paths must be virtual /mnt/user-data paths, '..' is rejected, and the resolved real path must stay inside the thread's workspace, uploads, or outputs directory. Host bash is removed from the tool list unless an operator opts in. The web_fetch tool is a general fetcher: any URL the model chooses is sent to Jina's reader with no local host allowlist, so it is an open outbound channel. MCP and plugin tools get no shared validation layer.
C4 Code-execution isolation
Minimal 0.33 / 1.00
By default DeerFlow uses its local sandbox provider with host bash switched off, so the model has no code-execution tool unless an operator enables one. If an operator sets allow_host_bash, commands run directly on the host (or inside the gateway container) as the gateway user, guarded only by a best-effort path filter the code itself says is not a security boundary, though secret-looking environment variables are scrubbed. The stronger option is the opt-in AIO Docker sandbox: per-thread containers with seccomp, memory/CPU/PID limits, and only the thread's directories mounted, but no forced non-root user and open network by default. Admin-configured MCP stdio servers always run on the gateway host.
C5 Untrusted input blast radius
Minimal 0.38 / 1.00
DeerFlow treats fetched web content and MCP results as untrusted and neutralizes framework tags (like a forged system reminder) before the model sees them, and it wraps user input in delimiters. This is detection-style hardening: nothing stops a hijacked agent from acting. Uploaded documents and file reads are not neutralized. In the default setup a successful injection from a web page can read the user's uploads, workspace, and memory and send them out through web_fetch to any URL with no human involved; there is no shell, email, or chat channel by default, so irreversible high-impact actions are limited.
C6 Memory, context & configuration integrity
Minimal 0.42 / 1.00
Long-term memory is on by default: after each turn a background model call extracts facts into a per-user memory file that is injected into later conversations. Extraction reads only user messages and the agent's final replies (tool results are excluded), injected memory is escaped and framed as user-managed data, stale facts are reviewed for removal, and users can view, edit, and delete facts in the UI. There is no human approval before a fact is saved, so injected content echoed in a final reply can persist and steer later tool use. When custom agents are enabled, the update_agent tool lets the agent rewrite its own persistent persona file and tool-group list without a check.
C7 Third-party extensions
Minimal 0.25 / 1.00
No third-party extension is enabled by default. Admins can add MCP servers (the API limits stdio launchers to npx/uvx and rejects eval-style flags), operators can list Python plugins in config.yaml, and users can install skill archives that pass a deterministic scan plus a model-based scan. None of these are pinned or integrity-checked; the shipped example launches MCP packages with 'npx -y' at whatever version is latest. Plugins load in-process with the gateway's full authority, and MCP stdio servers run as the gateway user, though only with the environment variables configured for them.
C8 Secrets & sensitive-data protection
Moderate 0.57 / 1.00
Secret handling is careful. Request-scoped secrets travel out of band and never enter prompts or tool arguments, secret-looking variables are scrubbed from sandbox subprocess environments, URLs and credential headers are redacted in logs by default, managed-model keys are typed as secrets, and config validation errors log only the error type. There is no telemetry, and tracing to LangSmith or Langfuse is opt-in. Keys still live in plaintext in .env and the gateway environment for the life of the process, with no secret manager or encryption at rest.
C9 Audit & traceability
Minimal 0.45 / 1.00
Every tool call the lead agent makes, with its arguments and result, is stored in the thread's LangGraph checkpoint, which defaults to a SQLite database outside the agent's reachable directories. A richer run journal records caller identity (lead agent vs. sub-agent) and token use, but its default backend is in memory and is lost on restart. Records are not tamper-evident and can be deleted with the thread; there is no OpenTelemetry or SIEM export by default.
C10 Limits & kill switch
Moderate 0.50 / 1.00
Each run is capped at 100 graph steps by default (clients can raise it up to a hard 1,000), loop detection stops repeated tool calls, web fetches time out after 10 seconds, and sub-agents are limited to 6 per run, 3 at once, with their own turn and time limits. The token budget exists but is off by default, and there is no run-level wall-clock limit. Stopping a run cancels its asyncio task; synchronous tool work already in flight finishes.