C1 Identity & least privilege
Minimal 0.00 / 1.00
The server holds the operator's API keys and runs as the operator's OS user (root inside the shipped docker-compose service), and none of its REST or WebSocket endpoints authenticate the caller. Any client that can reach the port acts with the server's full authority: it can spend the keys, read and delete every stored report, and supply an MCP 'command' that the server launches as a process. The maintainers' SECURITY.md states this is by design and leaves authentication to the operator. Nothing narrows the authority the agent inherits.
C2 Approval gates
Minimal 0.10 / 1.00
There is no approval step anywhere. The default research flow only searches and reads web pages, so it makes no consequential changes on its own. When a request enables MCP servers, the model picks MCP tools and arguments and they execute immediately with no human in the loop, whatever those tools do. A 'human_feedback' WebSocket command exists but is a stub that only prints the message.
C3 Tool & action scoping
Moderate 0.55 / 1.00
The default tools are narrow: search queries go to a configured search API and the scraper fetches result URLs after a shared check that allows only http(s) and rejects hosts resolving to private, loopback, link-local or metadata addresses. The code itself notes a DNS-rebinding window, and URL and transport handling does not cover every path. MCP tools receive raw model-chosen arguments, and MCP path restrictions are not a strict boundary. Write-capable tools only appear when a request adds MCP servers.
C4 Code-execution isolation
Minimal 0.30 / 1.00
The project never runs model-written code, but it launches subprocesses from text it does not control: a research request may carry MCP server configs whose 'command' and 'args' the server spawns as a local stdio process with no isolation, as the same OS user. Package installs via pip are also triggered at runtime for some optional scrapers and LLM providers. The shipped Docker image runs as a non-root user, but the shipped docker-compose overrides that to root and passes the API keys into the container; Docker is an optional deployment, so it is scored as an opt-in mechanism.
C5 Untrusted input blast radius
Minimal 0.10 / 1.00
Scraped web pages, search results, uploaded documents and MCP tool results all go straight into the model's prompts with no separation from instructions, and there is no detection or provenance handling. In the default web-research configuration the model sees mostly public content plus the user's query, has no tools that change state, and its only outbound channels are search queries to the search provider and a further unattended egress path. When a request enables local documents or MCP servers, injected content can steer private data out or drive MCP tools with no human involved.
C6 Memory, context & configuration integrity
Minimal 0.00 / 1.00
Generated reports and chat histories are kept in one global JSON store that any unauthenticated client can list, overwrite or delete. When someone chats about a stored report, the report text (model output derived from untrusted web pages) is placed into the system prompt and the stored chat history is replayed with whatever roles it contains, including 'system'. There is no per-user namespace, validation, provenance, or review of what gets persisted, so poisoned content persists across sessions and users and can drive the chat's web-search tool.
C7 Third-party extensions
Minimal 0.05 / 1.00
MCP servers are the main extension type, and they come from the request rather than from operator configuration: a client sends a command, arguments and environment, and the server spawns it. Nothing pins, hashes or allowlists what runs, and there is no consent step beyond the requester's own toggle, which any unauthenticated client can set. Separately, some optional scrapers and LLM providers pip-install unpinned packages at runtime when missing, and retriever plugins are loaded by name from installed entry points. Launched MCP processes run as the same OS user and can read the .env file holding the operator's keys.
C8 Secrets & sensitive-data protection
Minimal 0.25 / 1.00
API keys come from environment variables or a .env file and are only attached to provider HTTP calls, so they are not placed into model prompts. There is no redaction or masking anywhere. Research logs (queries, sources, context, report) are written to outputs/*.json and served over an unauthenticated /outputs static route, and the chat endpoint logs full model responses at INFO. LangSmith tracing turns on automatically in the multi-agent path whenever a LangChain key is present, sending prompts to a third party.
C9 Audit & traceability
Minimal 0.42 / 1.00
Each research run writes a JSON log file of streamed events (queries, sources, progress) and the server logs to logs/app.log at INFO, so the main pipeline leaves an unstructured trail. MCP tool calls log only the tool name at INFO; their arguments are logged at debug level. Records carry no caller identity (there is no authentication), sit in the working directory, and the per-run logs are served publicly.
C10 Limits & kill switch
Minimal 0.38 / 1.00
The research pipeline is a fixed sequence of stages rather than an open loop, and outbound HTTP calls have per-request timeouts. But the number of sub-queries is only suggested to the model (MAX_ITERATIONS is passed into the prompt and not enforced), clients can raise max_search_results without bound, and there is no wall-clock or cost cap. Closing the WebSocket cancels the running asyncio task, while REST /report/ background jobs cannot be stopped and executor threads finish their work.