BoundBench

IDA Pro MCP

MCP server bridging IDA Pro with LLMs for reverse engineering (incl. debugger ops)

github.com/mrexodia/ida-pro-mcp · 2026-10-03 · c133c38

Defense-in-depth score

3.8 / 10

Minimal

As the Claude Code plugin launches it, IDA Pro MCP is a local stdio server with no authentication and no per-call gating: the code-execution and debugger tools are withheld, but the database-editing and file-saving tools are on, and nothing separates attacker-controlled binary strings from instructions. Enforcement of that tier does not cover every path, and workers keep running after the supervisor exits. A read-only profile exists but is opt-in; the legacy GUI plugin enables every tool, including in-process Python, by default.

Key gaps (1)

  1. Python execution tools run model code in-process with no isolation, and spawned workers inherit the full environment. C4 · Code-execution isolation

Criteria

C1 Identity & least privilege

Minimal 0.15 / 1.00

The server holds no credentials of its own and never authenticates a caller: the stdio supervisor and the loopback HTTP workers it spawns accept any request, and the workers keep running after the supervisor exits. Nothing narrows the operating-system user's authority; idb_open will open any path the user can read and idb_save will write a database copy to any path they can write. The only default narrowing is that the code-execution and debugger tools are withheld unless an operator passes a command-line flag. A hijacked agent therefore acts with the full file authority of the user, limited to what the exposed tools can do.

C2 Approval gates

Minimal 0.40 / 1.00

As a tool server the project leaves approval to the host, so what matters is the risk signalling it gives the host. It publishes no readOnlyHint or destructiveHint annotations, so a host cannot tell the byte-patching, renaming and save tools from the read-only ones. The one server-enforced tier is the unsafe set (Python execution and the debugger), which is withheld by default in the headless server; however enforcement of that tier does not cover every path. Mutating tools such as patch, patch_asm, put_int and idb_save are not in the unsafe set, and only rename offers a dry run.

C3 Tool & action scoping

Moderate 0.50 / 1.00

Tools have typed schemas generated from Python type hints and addresses are parsed through a helper, but there is no allowlist validation of paths: idb_open accepts any path, idb_save writes to any path, and py_exec_file (unsafe, withheld by default) executes any file. In the default headless configuration the code-execution and debugger tools are removed, which leaves analysis and database-editing tools; the editing tools (patch, rename, set_type, patch_asm) are on by default. Everything is scoped to the open IDA databases, with full write access inside them.

C4 Code-execution isolation

Minimal 0.05 / 1.00

There is no containment for code execution. When the Python tools are enabled, model-supplied code runs with exec and eval inside the IDA process with the full builtins and the process's complete environment and file access; the debugger tools launch the analyzed binary under IDA. By default the headless server withholds those tools, which is a real reduction in exposure but not isolation, and that enforcement does not cover every path. The spawned worker processes inherit the whole parent environment, so any code that does run can read the user's secrets.

C5 Untrusted input blast radius

Minimal 0.20 / 1.00

Everything the server returns is derived from the binary under analysis, which is attacker-controlled in malware work: strings, symbol names, comments and decompiled text. Outputs are structured JSON with schemas, but nothing marks content as untrusted or records its source, and the server offers no mode that cuts off state changes after untrusted content has been read, other than an opt-in read-only profile. The server has no network tool and holds no credentials, so exfiltration depends on the host; irreversible effects are limited to overwriting a file via idb_save and editing the database.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

Model edits persist: names, comments, types and patches are saved into the IDA database and re-appear in later sessions' tool output with no provenance, and every call is logged into the same database. The server also keeps some of its own configuration inside the IDA database, which is the untrusted artifact under analysis, and that configuration is not integrity-protected.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The server loads no third-party plugins, models, packages or remote tools at runtime; the supervisor launches only its own worker module and, for GUI sessions, can adopt instances listed in files under the user's home directory (a local, same-user mechanism noted below). The absence is by design rather than a control, so the criterion is marked as structurally absent.

C8 Secrets & sensitive-data protection

Minimal 0.40 / 1.00

The server accepts no API keys or tokens and has no telemetry or crash reporting, so there is little secret material of its own to leak. It does not scrub anything either: spawned workers inherit the parent's whole environment, request parameters are logged truncated at debug level, and every tool call with its full arguments and result is stored in the IDA database. The model can also read any file the user can open by loading it with idb_open.

C9 Audit & traceability

Minimal 0.45 / 1.00

Every tools/call handled by a worker is recorded with tool name, arguments, structured result, duration and timestamp into an append-only gzipped log inside the IDA database, and this tracer is always on. It does not record who asked or who approved, the supervisor's own idb_open and idb_close calls are not recorded, and records are buffered in batches, with write errors swallowed silently, so a crash can lose recent entries. The log lives in the database file the same process edits.

C10 Limits & kill switch

Minimal 0.40 / 1.00

The server bounds its own work: tool calls on the worker have a 60 second default timeout with native cancellation, outputs over 50,000 characters are truncated into a cached download, request bodies are capped at 10 MB, and the supervisor limits simultaneous databases to four and gives each forwarded call 900 seconds. There are no rate limits, and workers are deliberately detached so they keep running with full tool authority after the supervisor stops, until an idle timeout the caller can lengthen.