BoundBench

ContextForge MCP Gateway

Open-source registry and proxy that federates MCP, A2A and REST/gRPC tools behind one authenticated MCP endpoint with RBAC, admin UI and observability.

github.com/IBM/mcp-context-forge · 2026-10-04 · 9bb9c2a

Defense-in-depth score

3.7 / 10

Minimal

ContextForge puts real access control in front of the tools it federates: every call is authenticated, checked against RBAC and team visibility, schema-validated, SSRF-filtered, rate-limited and timed out, and stored upstream credentials are encrypted and never shown to the model. It is still a relay rather than a firewall: tool results and upstream tool descriptions reach clients untouched, any user can register a new upstream server that is shared with everyone by default, and a client-supplied X-Upstream-Authorization header is always forwarded as the upstream credential. The audit trail and every content guardrail plugin ship switched off, so out of the box you cannot reconstruct who ran which tool with which arguments.

Key gaps (3)

  1. Client-supplied X-Upstream-Authorization is always forwarded upstream as the Authorization header (MCP token passthrough); other upstream credential forwarding is not strictly scoped. C1 · Identity & least privilege
  2. An escape from the in-process Jinja2 template sandbox lands in the gateway process that holds the JWT secret, the credential encryption key and every stored upstream credential. C4 · Code-execution isolation
  3. Upstream tool output and descriptions are relayed with no provenance or filtering by default, so a hijacked client can read through one registered server and act or exfiltrate through another with shared stored credentials. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

Every tool call on both the /rpc and /mcp paths is authenticated and checked against the caller's RBAC permission and team-scoped visibility before any upstream credential is attached, which is a genuine per-principal authorization layer. The credentials themselves are mostly static per upstream server and shared by every user who can see that server, and by default a newly registered server is public and every team role, including viewer, may execute tools. The gateway also always forwards a client-supplied X-Upstream-Authorization header as the upstream Authorization, which is MCP token passthrough and caps this criterion. A compromise exposes write access to every registered upstream system.

C2 Approval gates

Minimal 0.20 / 1.00

As a gateway, ContextForge leaves approval to the MCP host and only relays whatever readOnly/destructive hints the upstream server or the registering user supplied; it never derives them, so REST-wrapped tools that POST or DELETE usually carry no risk signal. It does offer a dry-run preview endpoint that validates arguments and resolves the target without calling it, but nothing forces a host to use it. Annotations can be set or changed by whoever registers or updates the tool, including the upstream server itself. Once a call is approved by the host, the gateway executes it against the upstream system with no undo.

C3 Tool & action scoping

Moderate 0.53 / 1.00

Every tool invocation, whatever the backend, is validated against the tool's JSON input schema in a killable worker before anything is sent, and REST tool URLs are re-validated against SSRF rules and pinned to the resolved IP after argument substitution, blocking cloud metadata, localhost and private ranges by default. Validation is only as tight as each tool's schema, and URL construction for upstream requests is not strictly constrained. The gateway ships no tools of its own, but anything registered is immediately callable by every user who can see it, and virtual servers are the only way to narrow the set.

C4 Code-execution isolation

Minimal 0.33 / 1.00

The gateway never launches MCP servers or shell commands itself; the code it interprets is prompt templates (rendered with Jinja2's in-process SandboxedEnvironment) and jq filters on tool output, which run in forked worker processes with the environment cleared and a 2-second kill timer. Both are restriction layers inside the gateway's own user account rather than an OS boundary, and on non-Linux hosts or with one config switch jq runs in-process with no timeout. A template-sandbox escape lands inside the gateway process, which holds the JWT signing key, the credential encryption key and every stored upstream credential.

C5 Untrusted input blast radius

Minimal 0.07 / 1.00

Tool results, upstream tool descriptions and A2A agent replies are relayed to clients exactly as received, with no provenance marking or untrusted-content flag the host could act on. The guardrail plugins that could filter content (deny lists, PII, moderation) are all disabled by default and are detection filters even when on. A client steered by malicious upstream content can use the same gateway session to read through one registered server and write or exfiltrate through another with stored credentials, and the gateway does nothing to stop that combination.

C6 Memory, context & configuration integrity

Minimal 0.35 / 1.00

What persists is the shared catalog of tools, prompts and resources, including tool descriptions copied from upstream servers when they are registered, plus per-user LLM chat history in Redis that expires after an hour. Writes to the catalog require an authenticated user with create permission, and upstream changes are not pulled in automatically, but stored descriptions carry no provenance and newly registered servers default to public visibility, so one user's registration is served to every user's model. Configuration comes from operator files (.env in the working directory, plugins/config.yaml), not from untrusted workspaces.

C7 Third-party extensions

Minimal 0.40 / 1.00

The extensions ContextForge loads are remote MCP servers and A2A agents that users register, plus optional Python plugins that are disabled by default. Remote servers run elsewhere and receive only their own configured credentials, so a malicious one cannot read the gateway's secrets, but nothing pins or verifies what a server serves, and any user who owns a personal team can register one that is public to everyone. Tool definitions are snapshotted at registration and only refreshed on request, which limits silent changes to the tool list but not to the server's behaviour.

C8 Secrets & sensitive-data protection

Moderate 0.65 / 1.00

Upstream credentials are encrypted at rest with an Argon2id-derived Fernet key, attached server-side at call time and never placed in model context; logs default to ERROR level, payload logging and tracing are opt-in, and trace exports redact token, password and authorization fields. The weak points are the credentials themselves: they are long-lived, broadly scoped upstream keys, and the client-supplied X-Upstream-Authorization header is forwarded to the upstream server.

C9 Audit & traceability

Minimal 0.38 / 1.00

Out of the box the gateway records only per-tool metrics rows (tool, timestamp, latency, success) with no caller identity or arguments, buffered and flushed once a minute; the audit trail, security event log and permission audit are all off, and the default ERROR log level drops the INFO lines that name each invocation. After an incident you could see that a tool ran, but not who ran it or with what input, unless the operator had turned the audit features on.

C10 Limits & kill switch

Moderate 0.57 / 1.00

Every tool call has a server-enforced timeout (60 seconds by default, overridable per tool), the MCP and tools endpoints are rate-limited per user at 100 requests a minute with lockout, and on the JSON-RPC path a cancellation request actually cancels the in-flight task. The advertised tool_rate_limit and tool_concurrent_limit settings are not enforced anywhere, there is no response-size cap on JSON tool results, and the embedded LLM chat agent sets no step limit of its own and registers a no-op cancel handler.