BoundBench

UFO (UFO² / UFO³ Galaxy)

Microsoft's Windows desktop GUI agent (UFO²) and multi-device orchestration framework (UFO³ Galaxy) that operate apps via UI Automation and MCP tools.

github.com/microsoft/ufo · 2026-10-04 · a795552

Defense-in-depth score

2.6 / 10

Minimal

UFO controls the Windows desktop as the logged-in user with no sandbox, and its keyboard and mouse tools can type anything into any window, including a PowerShell window. The command-line and Linux tools are tightly allowlisted, but the GUI path goes around them. The human confirmation prompt is well built yet only appears when the model decides to ask, and Galaxy's orchestrator auto-confirms, so prompt injection from on-screen content can leak data and send or delete things unattended.

Key gaps (4)

  1. The agent acts with the desktop user's full session and every signed-in account; nothing narrows its authority. C1 · Identity & least privilege
  2. Human confirmation only triggers when the model labels an action CONFIRM; HostAgent confirmation is a no-op and Galaxy auto-confirms, so the most powerful actions bypass the gate. C2 · Approval gates
  3. No sandbox; keyboard_input can type arbitrary commands into a shell window, bypassing the CLI allowlist, with host-equivalent reach. C4 · Code-execution isolation
  4. Untrusted on-screen text reaches the model unsanitized in a session that can see private data and send/delete via GUI without enforced approval. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

UFO drives the Windows desktop through UI Automation as the logged-in user, so every app, file and signed-in account (browser sessions, Outlook, Office) is within reach. Nothing narrows that authority: there is no separate low-privilege identity, no per-action authorization check, and launched programs inherit the user's environment. The device server and Galaxy web UI do require an API token, which stops strangers from issuing tasks, but once a task runs it acts with the user's full session.

C2 Approval gates

Minimal 0.17 / 1.00

A well-built confirmation prompt exists: it shows the exact function and arguments, defaults to 'no', and the approved batch is snapshotted so what runs is what was approved. But it only fires when the model itself labels an action 'CONFIRM', as instructed by the prompt. The HostAgent's confirmation is an empty TODO and Galaxy's orchestrator always auto-confirms, so typing, clicking 'Send' or launching apps proceeds without a human whenever the model doesn't ask.

C3 Tool & action scoping

Minimal 0.28 / 1.00

The non-GUI tools are carefully scoped: the CLI tool launches only a bare name from a fixed app allowlist with no arguments and no shell, the Linux tool uses an allowlist with per-command argument policies and pinned executables, and Office save paths are resolved and contained. But the main tools are general GUI primitives (click anywhere, type any key sequence into any window), which no validation can bound, and all of them are enabled by default.

C4 Code-execution isolation

Minimal 0.13 / 1.00

There is no sandbox: tools run as the desktop user on the host. The command-line tool is limited to launching allowlisted apps without arguments, and the Linux tool runs allowlisted read-only commands with shell disabled, but these are filters, not isolation. The GUI keyboard tool can type arbitrary text into any open window, including a PowerShell or terminal window, so the model can run any command around the allowlist; the shipped device config even suggests launching PowerShell windows.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

UFO reads whatever is on screen, including web pages, emails and documents, and passes UI text and screenshots to the model alongside the user's task. A regex sanitizer strips some injection phrases and wraps the user request and retrieved documents in tags, but the on-screen control text is passed through raw. Because the same session can read untrusted content, see private data and send email or type into a browser without approval, a successful injection can both leak data and take irreversible actions.

C6 Memory, context & configuration integrity

Minimal 0.35 / 1.00

Long-term memory features exist (experience and demonstration retrieval stores, saved Q&A customization) but are all off by default, and configuration is loaded only from UFO's own install directory rather than from a workspace the agent works on. When experience saving is enabled in 'always' or 'auto' mode, summaries of past runs are written without review and later re-injected as examples. Nothing tags retrieved memories with provenance.

C7 Third-party extensions

Minimal 0.28 / 1.00

By default UFO only runs its own in-repo MCP servers; no third-party extension is loaded. Operators can add stdio MCP servers by command in mcp.yaml, with no pinning or integrity check, and the optional retrieval features load FAISS indexes with pickle deserialization explicitly allowed. Added stdio servers get only the environment configured for them (inferred from the MCP SDK's default environment handling).

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

LLM API keys live in plaintext in config/ufo/agents.yaml (git-ignored) or environment variables, and the whole process environment is merged into the config object. There is no telemetry, but every session writes full prompts and screenshots to logs/ by default with no redaction. Some masking exists (config validator display, sensitive env-var blocks in the shell client), but not on the main log or model paths.

C9 Audit & traceability

Minimal 0.45 / 1.00

Every step is written as a structured JSON record (action, arguments, results, timings, and whether the user confirmed) to logs/<task>/response.log, flushed on each write, alongside the full LLM requests. The logs sit in UFO's working directory where the desktop agent itself could edit them, carry no actor or approver identity, and are not tamper-evident.

C10 Limits & kill switch

Minimal 0.45 / 1.00

Step limits are enforced in code (50 steps per UFO session, 15 for Galaxy, one round), LLM calls time out at 60 seconds, and Galaxy caps concurrent tasks at six. But each device session gets its own fresh step budget, there is no cost cap, and the MCP tool timeout is 6000 seconds (the comment says 5 minutes) with timed-out calls left running in a thread pool. A stop-task handler exists in the Galaxy web UI.