C1 Identity & least privilege
Minimal 0.00 / 1.00
UFO drives the Windows desktop through UI Automation as the logged-in user, so every app, file and signed-in account (browser sessions, Outlook, Office) is within reach. Nothing narrows that authority: there is no separate low-privilege identity, no per-action authorization check, and launched programs inherit the user's environment. The device server and Galaxy web UI do require an API token, which stops strangers from issuing tasks, but once a task runs it acts with the user's full session.
C2 Approval gates
Minimal 0.17 / 1.00
A well-built confirmation prompt exists: it shows the exact function and arguments, defaults to 'no', and the approved batch is snapshotted so what runs is what was approved. But it only fires when the model itself labels an action 'CONFIRM', as instructed by the prompt. The HostAgent's confirmation is an empty TODO and Galaxy's orchestrator always auto-confirms, so typing, clicking 'Send' or launching apps proceeds without a human whenever the model doesn't ask.
C3 Tool & action scoping
Minimal 0.28 / 1.00
The non-GUI tools are carefully scoped: the CLI tool launches only a bare name from a fixed app allowlist with no arguments and no shell, the Linux tool uses an allowlist with per-command argument policies and pinned executables, and Office save paths are resolved and contained. But the main tools are general GUI primitives (click anywhere, type any key sequence into any window), which no validation can bound, and all of them are enabled by default.
C4 Code-execution isolation
Minimal 0.13 / 1.00
There is no sandbox: tools run as the desktop user on the host. The command-line tool is limited to launching allowlisted apps without arguments, and the Linux tool runs allowlisted read-only commands with shell disabled, but these are filters, not isolation. The GUI keyboard tool can type arbitrary text into any open window, including a PowerShell or terminal window, so the model can run any command around the allowlist; the shipped device config even suggests launching PowerShell windows.
C5 Untrusted input blast radius
Minimal 0.25 / 1.00
UFO reads whatever is on screen, including web pages, emails and documents, and passes UI text and screenshots to the model alongside the user's task. A regex sanitizer strips some injection phrases and wraps the user request and retrieved documents in tags, but the on-screen control text is passed through raw. Because the same session can read untrusted content, see private data and send email or type into a browser without approval, a successful injection can both leak data and take irreversible actions.
C6 Memory, context & configuration integrity
Minimal 0.35 / 1.00
Long-term memory features exist (experience and demonstration retrieval stores, saved Q&A customization) but are all off by default, and configuration is loaded only from UFO's own install directory rather than from a workspace the agent works on. When experience saving is enabled in 'always' or 'auto' mode, summaries of past runs are written without review and later re-injected as examples. Nothing tags retrieved memories with provenance.
C7 Third-party extensions
Minimal 0.28 / 1.00
By default UFO only runs its own in-repo MCP servers; no third-party extension is loaded. Operators can add stdio MCP servers by command in mcp.yaml, with no pinning or integrity check, and the optional retrieval features load FAISS indexes with pickle deserialization explicitly allowed. Added stdio servers get only the environment configured for them (inferred from the MCP SDK's default environment handling).
C8 Secrets & sensitive-data protection
Minimal 0.25 / 1.00
LLM API keys live in plaintext in config/ufo/agents.yaml (git-ignored) or environment variables, and the whole process environment is merged into the config object. There is no telemetry, but every session writes full prompts and screenshots to logs/ by default with no redaction. Some masking exists (config validator display, sensitive env-var blocks in the shell client), but not on the main log or model paths.
C9 Audit & traceability
Minimal 0.45 / 1.00
Every step is written as a structured JSON record (action, arguments, results, timings, and whether the user confirmed) to logs/<task>/response.log, flushed on each write, alongside the full LLM requests. The logs sit in UFO's working directory where the desktop agent itself could edit them, carry no actor or approver identity, and are not tamper-evident.
C10 Limits & kill switch
Minimal 0.45 / 1.00
Step limits are enforced in code (50 steps per UFO session, 15 for Galaxy, one round), LLM calls time out at 60 seconds, and Galaxy caps concurrent tasks at six. But each device session gets its own fresh step budget, there is no cost cap, and the MCP tool timeout is 6000 seconds (the comment says 5 minutes) with timed-out calls left running in a thread pool. A stop-task handler exists in the Galaxy web UI.