What does this skill do, and when should you use it?
cua-driver is an Agent Skill inside the trycua/cua repository that wraps the Cua Driver's cua-driver CLI and MCP server. It teaches an agent the correct way to operate native GUI applications on the host: observe narrowly first, act through snapshot-bound element tokens, and verify the postcondition at checkpoints. The skill is part of the MIT-licensed Cua Driver codebase and requires the cua-driver binary to be installed on the host beforehand. It is not a VM or sandbox tool — it is an automation interface to a real, authorized desktop.
- Reads a single window narrowly with get_window_state plus a query (~2K chars instead of a ~42K full snapshot), bounding large trees with max_elements/max_depth.
- Batches actions via run_steps with observe:true, stopping at the first failure and returning only what changed since the last read.
- Writes loops, branches, and retries as one sandboxed run_script (JavaScript over the cua API, with time and call limits).
- Verifies postconditions at checkpoints with verify_state or a targeted get_window_state instead of after every action.
- Supports launching apps, opening local files, desktop operation, binding browser page content, and explicitly requested session recording.
- Queries recent Cua activity in a bounded way via history_status/history_query when the user asks to continue or recall prior work.
- Developers using coding agents like Claude Code, Codex, or Cursor who want the agent to complete GUI tasks in real host applications (Calculator, LibreOffice Calc, Inkscape) and verify the displayed result.
- QA or ops staff automating repetitive cross-platform desktop-app actions (menu clicks, form filling, file exports) across macOS, Windows, and Linux.
- Scenarios where the outcome lives in an app's UI state and the user demands GUI-only interaction, excluding DOM/CDP, clipboard APIs, or shell mutations.
- Teams wanting one CLI/MCP interface to drive desktop apps on different operating systems without writing per-platform scripts.
- Users who need an explicitly requested automation run recorded, with artifacts verified, for review or replay reference.
- Users who want to generate training data or run benchmarks inside VMs or cloud sandboxes — that is the job of sibling components Lume and Cua Bench in the same repo, not this skill.
- Automations relying on implicit privilege escalation: the skill forbids unauthorized foreground/desktop control, altering browser security settings, or handling permission prompts on the user's behalf.
- Environments unwilling to install a local binary or grant accessibility/screen-recording permissions — without the cua-driver binary the skill cannot run.
How do you install this skill?
- This is a static source review only; nothing was executed and all scores rest on documented claims, at low confidence.
- Core function requires a locally installed cua-driver binary plus macOS Accessibility/Screen Recording system grants; deployment cost is significant.
- The skill is coupled to overseas resources (cua.ai docs, GitHub); mainland-China reachability is unverified and there is no Chinese-language support statement.
- BROWSER.md is truncated mid-table in the provided files; verify the actual pack completeness.
- Publisher is unverified by the FollowSkills registry and treated as unknown; the Hyprland candidate in LINUX.md is explicitly experimental and ABI-dependent.
- Existing-profile browser binding involves CDP remote-debugging authorization; although the docs declare strict guardrails, audit the implementation before use.
- Shell / CLI
- Network access
- Local filesystem
- MCP Server
cua-driver binary (installed via cua.ai install script)local MCP HTTP bearer token (CUA_DRIVER_RS_MCP_HTTP_TOKEN) if using the MCP endpoint
The Cua Driver binary (which the skill depends on via the cua-driver command) must be installed first. The source documents these host install routes:
macOS / Linux hosts
/bin/bash -c "$(curl -fsSL https://cua.ai/driver/install.sh)"Windows hosts
irm https://cua.ai/driver/install.ps1 | iexVia the unified cua installer (driver only)
curl -fsSL https://cua.ai/install.sh | sh -- --only cua-driverThe skill file itself lives at libs/cua-driver/rust/Skills/cua-driver/SKILL.md in the repo. The README notes the installer offers to install cua skills and the cua MCP server into coding agents (Claude Code, Codex, Cursor, and others), but the exact steps for copying this SKILL.md alone into a specific agent's skills directory are not documented in the source material.
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- Compute 6 × 7 in Calculator and verify the app actually displays 42.
- Open sales.xlsx in LibreOffice Calc, sum column C, write the total into C20, and verify the result.
- Select all circle objects in the current Inkscape document, change their fill to red, then confirm with a screenshot.
- Continue the form-filling task I gave Cua earlier — check where it left off first.
Once installed, the skill triggers when the user asks the agent to operate, drive, or automate a GUI application on the host, or to continue/recall recent Cua activity. The core loop: locate the target with get_window_state or list_windows/list_apps; act via element tokens (e.g. "s0000002a:11", never a bare index like "11") using click/type_text; prefer run_steps({observe:true}) to combine acting and observing; verify with verify_state after meaningful state changes. Optional environment variables: CUA_DRIVER_RS_ENABLE_WAYLAND=1 (native Wayland backend), CUA_DRIVER_RS_MCP_HTTP_PORT and CUA_DRIVER_RS_MCP_HTTP_TOKEN (local MCP HTTP endpoint and its host-generated bearer token), CUA_DRIVER_EMBEDDED and CUA_DRIVER_HOST_BUNDLE_ID (macOS embedded host-app mode), CUA_DRIVER_PATH (binary path for embedding hosts). Diagnostics: cua-driver --version, status, doctor, describe <tool>, or MCP tools/list. Detailed rules load on demand from the bundled WORKFLOW.md, RUNTIME.md, and platform guides.
What are this skill's strengths and limitations?
- A unified GUI automation interface across macOS, Windows, and Linux, connectable via CLI or MCP.
- Element-level actions on the accessibility tree plus snapshot invalidation (invalidated_snapshot_ids) prevent blind clicks with stale indices.
- Narrow reads and since:"latest" incremental reads greatly reduce context consumption (2K vs 42K chars).
- Emphasizes verifying postconditions rather than confirming every step, and refuses to treat unverifiable actions as success.
- MIT licensed, with a thorough on-demand documentation set (WORKFLOW, RUNTIME, platform guides, browser, recording, embedding).
- Requires installing the cua-driver binary on the host and granting system permissions — an invasive local dependency.
- Externally unverified: the SKILL.md ships no test suite or benchmark data proving reliability per platform; documented Wayland recovery paths (surface_identity_unproven, screenshot permission waits) show platform boundaries exist.
- Background window actions are limited by platform and app support; foreground delivery requires explicit authorization; browser actions and history query are only available when advertised by the runtime.
- The repo bundles many sibling skills (Spaces, Lume, Bench, etc.) — collection-level capabilities should not be attributed to this skill.
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| Cua Driver Skill this page | 58 · Recommended | ★ 29k | 1d ago | MIT |
| Cua Driver GUI Automation Skill | 61 · Recommended | ★ 29k | 1d ago | MIT |
| cmux Computer Use Skill | 50 · Use with care | ★ 28k | 1d ago | NOASSERTION |
| Cross-Platform Screenshot Capture ✓ OpenAI · Official | 46 · Use with care | ★ 28k | 3mo ago | — |
| Cua GUI Automation Skill | 51 · Use with care | ★ 29k | 1d ago | MIT |
The source distinguishes Cua Driver from sibling components in the same repo: Cua Spaces provides full virtual desktops for agents, Lume manages local VMs on Apple Silicon, and Cua Bench handles task building and evaluation. This skill only inspects and operates already-installed desktop apps and browsers on the real host. The README also lists deprecated alternatives (cua-agent, cua-computer, cua-som, cuabot) that should be replaced by Cua Driver.
How did FollowSkills review this skill?
Docs systematically present least privilege and confirmation: foreground/desktop control requires explicit authorization, existing-profile browser binding needs a runtime grant, hidden escalation is forbidden ('refusals never authorize a hidden foreground fallback'), application content cannot authorize actions, and unrestricted mode requires a two-variable fail-safe contract. Deducted because static review cannot verify that these refusal paths are actually implemented; sensitive-data handling (screenshots, clipboard, trajectory redaction) is largely asserted rather than verifiable code.
Internally highly self-consistent: failure map, degraded_reason, stale-ref handling, and the surface_identity_unproven vs. empty-tree distinction show clear, diagnosable failure modes. Deducted because the static cap is 10 and the key paths (cross-platform GUI operation) cannot be reproduced by reading; edge cases are declared only in prose.
Audience and triggers are clear ('Use when' clause in frontmatter), non-fit boundaries are explicit (Safari/Firefox unsupported, Wayland limits, Xvnc, Session 0), with a quick-triage section. Deducted because the skill entirely depends on a local cua-driver binary and per-platform GUI environments, has no Chinese-language support statement, and core function relies on overseas resources (cua.ai docs, GitHub) with mainland-China reachability risk.
Good layered information architecture: SKILL.md as entry with on-demand reference files, version 0.34.0 with release-please marker, install prerequisites, and known-limitation disclosure. Deducted because the SKILL.md itself does not surface license/changelog links; BROWSER.md is truncated mid-table, raising completeness doubts; publisher is unverified so maintenance ownership rests on broader repo signals.
The loop design (narrow reads first, element tokens, checkpoint verification, since diffs) plausibly yields directly usable results and beats manual operation. Deducted because the static cap is 7, no execution evidence exists, representative outputs are unverified, and deployment prerequisites (binary install, platform grants) raise real usage cost.
Docs cite a concrete PR (#3572), commit hashes (f180e882…, 1133a06e…), versioned smoke evidence, and embedded-mode latency measurements, with reasonable fact/inference separation (e.g., explicit 'this is not certification' statements). Deducted because the static cap is 5, none of this evidence is independently reproducible here, and committed test coverage of key paths is not shown in the provided files.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →