What does this skill do, and when should you use it?
The cua-driver skill hands an agent control of a real desktop application. It connects through the cua-driver CLI (default) or an MCP server, snapshots the app's accessibility tree, and acts through snapshot-bound element tokens, native menu paths, exact window geometry, or pixel coordinates. The workflow is deliberately observe-act-verify: it reads fresh state before input, performs one action, and confirms the user's postcondition from new state. It lives in the trycua/cua monorepo alongside VM, benchmark, and SDK siblings, but this skill itself covers only driving GUI apps on the host.
- Finds and opens the target app via list_apps / list_windows / launch_app
- Captures a window's accessibility tree with get_window_state
- Performs input (click, type_text) using element_token values returned by the snapshot
- Falls back to fresh-screenshot pixel coordinates when semantics cannot reach a control
- Verifies the postcondition with verify_state or a fresh snapshot read
- Supports desktop-level input, bound browser page actions, and explicitly requested recordings
- Agent developers who need an agent to complete a task in a real app (Calculator, LibreOffice, Inkscape) and prove the result on screen
- Operators driving desktop apps in the background without stealing pointer or focus, where the app and platform allow it
- Users resuming an interrupted Cua task via history_status/history_query recall
- Testers who want accessibility-tree-driven GUI automation instead of brittle coordinate scripting
- Headless servers or environments without a GUI — the skill targets real application windows
- Scenarios hoping to bypass permission prompts or alter browser security settings covertly — the skill's rules explicitly forbid this
- Targets outside the authorized desktop; unauthorized foreground/desktop control is refused, not escalated
How do you install this skill?
- This skill genuinely controls the user's desktop and applications; foreground/desktop input and existing-browser-profile binding are high-impact. Enable only with explicit user authorization; this was a static review with no execution.
- Publisher identity is unverified (unknown, not suspicious). Reference docs (WORKFLOW/RUNTIME, etc.) were not in the evidence set, so actual behavior may diverge from documentation.
- Docs are English-only with no Chinese support statement; mainland-China network reachability is unverified. Rollback is incomplete: a crash can leave browser remote-debugging settings changed, requiring manual recovery.
- Several capabilities (Hyprland plugin, KDE foreground activation, legacy page tool) are experimental or sharply bounded; do not assume full coverage.
- Shell / CLI
- Network access
- MCP Server
cua-driver binarymacOS/Windows/Linux host with GUIscreen/accessibility permissions
No per-client manual steps are documented beyond the installers; the skill ships with the Cua tooling. Install the cua-driver binary:
macOS / Linux terminal
/bin/bash -c "$(curl -fsSL https://cua.ai/driver/install.sh)"Windows PowerShell
irm https://cua.ai/driver/install.ps1 | iexBundled with the cua CLI into your agents
curl -fsSL https://cua.ai/install.sh | sh
sh -s -- --select cua-driverThe installer offers the cua-driver skill and MCP server into Claude Code, Codex, Cursor, and other agents; other clients are not documented.
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- Open Calculator, compute 6 × 7, and verify the app displays 42
- Select cell A1 in LibreOffice Calc and fill in today's date, then confirm the cell contents
- Continue my interrupted Cua task — check the history first
- Select all objects in Inkscape in the background without taking my mouse focus
The skill triggers when the user asks the agent to operate, drive, or automate a GUI task in a real application on the host, or to continue recent Cua activity. The flow: run status/doctor preflight, locate the app with list_apps, take get_window_state, act once with the returned element_token, then verify with verify_state or a fresh snapshot. Optional settings include CUA_DRIVER_RS_MCP_HTTP_PORT plus CUA_DRIVER_RS_MCP_HTTP_TOKEN for a local MCP HTTP endpoint, and CUA_DRIVER_RS_ENABLE_WAYLAND=1 for the native Wayland backend on Linux. Platform guides (MACOS/WINDOWS/LINUX) and workflow references load on demand.
What are this skill's strengths and limitations?
- Snapshot-bound element tokens avoid brittle coordinate guessing and keep actions precise
- Observe-act-verify loop explicitly separates a successful exit code from task success
- Concrete failure map and permission boundaries, e.g. unverifiable effects are never treated as success
- Requires the cua-driver binary plus OS accessibility/screen permissions, a nontrivial setup
- Background window actions depend on app and platform support; Wayland has extra capture-recovery failure paths
- Shared-desktop focus, keyboard, and snapshot caches are not isolated across sessions, so concurrent agents must coordinate
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| Cua Driver GUI Automation Skill this page | 61 · Recommended | ★ 29k | 1d ago | MIT |
| Cua Driver Skill | 58 · Recommended | ★ 29k | 1d ago | MIT |
| Cross-Platform Screenshot Capture ✓ OpenAI · Official | 46 · Use with care | ★ 28k | 3mo ago | — |
| cmux Computer Use Skill | 50 · Use with care | ★ 28k | 1d ago | NOASSERTION |
| Cua Spaces (cua-spaces skill) | 46 · Use with care | ★ 29k | 1d ago | MIT |
The README positions Cua Driver as the replacement for deprecated packages (cua-computer, cua-som, cua-agent): use Cua Driver for the machine you are on and the cua SDK for sandboxes and remote machines.
How did FollowSkills review this skill?
SKILL.md states least privilege and authorization boundaries: foreground/desktop input requires explicit authorization, no hidden fallbacks, no security-setting changes, refusals instead of silent execution; browser accepts only http/https/about, CDP method policy is restricted, existing-profile binding needs an explicit expiring grant discarded at runtime, sensitive-data handling (screenshot opt-in, redacted file paths) is disclosed. Deducted: controlling the user's desktop is inherently high-impact; rollback is only partial (crash can leave browser settings changed, manual recovery documented), publisher identity unverified, and authorization implementation cannot be confirmed statically.
Internally highly self-consistent: failure maps cover stale tokens, empty AX trees, Wayland surface_identity_unproven, popup keyboard grabs, Session 0, missing D-Bus session bus, with named error codes and diagnosable refusals. Deducted: static review only — no committed test suite or CI evidence covering key paths was shown in the files; some capabilities are labeled experimental (Hyprland plugin, KDE foreground activation), so beyond-happy-path behavior rests on self-description.
Trigger description is precise (user asks to operate a GUI app, or continue/resume prior Cua work); platform boundaries and non-fit ranges (Safari/Firefox, Xvnc, containers without uinput, Session 0) are declared; doctor handles environment detection. Deducted: docs are English-only with no Chinese support statement; mainland-China reachability depends mainly on the cua.ai docs site and is unverified.
Well-layered docs (SKILL → WORKFLOW/RUNTIME/platform/browser/recording/embedding) loaded on demand, version 0.31.0 with release-please marker, MIT license, complete env-var tables, extensive known-limitation disclosure, embedding guide kept in sync with a repo example. Deducted: changelog and maintenance ownership not evidenced inside the assessed files; reference files (WORKFLOW/RUNTIME) and table tails were not in the evidence set.
Claimed outputs (verifiable postconditions, typed browser actions, state verification) have clear marginal value over manual osascript/xdotool: background automation without focus stealing. Deducted: not executed in this review; direct usability rests on self-reported smoke evidence (KDE Wayland timing, Hyprland PR validation), representative outputs not independently reproduced, capped at 7.
Better than marketing-only: dated PR validation record (#3572), source hashes (f180e882…, 1133a06e…), in-repo example code paths, and measured macOS latency figures. Deducted: all author-supplied; no third-party execution evidence or committed tests appear in the assessed files, cross-source corroboration is limited, and the 5 cap is not reached.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →