Automation & Ops

Cua Driver Skill

Let agents operate real desktop applications: drive native GUIs on macOS, Windows, and Linux via accessibility trees and element tokens, then verify from fresh state.

58/ 100
Recommended

Generally reliable with disclosed limitations; trial as directed and keep a rollback path.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 29k
Last updated
1d ago
License
MIT
desktop-automationcomputer-usegui-automationmcp-server
+4accessibility-treemacoswindowslinux

What does this skill do, and when should you use it?

cua-driver is an Agent Skill inside the trycua/cua repository that wraps the Cua Driver's cua-driver CLI and MCP server. It teaches an agent the correct way to operate native GUI applications on the host: observe narrowly first, act through snapshot-bound element tokens, and verify the postcondition at checkpoints. The skill is part of the MIT-licensed Cua Driver codebase and requires the cua-driver binary to be installed on the host beforehand. It is not a VM or sandbox tool — it is an automation interface to a real, authorized desktop.

  • Reads a single window narrowly with get_window_state plus a query (~2K chars instead of a ~42K full snapshot), bounding large trees with max_elements/max_depth.
  • Batches actions via run_steps with observe:true, stopping at the first failure and returning only what changed since the last read.
  • Writes loops, branches, and retries as one sandboxed run_script (JavaScript over the cua API, with time and call limits).
  • Verifies postconditions at checkpoints with verify_state or a targeted get_window_state instead of after every action.
  • Supports launching apps, opening local files, desktop operation, binding browser page content, and explicitly requested session recording.
  • Queries recent Cua activity in a bounded way via history_status/history_query when the user asks to continue or recall prior work.
Good fit
  • Developers using coding agents like Claude Code, Codex, or Cursor who want the agent to complete GUI tasks in real host applications (Calculator, LibreOffice Calc, Inkscape) and verify the displayed result.
  • QA or ops staff automating repetitive cross-platform desktop-app actions (menu clicks, form filling, file exports) across macOS, Windows, and Linux.
  • Scenarios where the outcome lives in an app's UI state and the user demands GUI-only interaction, excluding DOM/CDP, clipboard APIs, or shell mutations.
  • Teams wanting one CLI/MCP interface to drive desktop apps on different operating systems without writing per-platform scripts.
  • Users who need an explicitly requested automation run recorded, with artifacts verified, for review or replay reference.
Not a fit
  • Users who want to generate training data or run benchmarks inside VMs or cloud sandboxes — that is the job of sibling components Lume and Cua Bench in the same repo, not this skill.
  • Automations relying on implicit privilege escalation: the skill forbids unauthorized foreground/desktop control, altering browser security settings, or handling permission prompts on the user's behalf.
  • Environments unwilling to install a local binary or grant accessibility/screen-recording permissions — without the cua-driver binary the skill cannot run.

How do you install this skill?

Before you use it
  • This is a static source review only; nothing was executed and all scores rest on documented claims, at low confidence.
  • Core function requires a locally installed cua-driver binary plus macOS Accessibility/Screen Recording system grants; deployment cost is significant.
  • The skill is coupled to overseas resources (cua.ai docs, GitHub); mainland-China reachability is unverified and there is no Chinese-language support statement.
  • BROWSER.md is truncated mid-table in the provided files; verify the actual pack completeness.
  • Publisher is unverified by the FollowSkills registry and treated as unknown; the Hyprland candidate in LINUX.md is explicitly experimental and ABI-dependent.
  • Existing-profile browser binding involves CDP remote-debugging authorization; although the docs declare strict guardrails, audit the implementation before use.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
  • MCP Server
Install first
  • cua-driver binary (installed via cua.ai install script)
  • local MCP HTTP bearer token (CUA_DRIVER_RS_MCP_HTTP_TOKEN) if using the MCP endpoint

The Cua Driver binary (which the skill depends on via the cua-driver command) must be installed first. The source documents these host install routes:

macOS / Linux hosts

/bin/bash -c "$(curl -fsSL https://cua.ai/driver/install.sh)"

Windows hosts

irm https://cua.ai/driver/install.ps1 | iex

Via the unified cua installer (driver only)

curl -fsSL https://cua.ai/install.sh | sh -- --only cua-driver

The skill file itself lives at libs/cua-driver/rust/Skills/cua-driver/SKILL.md in the repo. The README notes the installer offers to install cua skills and the cua MCP server into coding agents (Claude Code, Codex, Cursor, and others), but the exact steps for copying this SKILL.md alone into a specific agent's skills directory are not documented in the source material.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • Compute 6 × 7 in Calculator and verify the app actually displays 42.
  • Open sales.xlsx in LibreOffice Calc, sum column C, write the total into C20, and verify the result.
  • Select all circle objects in the current Inkscape document, change their fill to red, then confirm with a screenshot.
  • Continue the form-filling task I gave Cua earlier — check where it left off first.

Once installed, the skill triggers when the user asks the agent to operate, drive, or automate a GUI application on the host, or to continue/recall recent Cua activity. The core loop: locate the target with get_window_state or list_windows/list_apps; act via element tokens (e.g. "s0000002a:11", never a bare index like "11") using click/type_text; prefer run_steps({observe:true}) to combine acting and observing; verify with verify_state after meaningful state changes. Optional environment variables: CUA_DRIVER_RS_ENABLE_WAYLAND=1 (native Wayland backend), CUA_DRIVER_RS_MCP_HTTP_PORT and CUA_DRIVER_RS_MCP_HTTP_TOKEN (local MCP HTTP endpoint and its host-generated bearer token), CUA_DRIVER_EMBEDDED and CUA_DRIVER_HOST_BUNDLE_ID (macOS embedded host-app mode), CUA_DRIVER_PATH (binary path for embedding hosts). Diagnostics: cua-driver --version, status, doctor, describe <tool>, or MCP tools/list. Detailed rules load on demand from the bundled WORKFLOW.md, RUNTIME.md, and platform guides.

What are this skill's strengths and limitations?

Pros
  • A unified GUI automation interface across macOS, Windows, and Linux, connectable via CLI or MCP.
  • Element-level actions on the accessibility tree plus snapshot invalidation (invalidated_snapshot_ids) prevent blind clicks with stale indices.
  • Narrow reads and since:"latest" incremental reads greatly reduce context consumption (2K vs 42K chars).
  • Emphasizes verifying postconditions rather than confirming every step, and refuses to treat unverifiable actions as success.
  • MIT licensed, with a thorough on-demand documentation set (WORKFLOW, RUNTIME, platform guides, browser, recording, embedding).
Limitations
  • Requires installing the cua-driver binary on the host and granting system permissions — an invasive local dependency.
  • Externally unverified: the SKILL.md ships no test suite or benchmark data proving reliability per platform; documented Wayland recovery paths (surface_identity_unproven, screenshot permission waits) show platform boundaries exist.
  • Background window actions are limited by platform and app support; foreground delivery requires explicit authorization; browser actions and history query are only available when advertised by the runtime.
  • The repo bundles many sibling skills (Spaces, Lume, Bench, etc.) — collection-level capabilities should not be attributed to this skill.

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
Cua Driver Skill this page 58 · Recommended ★ 29k 1d ago MIT
Cua Driver GUI Automation Skill 61 · Recommended ★ 29k 1d ago MIT
cmux Computer Use Skill 50 · Use with care ★ 28k 1d ago NOASSERTION
Cross-Platform Screenshot Capture ✓ OpenAI · Official 46 · Use with care ★ 28k 3mo ago —
Cua GUI Automation Skill 51 · Use with care ★ 29k 1d ago MIT

The source distinguishes Cua Driver from sibling components in the same repo: Cua Spaces provides full virtual desktops for agents, Lume manages local VMs on Apple Silicon, and Cua Bench handles task building and evaluation. This skill only inspects and operates already-installed desktop apps and browsers on the real host. The README also lists deprecated alternatives (cua-agent, cua-computer, cua-som, cuabot) that should be replaced by Cua Driver.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Recommended
58/ 100 5-point scale 2.9 / 5
1Trust18 / 25 · 3.6/5

Docs systematically present least privilege and confirmation: foreground/desktop control requires explicit authorization, existing-profile browser binding needs a runtime grant, hidden escalation is forbidden ('refusals never authorize a hidden foreground fallback'), application content cannot authorize actions, and unrestricted mode requires a two-variable fail-safe contract. Deducted because static review cannot verify that these refusal paths are actually implemented; sensitive-data handling (screenshots, clipboard, trajectory redaction) is largely asserted rather than verifiable code.

2Reliability9 / 20 · 2.3/5

Internally highly self-consistent: failure map, degraded_reason, stale-ref handling, and the surface_identity_unproven vs. empty-tree distinction show clear, diagnosable failure modes. Deducted because the static cap is 10 and the key paths (cross-platform GUI operation) cannot be reproduced by reading; edge cases are declared only in prose.

3Adaptability10 / 15 · 3.3/5

Audience and triggers are clear ('Use when' clause in frontmatter), non-fit boundaries are explicit (Safari/Firefox unsupported, Wayland limits, Xvnc, Session 0), with a quick-triage section. Deducted because the skill entirely depends on a local cua-driver binary and per-platform GUI environments, has no Chinese-language support statement, and core function relies on overseas resources (cua.ai docs, GitHub) with mainland-China reachability risk.

4Convention11 / 15 · 3.7/5

Good layered information architecture: SKILL.md as entry with on-demand reference files, version 0.34.0 with release-please marker, install prerequisites, and known-limitation disclosure. Deducted because the SKILL.md itself does not surface license/changelog links; BROWSER.md is truncated mid-table, raising completeness doubts; publisher is unverified so maintenance ownership rests on broader repo signals.

5Effectiveness6 / 15 · 2.0/5

The loop design (narrow reads first, element tokens, checkpoint verification, since diffs) plausibly yields directly usable results and beats manual operation. Deducted because the static cap is 7, no execution evidence exists, representative outputs are unverified, and deployment prerequisites (binary install, platform grants) raise real usage cost.

6Verifiability4 / 10 · 2.0/5

Docs cite a concrete PR (#3572), commit hashes (f180e882…, 1133a06e…), versioned smoke evidence, and embedded-mode latency measurements, with reasonable fact/inference separation (e.g., explicit 'this is not certification' statements). Deducted because the static cap is 5, none of this evidence is independently reproducible here, and committed test coverage of key paths is not shown in the provided files.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision e32127764436 Review evidence[1][2][3][4][5]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

What permissions does it need? Will the agent mess with my computer?
The skill's rules are explicit: foreground delivery and desktop input require authorization for visible control; background window actions must be non-interfering. User/system permission prompts belong to the user or trusted host, and the skill forbids altering browser profiles or security settings as hidden setup. Only one controller should operate a shared desktop to avoid focus and input conflicts.
What happens when an action fails?
The skill includes a failure map: missing binary or daemon mismatch goes to RUNTIME.md preflight; stale tokens or ambiguous windows are fixed by refreshing list_windows/get_window_state; large or sparse trees use incremental since-diffs; Wayland capture issues follow LINUX.md capture recovery; refused browser bindings follow BROWSER.md recovery rules. When text did not visibly change, reobserve before retrying.
Is it free? Any licensing caveats?
Cua Driver and this skill are part of the MIT-licensed portion of the repo and free to use. However, the optional cua-perception extension is not MIT — it combines an AGPL-3.0 OmniParser model artifact with Apache-2.0 PP-OCR artifacts, and redistributing it or offering it over a network can trigger AGPL-3.0 source obligations.
How does this differ from Cua Spaces or Lume?
This skill only drives real GUI apps already installed on the host. Cua Spaces provides full virtual desktops for agents (macOS/Linux VMs, FSL-1.1-MIT licensed), Lume manages local VMs via Virtualization.Framework, and Cua Bench builds and evaluates computer-use tasks. Those are separate components in the repo and out of scope for this skill.

More skills from this repository

All from trycua/cua

Automation & Ops

Cua Driver GUI Automation Skill

Lets an AI agent operate real native app windows on macOS, Windows, and Linux: observe state, act precisely, and verify the outcome.

★ 29k FS 61 Recommended 1d ago
Automation & Ops

Cua GUI Automation Skill

Give AI agents eyes and hands on a real computer: click buttons, fill forms, and run end-to-end visual QA on any application's GUI.

★ 29k FS 51 Use with care 1d ago
Automation & Ops

Cua Spaces (cua-spaces skill)

Through the cua MCP server, lets your agent work inside a watchable remote or local computer — running commands, moving files, spawning coding agents and sharing the host network.

★ 29k FS 46 Use with care 1d ago
Dev & Engineering

jev-use — A Bounded Computer-Use Loop over Cua Driver

Build a tightly bounded computer-use loop on top of Cua Driver: the driver only observes and acts, TypeSafe Jev picks from application-owned candidate IDs, and the caller verifies every action.

★ 29k FS 58 Recommended 1d ago
Automation & Ops

Cua Volume Skill

Give AI agents one versioned, user-level shared volume that persists files across Spaces, syncs to your Mac in seconds, and shares safely with other agents.

★ 29k FS 52 Use with care 1d ago
Dev & Engineering

Cua Sandboxes

Spin up disposable Linux or macOS machines locally or in the Cua cloud, so agents can run code, test apps, drive a desktop GUI, and browse the web without ever touching your own computer.

★ 29k FS 52 Use with care 1d ago
Dev & Engineering

Poll GitHub Work

Polls and ranks open GitHub issues, RFCs, and PRs so maintainers know what to work on next — or starts one explicitly selected item.

★ 29k FS 52 Use with care 1d ago

Related skills