What does this skill do, and when should you use it?
gui-automation is an Agent Skill inside the trycua/cua repository (MIT licensed) that gives agents "eyes and hands on a real computer" through the `cua` CLI: screenshots, clicks, typing, drags, and window management. It targets visual interaction that can't be done via shell or API — button testing, form filling, layout verification, and web fuzzing. The skill connects to sandboxes, relay machines, spacesd URLs, or your local machine, and auto-records every action into a trajectory you can replay. With an optional ANTHROPIC_API_KEY, `cua do snapshot` returns an AI-annotated screen with element coordinates.
- Screenshots the target screen (
cua do screenshot), optionally zooming into a single window first (cua do zoom) for higher-precision clicks - Performs mouse and keyboard actions: click, double-click, type, key presses, hotkeys, scroll, drag, cursor move
- Calls
cua do snapshot(with ANTHROPIC_API_KEY) to get an AI-annotated screen with element coordinates for precise targeting - Connects to multiple targets: sandboxes (
cua sb ls), relay Spaces, spacesd URLs, or the local host after one-time consent - Auto-records every action to ~/.cua/trajectories/ and opens a replay with
cua trajectory view - Lists, focuses, and manages application windows (
cua do window ls/focus)
- A QA engineer end-to-end tests a desktop or web app with no API, verifying buttons and UI behavior through a screenshot-click-screenshot loop
- A tester automates full user flows — login, file upload, form submission — inside cloud VMs, Docker containers, or sandboxes
- A security team fuzzes a web form with XSS or SQL injection payloads, screenshotting afterward to check for errors, crashes, or unexpected behavior
- An automation scenario that must fill multiple form fields, tabbing between them
- A developer needing precise clicks on small, dense UI elements on the host machine, zooming into the window first
- Tasks doable via shell, API, or headless tooling — SKILL.md explicitly scopes this skill to visual interaction that can't be done via shell or API
- Users unwilling to install the cua CLI and its runtime (sandboxes/Spaces/driver) — the skill is entirely dependent on the `cua` command
- Teams wanting AI-annotated screens without providing an ANTHROPIC_API_KEY — without a key you must read raw screenshots yourself
How do you install this skill?
- This is a static source review (low confidence); no commands were executed and example usability is unverified.
- The install command curl -fsSL https://cua.ai/install.sh | sh executes a remote script; review its contents first. Reachability of cua.ai from mainland China is unverified.
- GUI automation is broad-permission: examples include typing a password (SecureP@ss123); use in controlled environments and note that trajectories are recorded by default under ~/.cua/trajectories/.
- The AI-annotated snapshot feature depends on ANTHROPIC_API_KEY and sends screen contents to a third-party API; be aware of sensitive-screen disclosure risk.
- Before using `switch host`, understand the scope of the consent mechanism; security of sandbox/relay targets depends on that infrastructure.
- The publisher is not verified by the FollowSkills registry and is treated as unknown; the skill has no version number or changelog, so track upstream updates yourself.
- Shell / CLI
- Network access
- Local filesystem
cua CLI (curl -fsSL https://cua.ai/install.sh | sh)optional ANTHROPIC_API_KEY for AI-annotated snapshots
Install the cua collection containing the skills (installs the cua CLI and, by default on macOS 26+, the Spaces app):
curl -fsSL https://cua.ai/install.sh | shWindows (PowerShell):
irm https://cua.ai/install.ps1 | iexThe installer runs cua auth login and offers to install cua skills and the MCP server into AI coding agents such as Claude Code, Codex, and Cursor:
cua auth loginThe skill file lives at libs/cua/skills/gui-automation/SKILL.md in the repo; SKILL.md itself documents no separate copy command for the skill alone.
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- Open the signup page at http://127.0.0.1:3000 in the browser, fill in the username and email, submit, and screenshot to verify the success message
- In LibreOffice, click File > Export and export the current document as PDF to /home/user/report.pdf, screenshotting each step to confirm
- Fuzz this login form: inject <script>alert(1)</script> and a SQL injection payload in turn, then screenshot to check for errors or crashes
- Switch to my dev sandbox, open the file upload dialog, pick /home/user/report.pdf, upload it, and verify the filename appears on the page
The skill triggers when the agent needs to visually operate a GUI. The core flow is Look → Act → Verify: screenshot, act, screenshot again, repeat until done, then share the trajectory replay. First connect to a target (sandbox, relay Space, spacesd URL, or host):
cua do switch my-sandbox
cua do screenshot
cua do click 450 280
cua do screenshot
cua trajectory viewKey options: re-screenshot after every UI change (coordinates go stale); use cua do zoom "App Name" before clicking small targets to make coordinates window-relative, then cua do unzoom; use cua do --no-record click x y to skip recording for a single action; the local host requires one-time cua do-host-consent.
What are this skill's strengths and limitations?
- Environment-agnostic: cloud VMs, Docker containers, sandboxes, and the local machine all work as targets
- Every action is auto-recorded with a replayable trajectory for auditing and reproduction
- Window zoom provides window-relative coordinates, solving precision clicks on small elements
- Thorough command reference covering click, type, drag, window management, and scrolling
- Coordinate-based clicking: coordinates go stale whenever the screen changes, demanding strict re-screenshot discipline
- AI-annotated snapshots require a separate ANTHROPIC_API_KEY and its associated API cost
- SKILL.md offers no automated test suite or success-rate evidence; real reliability must be validated yourself
- The Cua Spaces app and cua-spacesd are FSL-1.1-MIT source-available (not MIT); offering hosted Spaces requires checking COMMERCIAL.md terms
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| Cua GUI Automation Skill this page | 51 · Use with care | ★ 29k | 1d ago | MIT |
| Cua Driver Skill | 58 · Recommended | ★ 29k | 1d ago | MIT |
| cmux Computer Use Skill | 50 · Use with care | ★ 28k | today | NOASSERTION |
| Cua Driver GUI Automation Skill | 61 · Recommended | ★ 29k | 1d ago | MIT |
| Cross-Platform Screenshot Capture ✓ OpenAI · Official | 46 · Use with care | ★ 28k | 3mo ago | — |
The README frames the repo as a "Computer-Use 2.0" approach: an agent moves between code, APIs, and graphical interfaces within the same task; compared with single-layer computer-use tools, Cua also ships a sandbox SDK, Lume local VMs, and the Cua Bench evaluation environment. SKILL.md names no specific competitors.
How did FollowSkills review this skill?
Positives: explicit one-time consent (cua do-host-consent) before controlling the local machine, sandbox/relay as default remote targets, automatic trajectory recording of every action with an opt-out flag, and reasonably clear data-flow disclosure. Deductions: install via curl | sh remote script (weak supply-chain transparency), GUI automation is inherently broad-permission (click, type, shell, clipboard), no risk guidance for sensitive input (the example types a password), and rollback relies on replayable trajectories rather than true undo. Scored 14.
SKILL.md and command-reference.md are mutually consistent; command syntax matches; the Look→Act→Verify workflow is a clear reproduction path; stale-coordinate edge cases are flagged. Deductions: static review cannot execute anything, no error-feedback examples shown (how failures are diagnosed), no test evidence covering this skill's key paths, and availability of sandbox/relay prerequisites is not covered within the skill. Capped at 9.
Trigger conditions (visual interaction not achievable via shell/API) and target environments (cloud VMs, Docker, local, sandboxes) are clearly declared; zoom guidance covers precision boundaries. Deductions: no explicit non-fit range (e.g., headless environments), mainland-China reachability of the cua.ai installer and ANTHROPIC_API_KEY-dependent snapshot is not disclosed, and Chinese-language support is unaddressed. Scored 8.
Good information architecture: concise SKILL.md with details pushed to references/command-reference.md (progressive disclosure); reference claims derivation from --help; MIT license and repo-level maintenance responsibility (SECURITY.md, contributing) are clear. Deductions: no skill-level version number or changelog, no version governance ensuring the reference file stays in sync, and known limitations (platform differences, permission requirements) incompletely disclosed. Scored 10.
Targets real needs (GUI testing, form filling, E2E QA); examples are directly usable and consistently formatted; trajectory replay gives users auditable output. Deductions: static review cannot verify actual example outcomes, marginal value depends on installing the cua CLI ecosystem, and no evidence of representative completed tasks or comparison with alternatives (Playwright, native a11y APIs). Scored 6.
The repository contains real CI workflows and test suites, but no evidence in the reviewed material shows tests covering this skill's key paths (GUI click/type) that are independently reproducible; the reference's claimed --help provenance is author-asserted. Facts and inference are reasonably separated, but third-party corroboration and executed reproduction are insufficient. Scored 4.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →