What does this skill do, and when should you use it?
cua-sandboxes is one of the skills bundled in the trycua/cua monorepo, teaching agents to create, use, and clean up sandboxes — disposable containers or VMs — via the `cua` CLI and the cua SDK. It covers gVisor containers, QEMU Linux VMs, and Lume macOS VMs on Apple silicon, plus any OCI image and direct connections to machines running cua-spacesd. The skill documents three entry paths: terminal commands, code (Python with equivalent TypeScript, Swift, and Kotlin bindings), and the cua MCP server, including a full in-sandbox web-browsing workflow with Chromium. Local sandboxes need no account; cloud sandboxes are metered and must be deleted when done.
- Creates sandboxes with
cua sb createas local gVisor containers, QEMU Linux VMs, or Lume macOS VMs (Apple silicon), in the Cua cloud, from any OCI image, or attached to an existing machine running cua-spacesd - Runs commands, interactive shells, screenshots, and browser-based desktop viewing via
cua sb exec/shell/screenshot/vnc - Exposes the same sandbox lifecycle from code: Python's
cua.embedded()plus matching objects in TypeScript, Swift, and Kotlin bindings - Drives in-sandbox web browsing through cua MCP tools (images_list, sandbox_create, browser_navigate, browser_click, browser_type) against a Chromium-shipping image
- Cleans up with
cua sb rm,cua sb suspend, andcua fleet pools gc, and pauses/resumes sandboxes
- A developer who needs an isolated machine to run untrusted code or test an app without polluting their own machine
- An automation agent that must click, fill forms, or take screenshots in a desktop GUI or on a website, keeping the user's real desktop untouched
- An engineer testing a website or a local web app by running it in the sandbox and opening http://127.0.0.1:<port>
- A tester who needs a macOS VM environment on Apple silicon via Lume
- A pipeline that reproduces an issue or generates data in a disposable environment and tears it down afterward
- Users unwilling to install the cua toolchain — every operation in this skill depends on the cua CLI/SDK
- Pure API model integrations with no shell or runtime to execute commands in
- Workflows needing a Windows sandbox image — the skill documents only Linux and macOS images and runtimes (gVisor, QEMU, Lume), even though the broader repo supports Windows
How do you install this skill?
- Cloud sandboxes depend on cua.ai authentication, relay, and ghcr.io image sources; mainland-China reachability is unverified and core functionality may be unavailable — validate in local mode (gVisor/QEMU) first.
- Token handling for `--on direct:... --token TOKEN` is undocumented; do not pass long-lived credentials in shared environments.
- Cloud sandboxes are metered; always run `cua sb rm` after the task as instructed to avoid ongoing charges.
- The skill assumes images shipping cua-spacesd; with plain OCI images, exec/screenshot/GUI silently degrade to lifecycle-only management.
- This is a static, non-executed review; all command behavior, image availability, and isolation strength claims are unverified.
- Publisher is not verified by the FollowSkills registry and is treated as unknown; verify repository and binary provenance independently before use.
- Shell / CLI
- Network access
- Local filesystem
- MCP Server
cua CLI (Rust)cua SDK (Python/TypeScript/Swift/Kotlin/Rust)cua MCP server for browser workflowcua-spacesd-capable images
Install the full cua toolchain first (this skill ships inside the libs/cua/skills collection):
macOS / Linux (terminal installer)
curl -fsSL https://cua.ai/install.sh | shWindows (PowerShell)
irm https://cua.ai/install.ps1 | iexThe installer walks you through cua auth login and can install cua skills and the cua MCP server into AI coding agents such as Claude Code, Codex, and Cursor.
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- Run `uname -a` in an isolated Linux sandbox and tell me the output
- Create a macOS sandbox and take a screenshot of its desktop
- Open example.com inside a sandbox, read the page, and fill in the search box
- Test my app on port 8080: start it in a sandbox, visit http://127.0.0.1:8080, and delete the sandbox when done
The skill triggers from its SKILL.md description: any task needing an isolated machine to run code, test an app, drive a GUI, or browse/test a website. Core flow:
cua --version # confirm the Rust CLI
cua auth status # cloud requires login
cua sb create linux --name dev # local gVisor container
cua sb exec dev -- uname -a # run a command
cua sb screenshot dev # save a screenshot
cua sb vnc dev # open desktop in a browser
cua sb rm dev --force # delete when doneOptions worth knowing: --kind vm for a QEMU Linux VM, cua sb create macos for a Lume macOS VM (Apple silicon), --on cloud for the metered Cua cloud, and -- on any command for machine-readable output. GUI work goes through cua do switch <name> then cua do click/screenshot. Web browsing uses cua MCP tools: images_list, sandbox_create with browser: true, then browser_navigate/browser_click/browser_type and get_browser_state for semantic snapshots. Local runtimes auto-install on first use (cua runtime setup does it explicitly), and logged-in-site teleport requires explicit user consent.
What are this skill's strengths and limitations?
- Local sandboxes work with no account, keeping everyday use friction-free
- One runtime and sandbox list shared by CLI and Python/TypeScript/Swift/Kotlin SDKs, so scripts and agents interoperate
- Supports both containers and VMs, including macOS VMs on Apple silicon via Lume
- Three entry paths (CLI, code, MCP server) cover human, programmatic, and agent-driven use
- Explicit hygiene rules — delete cloud sandboxes promptly and report the sandbox name used
- Heavy prerequisites: the cua CLI, possibly local runtimes, and optionally the MCP server must all be installed
- Cloud sandboxes are metered; forgetting to delete them incurs cost
- Browser automation requires the cua MCP server and an image shipping cua-spacesd — plain images only get lifecycle, logs, and port-forward
- The source provides no independent test results or performance benchmarks
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| Cua Sandboxes this page | 52 · Use with care | ★ 29k | 1d ago | MIT |
| OpenChamber Performance Engineering Skill | 62 · Recommended | ★ 11k | 3d ago | MIT |
| cmux Workspace Skill | 64 · Recommended | ★ 28k | 1d ago | NOASSERTION |
| cmux-browser: Browser Automation Skill for cmux | 60 · Recommended | ★ 28k | 1d ago | NOASSERTION |
| playwright-cli Browser Automation Skill | 56 · Use with care | ★ 11k | 13d ago | MIT |
Within the same repo, Lume only manages local macOS/Linux VMs on Apple silicon, while this skill additionally covers gVisor containers, cloud sandboxes, direct-attached machines, and in-sandbox browser automation through one CLI; Cua Driver, another sibling, automates the machine you are actually on, whereas this skill emphasizes isolation.
How did FollowSkills review this skill?
SKILL.md explicitly advocates isolation (run in sandbox, not on the user's machine), requires explicit user consent for moving logged-in sessions, mandates deleting cloud sandboxes and reporting names for inspection; data flow is largely transparent. Deducted for: undocumented credential lifecycle for --token, incomplete disclosure of data paths via the cua.ai relay, unverified isolation strength claims for local gVisor/QEMU, and no rollback/recovery guidance.
Instructions are internally self-consistent across CLI, SDK, and MCP entry points, with diagnosable failure behavior ('invalid placement' lists valid values, read-only runtime doctor). Deducted for: static review cannot reproduce key paths (create/exec/screenshot/browse), no test evidence specific to this skill, and incomplete failure feedback for network or image-pull errors and abnormal input.
Trigger conditions and scenarios are clear (isolated machine to run code, test apps, drive GUI, fill forms, take screenshots) with explicit non-fit boundaries (never run destructive commands on the user's own machine). Deducted for: dependence on cua.ai and ghcr.io without mainland-China reachability disclosure, no Chinese-language support statement, and only partial enumeration of unsupported platform/runtime combinations.
Well-layered documentation (prerequisites, create, use, clean up, SDK, MCP, rules), clear MIT licensing with detailed FSL/MIT boundary disclosure in README, good progressive disclosure. Deducted for: no version, changelog, or owner attribution inside the skill itself; hidden assumptions in the cua.ai install-script dependency; known limitations (e.g., Lume restricted to Apple silicon) not consolidated in the skill.
Commands and SDK examples are directly reusable, the browser automation flow is complete with cleanup steps, and clear marginal value over manually provisioning VMs. Deducted for: static review cannot verify that representative outputs are actually usable, value claims (metering, image availability) depend on external service state, and there is no verified output evidence.
The repository contains SECURITY.md, real CI workflows, and test suites, but they target cua-driver/perception components, not this skill's key paths (sandbox create/exec/browse). Deducted for: skill claims rest mainly on self-authored documentation with no independent, reproducible third-party execution evidence.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →