Dev & Engineering

Cua Sandboxes

Spin up disposable Linux or macOS machines locally or in the Cua cloud, so agents can run code, test apps, drive a desktop GUI, and browse the web without ever touching your own computer.

52/ 100
Use with care

Useful, but reliability, evidence or controls still have material gaps.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 29k
Last updated
1d ago
License
MIT
sandboxvirtualizationcontainersgvisor
+4qemumacos-vmweb-automationcua-cli

What does this skill do, and when should you use it?

cua-sandboxes is one of the skills bundled in the trycua/cua monorepo, teaching agents to create, use, and clean up sandboxes — disposable containers or VMs — via the `cua` CLI and the cua SDK. It covers gVisor containers, QEMU Linux VMs, and Lume macOS VMs on Apple silicon, plus any OCI image and direct connections to machines running cua-spacesd. The skill documents three entry paths: terminal commands, code (Python with equivalent TypeScript, Swift, and Kotlin bindings), and the cua MCP server, including a full in-sandbox web-browsing workflow with Chromium. Local sandboxes need no account; cloud sandboxes are metered and must be deleted when done.

  • Creates sandboxes with cua sb create as local gVisor containers, QEMU Linux VMs, or Lume macOS VMs (Apple silicon), in the Cua cloud, from any OCI image, or attached to an existing machine running cua-spacesd
  • Runs commands, interactive shells, screenshots, and browser-based desktop viewing via cua sb exec / shell / screenshot / vnc
  • Exposes the same sandbox lifecycle from code: Python's cua.embedded() plus matching objects in TypeScript, Swift, and Kotlin bindings
  • Drives in-sandbox web browsing through cua MCP tools (images_list, sandbox_create, browser_navigate, browser_click, browser_type) against a Chromium-shipping image
  • Cleans up with cua sb rm, cua sb suspend, and cua fleet pools gc, and pauses/resumes sandboxes
Good fit
  • A developer who needs an isolated machine to run untrusted code or test an app without polluting their own machine
  • An automation agent that must click, fill forms, or take screenshots in a desktop GUI or on a website, keeping the user's real desktop untouched
  • An engineer testing a website or a local web app by running it in the sandbox and opening http://127.0.0.1:<port>
  • A tester who needs a macOS VM environment on Apple silicon via Lume
  • A pipeline that reproduces an issue or generates data in a disposable environment and tears it down afterward
Not a fit
  • Users unwilling to install the cua toolchain — every operation in this skill depends on the cua CLI/SDK
  • Pure API model integrations with no shell or runtime to execute commands in
  • Workflows needing a Windows sandbox image — the skill documents only Linux and macOS images and runtimes (gVisor, QEMU, Lume), even though the broader repo supports Windows

How do you install this skill?

Before you use it
  • Cloud sandboxes depend on cua.ai authentication, relay, and ghcr.io image sources; mainland-China reachability is unverified and core functionality may be unavailable — validate in local mode (gVisor/QEMU) first.
  • Token handling for `--on direct:... --token TOKEN` is undocumented; do not pass long-lived credentials in shared environments.
  • Cloud sandboxes are metered; always run `cua sb rm` after the task as instructed to avoid ongoing charges.
  • The skill assumes images shipping cua-spacesd; with plain OCI images, exec/screenshot/GUI silently degrade to lifecycle-only management.
  • This is a static, non-executed review; all command behavior, image availability, and isolation strength claims are unverified.
  • Publisher is not verified by the FollowSkills registry and is treated as unknown; verify repository and binary provenance independently before use.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
  • MCP Server
Install first
  • cua CLI (Rust)
  • cua SDK (Python/TypeScript/Swift/Kotlin/Rust)
  • cua MCP server for browser workflow
  • cua-spacesd-capable images

Install the full cua toolchain first (this skill ships inside the libs/cua/skills collection):

macOS / Linux (terminal installer)

curl -fsSL https://cua.ai/install.sh | sh

Windows (PowerShell)

irm https://cua.ai/install.ps1 | iex

The installer walks you through cua auth login and can install cua skills and the cua MCP server into AI coding agents such as Claude Code, Codex, and Cursor.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • Run `uname -a` in an isolated Linux sandbox and tell me the output
  • Create a macOS sandbox and take a screenshot of its desktop
  • Open example.com inside a sandbox, read the page, and fill in the search box
  • Test my app on port 8080: start it in a sandbox, visit http://127.0.0.1:8080, and delete the sandbox when done

The skill triggers from its SKILL.md description: any task needing an isolated machine to run code, test an app, drive a GUI, or browse/test a website. Core flow:

cua --version            # confirm the Rust CLI
cua auth status          # cloud requires login
cua sb create linux --name dev   # local gVisor container
cua sb exec dev -- uname -a      # run a command
cua sb screenshot dev            # save a screenshot
cua sb vnc dev                   # open desktop in a browser
cua sb rm dev --force            # delete when done

Options worth knowing: --kind vm for a QEMU Linux VM, cua sb create macos for a Lume macOS VM (Apple silicon), --on cloud for the metered Cua cloud, and -- on any command for machine-readable output. GUI work goes through cua do switch <name> then cua do click/screenshot. Web browsing uses cua MCP tools: images_list, sandbox_create with browser: true, then browser_navigate/browser_click/browser_type and get_browser_state for semantic snapshots. Local runtimes auto-install on first use (cua runtime setup does it explicitly), and logged-in-site teleport requires explicit user consent.

What are this skill's strengths and limitations?

Pros
  • Local sandboxes work with no account, keeping everyday use friction-free
  • One runtime and sandbox list shared by CLI and Python/TypeScript/Swift/Kotlin SDKs, so scripts and agents interoperate
  • Supports both containers and VMs, including macOS VMs on Apple silicon via Lume
  • Three entry paths (CLI, code, MCP server) cover human, programmatic, and agent-driven use
  • Explicit hygiene rules — delete cloud sandboxes promptly and report the sandbox name used
Limitations
  • Heavy prerequisites: the cua CLI, possibly local runtimes, and optionally the MCP server must all be installed
  • Cloud sandboxes are metered; forgetting to delete them incurs cost
  • Browser automation requires the cua MCP server and an image shipping cua-spacesd — plain images only get lifecycle, logs, and port-forward
  • The source provides no independent test results or performance benchmarks

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
Cua Sandboxes this page 52 · Use with care ★ 29k 1d ago MIT
OpenChamber Performance Engineering Skill 62 · Recommended ★ 11k 3d ago MIT
cmux Workspace Skill 64 · Recommended ★ 28k 1d ago NOASSERTION
cmux-browser: Browser Automation Skill for cmux 60 · Recommended ★ 28k 1d ago NOASSERTION
playwright-cli Browser Automation Skill 56 · Use with care ★ 11k 13d ago MIT

Within the same repo, Lume only manages local macOS/Linux VMs on Apple silicon, while this skill additionally covers gVisor containers, cloud sandboxes, direct-attached machines, and in-sandbox browser automation through one CLI; Cua Driver, another sibling, automates the machine you are actually on, whereas this skill emphasizes isolation.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Use with care
52/ 100 5-point scale 2.6 / 5
1Trust15 / 25 · 3.0/5

SKILL.md explicitly advocates isolation (run in sandbox, not on the user's machine), requires explicit user consent for moving logged-in sessions, mandates deleting cloud sandboxes and reporting names for inspection; data flow is largely transparent. Deducted for: undocumented credential lifecycle for --token, incomplete disclosure of data paths via the cua.ai relay, unverified isolation strength claims for local gVisor/QEMU, and no rollback/recovery guidance.

2Reliability9 / 20 · 2.3/5

Instructions are internally self-consistent across CLI, SDK, and MCP entry points, with diagnosable failure behavior ('invalid placement' lists valid values, read-only runtime doctor). Deducted for: static review cannot reproduce key paths (create/exec/screenshot/browse), no test evidence specific to this skill, and incomplete failure feedback for network or image-pull errors and abnormal input.

3Adaptability9 / 15 · 3.0/5

Trigger conditions and scenarios are clear (isolated machine to run code, test apps, drive GUI, fill forms, take screenshots) with explicit non-fit boundaries (never run destructive commands on the user's own machine). Deducted for: dependence on cua.ai and ghcr.io without mainland-China reachability disclosure, no Chinese-language support statement, and only partial enumeration of unsupported platform/runtime combinations.

4Convention9 / 15 · 3.0/5

Well-layered documentation (prerequisites, create, use, clean up, SDK, MCP, rules), clear MIT licensing with detailed FSL/MIT boundary disclosure in README, good progressive disclosure. Deducted for: no version, changelog, or owner attribution inside the skill itself; hidden assumptions in the cua.ai install-script dependency; known limitations (e.g., Lume restricted to Apple silicon) not consolidated in the skill.

5Effectiveness6 / 15 · 2.0/5

Commands and SDK examples are directly reusable, the browser automation flow is complete with cleanup steps, and clear marginal value over manually provisioning VMs. Deducted for: static review cannot verify that representative outputs are actually usable, value claims (metering, image availability) depend on external service state, and there is no verified output evidence.

6Verifiability4 / 10 · 2.0/5

The repository contains SECURITY.md, real CI workflows, and test suites, but they target cua-driver/perception components, not this skill's key paths (sandbox create/exec/browse). Deducted for: skill claims rest mainly on self-authored documentation with no independent, reproducible third-party execution evidence.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision e32127764436 Review evidence[1][2][3][4][5][6][7][8][9][10][11]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Does using this skill cost anything?
Local sandboxes are free and need no account. Cua cloud sandboxes are metered, and the skill's rules require deleting them as soon as the task ends.
Can I use a plain image without cua-spacesd?
Only for lifecycle management, logs, and port-forward. Features that go through cua-spacesd — exec, shell, screenshots, GUI control — need an image that ships it, like the canonical linux alias.
Do I need extra browser tooling for web automation?
No. With the cua MCP server connected, the canonical image ghcr.io/trycua/linux:24.04 ships Chromium; sandbox_create with browser: true is all the setup needed.
Can I attach an existing machine instead of creating one?
Yes. Use `--on direct:<addr> --token TOKEN` to connect to a machine already running cua-spacesd.

More skills from this repository

All from trycua/cua

Automation & Ops

Cua GUI Automation Skill

Give AI agents eyes and hands on a real computer: click buttons, fill forms, and run end-to-end visual QA on any application's GUI.

★ 29k FS 51 Use with care 1d ago
Automation & Ops

Cua Spaces (cua-spaces skill)

Through the cua MCP server, lets your agent work inside a watchable remote or local computer — running commands, moving files, spawning coding agents and sharing the host network.

★ 29k FS 46 Use with care 1d ago
Automation & Ops

Cua Driver GUI Automation Skill

Lets an AI agent operate real native app windows on macOS, Windows, and Linux: observe state, act precisely, and verify the outcome.

★ 29k FS 61 Recommended 1d ago
Automation & Ops

Cua Driver Skill

Let agents operate real desktop applications: drive native GUIs on macOS, Windows, and Linux via accessibility trees and element tokens, then verify from fresh state.

★ 29k FS 58 Recommended 1d ago
Dev & Engineering

jev-use — A Bounded Computer-Use Loop over Cua Driver

Build a tightly bounded computer-use loop on top of Cua Driver: the driver only observes and acts, TypeSafe Jev picks from application-owned candidate IDs, and the caller verifies every action.

★ 29k FS 58 Recommended 1d ago
Automation & Ops

Cua Volume Skill

Give AI agents one versioned, user-level shared volume that persists files across Spaces, syncs to your Mac in seconds, and shares safely with other agents.

★ 29k FS 52 Use with care 1d ago
Dev & Engineering

Poll GitHub Work

Polls and ranks open GitHub issues, RFCs, and PRs so maintainers know what to work on next — or starts one explicitly selected item.

★ 29k FS 52 Use with care 1d ago

Related skills