What does this skill do, and when should you use it?
cmux-browser is one of 25 skills bundled in the manaflow-ai/cmux repository (skills/cmux-browser/SKILL.md) for driving cmux's built-in browser, whose scriptable API is ported from agent-browser. Its defining discipline is explicit targeting: discover browser surfaces with read-only identify/tree commands, then act through an explicit --surface handle instead of guessing from visual focus. The skill covers creating surfaces, waiting on page state, interactive snapshots, click/fill interactions, exact viewport sizing, and download-record queries, and it is honest about WKWebView limits such as no network interception. It suits macOS developers already using cmux who want agents to verify web changes against a dev server in the background.
- Creates a browser surface with
cmux -- browser open <url> --focus falseand extracts the surface_ref from the JSON response - Discovers existing browser surfaces via
cmux tree --all --piped through jq, matching on URL or title while never logging the sensitive value - Runs surface-bound operations like
get url,snapshot --interactive,fill,click, andwait(selector/text/url/load-state/function) against an explicit handle - Sets an exact 1–4096 CSS-pixel logical viewport with
viewport <width> <height>, aspect-fitted inside the existing pane - Queries download records with
download list --(download_id, filename, saved path, status, byte count) without opening files or consuming waiters - Falls back to
get text body/get html bodywhen snapshot or eval hits a js_error on complex pages
- A developer running multiple parallel Claude Code sessions in cmux who wants an agent to open a page in the background and verify web changes without stealing focus from the pane you're working in
- Users automating form flows: the skill ships templates/form-automation.sh, a snapshot/ref fill loop requiring an explicit surface
- Users who need to log in once and reuse session state: templates/authenticated-session.sh covers OAuth/2FA patterns and state save/load (details in references/authentication.md)
- Anyone who knows a page by URL or title and needs to locate that exact open browser tab — the jq matching script against `cmux tree --all --` is ready to copy
- Developers needing a fixed viewport for screenshots or responsive testing, achievable without disturbing pane layout
- Users tracking a browser download's progress or result on a specific surface via JSON download records
- Non-macOS users: cmux is a macOS-only native Swift/AppKit app and the browser runs on WKWebView, so Linux/Windows are out of scope entirely
- Anyone needing network route interception, offline emulation, trace/screencast recording, or raw input injection — these Chrome/CDP-only APIs return not_supported on WKWebView
- Users without the cmux app and its CLI installed: every command depends on the `cmux` binary, and the matching scripts also require jq
How do you install this skill?
- Static review only; no commands were executed. Verify syntax against the actual binary via cmux browser --help.
- Hard prerequisites of macOS and the cmux app (WKWebView); unusable elsewhere; no Chinese documentation.
- Auth state files contain cookies/tokens: store with restrictive permissions as documented, clean up after tasks, never commit or paste them.
- The installer pulls [email protected] via npx — an external supply-chain dependency; pin and review changes.
- Repository license is GPL-3.0-or-later mixed with BUSL directories (web/ etc.); registry metadata is NOASSERTION — verify licensing before commercial use.
- Publisher is unverified by FollowSkills and identity is unknown; cross-check downloads and updates against the pinned revision.
- Shell / CLI
- Network access
- Local filesystem
cmux app (macOS) with its CLIjq
The skill is distributed via the Vercel skills installer; the repository copies (.claude/skills/cmux-browser and .agents/skills/cmux-browser) point at skills/cmux-browser and are the source of truth. Three install routes are documented:
Claude Code and Codex (from a local checkout, while developing the skill)
npx --yes [email protected] add . --global --yes --skill cmux-browser --agent claude-code codex --copyClaude Code and Codex (from the published repository)
npx --yes [email protected] add manaflow-ai/cmux --global --yes --skill cmux-browser --agent claude-code codex --copyCodex-only destination (skills.sh)
./skills.sh --dest "$HOME/.codex/skills" --skill cmux-browserRestart your agent session after a refresh if it cached the previous document. Prerequisite: install the cmux app itself (DMG from GitHub Releases, or brew tap manaflow-ai/cmux && brew install --cask cmux).
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- Open http://localhost:3000/dashboard in cmux in the background, wait for load, and take an interactive snapshot without stealing my focus
- Find the browser tab whose title is 'Checkout – Staging' and read back its URL
- Use cmux browser to fill in the username and password on the login page and submit, re-snapshotting after each step
- Set this browser surface's viewport to 1024x768 and capture the page
The skill is triggered whenever an agent needs browser automation, and it prescribes a fixed flow: first verify the CLI contract with cmux browser --help and cmux --version; then discover surfaces with read-only cmux identify -- or cmux tree --all -- (never via focus/select commands). The core loop is: open (--focus false) → extract surface_ref from the JSON response → pass --surface explicitly for every get url / wait / snapshot / fill / click → re-snapshot after navigation or major DOM changes since refs go stale. Two conventions matter: prefer the flag form in scripts (--surface "$SURFACE") so the target is unmissable, and use get url and snapshot --interactive in new documentation rather than aliases like url or -i. Before viewport emulation, close or detach the browser inspector — an attached inspector resets emulation to native sizing. Deeper references include commands.md (full alias mapping and viewport error codes), snapshot-refs.md, session-management.md, and proxy-support.md.
What are this skill's strengths and limitations?
- Strong discipline: read-only discovery plus explicit surface handles eliminates the most common automation failure mode — acting on the wrong focused pane
- Copy-paste-ready scripts, including a jq matcher that locates a surface by URL/title without ever logging the sensitive value
- Honest capability boundaries: an explicit not_supported list for WKWebView and a documented js_error troubleshooting path
- Backed by 7 deep-dive reference docs and 3 ready-to-use templates (form automation, authenticated sessions, capture workflow)
- Hard-wired to the cmux ecosystem: requires the macOS-only cmux app; useless without it
- WKWebView has real gaps versus Chrome/CDP: no network interception, no trace recording, no offline emulation
- The skill itself warns the CLI may evolve and that SKILL.md and the binary can disagree — expect occasional version drift requiring a refresh
- The repo's License field is NOASSERTION: cmux core is GPL-3.0-or-later but web/ server components are BSL 1.1, so verify licensing scope before adoption
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| cmux-browser: Browser Automation Skill for cmux this page | 60 · Recommended | ★ 28k | 1d ago | NOASSERTION |
| cmux Workspace Skill | 64 · Recommended | ★ 28k | 1d ago | NOASSERTION |
| cmux Diagnostics | 51 · Use with care | ★ 28k | 1d ago | NOASSERTION |
| cmux Shared Behavior Rules | 48 · Use with care | ★ 28k | 1d ago | NOASSERTION |
| cmux Settings Management Skill (cmux-settings) | 58 · Recommended | ★ 28k | 1d ago | NOASSERTION |
cmux's browser API is explicitly ported from vercel-labs/agent-browser, and the README contrasts cmux with tmux — tmux is a multiplexer inside any terminal, while cmux is a native macOS GUI app with a built-in browser, vertical tabs, and a socket API. If you don't use cmux and just want general-purpose browser automation, Playwright or agent-browser itself would be more direct choices.
How did FollowSkills review this skill?
The skill repeatedly enforces least privilege and data protection: read-only discovery, explicit surface handles, credentials via env vars, umask 077/chmod 600 on state files, no committing cookies, redacted logging, and cleanup instructions. No red-line risks. Deducted because static review cannot verify actual CLI behavior, isolation guarantees are thin, publisher is unverified, and rollback is limited to skill refresh.
Internally consistent documentation (command shapes, aliases, surface contract cross-referenced), scripts include error checks and explicit stderr feedback, plus troubleshooting for js_error and stale refs. Deducted: static review caps this at 10; no committed tests or third-party execution evidence, and command syntax relies entirely on self-verification via the binary's --help.
Trigger conditions are clear in the description (open sites, inspect browser, extract data) with well-disclosed boundaries (macOS/WKWebView only, not_supported capability list, no network interception). Deducted: hard prerequisite of macOS plus the cmux app narrows the audience; no Chinese documentation; limited fit for Chinese-language scenarios.
Good layered architecture (SKILL.md + 7 references + templates + AGENTS.md), pinned installer at [email protected], clear mirror provenance. Deducted: registry license metadata is NOASSERTION (repo is GPL/BUSL mixed but the skill directory declares no license); no per-skill version or changelog; maintenance ownership only implicit in the repo.
Workflows cover discovery→open→interact→wait→download→cleanup with directly usable command templates that clearly beat manually probing the CLI. Deducted: static review caps this at 7; zero execution validation, so doc/binary consistency and output usability remain inferred.
Auditable primary material exists (detailed CLI reference mirroring the repo's own binary, agent-browser mapping notes, structured error codes). Deducted: capped at 5; no test suite or CI evidence covering the skill's key paths, so all conclusions are documentary inference rather than independent reproduction.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →