What does this skill do, and when should you use it?
This is the cmux-testing skill (one of 25 bundled skills) in the manaflow-ai/cmux monorepo, aimed at developers contributing to cmux, a Swift/AppKit macOS terminal app. It guides you through choosing scoped checks with scripts/verify-local.py, writing failing regression tests before fixes (two commits with recorded evidence), and wiring new cmuxTests/*.swift files into project.pbxproj via scripts like sync-test-wiring. It also codifies Swift Testing conventions, native test-evidence requirements, and PR CI label rules. Useful for cmux contributors; not for end users of the terminal app.
- Runs python3 scripts/verify-local.py (with --all, --only swift-syntax --swift-changed, or --only test-wiring) from a trusted checkout to select and execute local verification
- Inspects UI test behavior frame-by-frame with scripts/ui-test, and renders view code to light/dark PNGs in seconds via scripts/ui-lab/ui-lab.py
- Drives the app from CI with scripts/run-e2e.sh --scenario JSON tours, producing screenshots, GIFs, and accessibility trees
- Enforces a two-commit regression discipline: failing behavioral test first, fix second, with command, SHAs, and outcomes recorded
- Reconciles cmuxTests PBXFileReference/Sources membership with ./scripts/sync-test-wiring and wires new app sources with ./scripts/wire-app-sources.py
- Mandates Swift Testing (import Testing, @Test, #expect) for Swift unit/integration targets while keeping UI tests on XCTest
- A contributor to cmux who wants the minimal set of local checks before opening a PR touching Swift files
- A contributor who added a new cmuxTests/*.swift file and must verify it is wired into the Xcode project to avoid a misleading zero-test pass
- A developer fixing a bug who needs to follow the prescribed failing-regression-test-then-fix commit discipline with evidence
- A reviewer who wants to read the CI dogfood comment's screenshots and GIFs before merging an app PR
- A contributor changing remote tmux sizing who needs the E2E recipe for validation
- End users who just want to install and use the cmux terminal app — this skill targets the repo's contribution/testing workflow
- Projects outside the cmux repo — scripts like verify-local.py and sync-test-wiring are bound to cmux's repo structure
- Non-macOS environments — cmux is a Swift/AppKit macOS app; UI tests and Xcode project wiring require the macOS toolchain
How do you install this skill?
- This is a static source-only review; no scripts or tests were executed. Confidence is low.
- Repository license metadata is NOASSERTION; actual licensing is GPL-3.0 plus BUSL-1.1 (web/ and related dirs) — verify scope before commercial or self-hosted use.
- The skill depends heavily on macOS, Xcode, tmux, and GitHub Actions (including Blacksmith runners); it is essentially unusable outside macOS, and CI dispatch may be slow or unreliable from mainland-China networks.
- English-only documentation; no Chinese support.
- Confirm a trusted checkout before running repository commands; even --help loads repository code.
- Publisher is not verified by the FollowSkills registry; treat identity as unknown.
- Shell / CLI
- Local filesystem
Python 3Xcode / cmux.xcodeprojSwift toolchaincmux repo checkout (scripts/verify-local.py, scripts/ui-test, scripts/ui-lab, scripts/sync-test-wiring, scripts/wire-app-sources.py)
This skill is a SKILL.md under the skills/ directory of the cmux monorepo; the source provides no standalone install commands. Obtain the skill collection by cloning the repo:
git clone https://github.com/manaflow-ai/cmux.git(The README's DMG/Homebrew instructions install the cmux app, not this skill.)
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- I changed a few Swift files in cmux — help me decide which local verifications to run before committing
- I added a new cmuxTests test file — check whether it's properly wired in project.pbxproj
- I'm fixing a cmux bug — set up the failing regression test commit and the fix commit per the discipline
- Look at this PR's CI dogfood comment screenshots and tell me whether the UI behavior is correct
Triggered when adding tests or deciding what local/CI evidence a change needs. Core flow: run verify-local.py to pick checks from a trusted checkout; when fixing bugs, keep a focused failing command, use two commits, and record SHAs and results; after creating/renaming/deleting cmuxTests/*.swift files run sync-test-wiring (--check validates without writing); wire app sources with wire-app-sources.py, UI tests with --target cmuxUITests --dir cmuxUITests. Common commands:
python3 scripts/verify-local.py --all
python3 scripts/verify-local.py --only test-wiring
./scripts/sync-test-wiring
./scripts/wire-app-sources.py --checkNote: parsing checks only validate syntax — no typecheck, no tests run; an app build does not compile test targets, so package/API changes need the relevant test target compiled and executed.
What are this skill's strengths and limitations?
- Provides a complete verification decision path from local static checks to CI dogfood screenshots
- Explicitly guards against common traps: zero-executed-test false passes, unwired test files, tests that merely mirror implementation details
- Clear rules on Swift Testing vs. XCTest migration boundaries, preventing unnecessary churn
- Only meaningful within the cmux repo; scripts and paths are tightly bound to it
- No independent test suite or cross-platform validation evidence for the skill itself in the source
- Some verification (UI tests, Xcode wiring, dogfood) requires macOS and Xcode — a high prerequisite bar
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| cmux Testing Skill this page | 60 · Recommended | ★ 28k | 1d ago | NOASSERTION |
| cmux Shared Behavior Rules | 48 · Use with care | ★ 28k | 1d ago | NOASSERTION |
| cmux Dev Workflow Skill | 53 · Use with care | ★ 28k | 1d ago | NOASSERTION |
| cmux Workspace Skill | 64 · Recommended | ★ 28k | 1d ago | NOASSERTION |
| cmux-browser: Browser Automation Skill for cmux | 60 · Recommended | ★ 28k | 1d ago | NOASSERTION |
The README compares cmux with tmux (cmux is a GUI-native macOS app, not a terminal multiplexer), but that is at the app level; the source names no direct alternative for this testing skill.
How did FollowSkills review this skill?
Evidence shows the skill explicitly restricts repository commands to a trusted checkout and discloses that even --help loads repository code; the skill itself is pure guidance, requests no credentials, and routes e2e/UI tests through commit-pinned CI dispatch rather than running untagged apps locally. Deducted: referenced scripts were not directly reviewed in this scope, so least-privilege and external-effect claims rest on the skill's descriptions; publisher unverified.
SKILL.md and references are highly self-consistent: they distinguish compiling test targets from executing tests, reject zero-test invocations as verification, warn about misleading zero-test passes from unwired files, and provide detailed failure diagnostics (shim-check, pane_grids introspection, concrete fuzz setup-failure fixes). Deducted: static review did not execute; runnability of key paths (verify-local.py, sync-test-wiring, fuzz scripts) and dependency availability (tmux, Xcode, GitHub runners) are unverified, so no score above 10.
Audience and scenarios are clear (cmux contributors/agents choosing tests and validation evidence), the description gives precise trigger conditions, and boundaries are disclosed (ui-lab states it is not the running app; the sizing suite states hermeticity requirements). Deducted: hard macOS/Xcode/GitHub Actions dependency without declared non-fit ranges for other environments; no Chinese-language support; CI dispatch relies on overseas GitHub Actions/Blacksmith runners with uncertain mainland-China reachability.
Information architecture is well layered: SKILL.md plus multiple deep reference guides, clear progressive disclosure, stable command naming and cross-links. Repository-level LICENSE and SECURITY.md exist. Deducted: repository license metadata is NOASSERTION and the actual licensing is a mixed GPL-3.0/BUSL-1.1 scheme; the skill itself declares no version or changelog, and maintenance/update responsibility is not stated inside the skill files.
The skill gives directly usable commands and decision rules for its core tasks (scoped verification, behavioral tests, wiring repair) with clear marginal value over unguided trial and error. Deducted: static review cannot verify output usability; value depends on the repository toolchain existing and being correct, and key validation ultimately costs 10-20 CI runner-minutes per the docs, so cost/benefit is unproven without execution.
The repository contains real CI workflows (e.g. app-host-test-rerun.yml with SHA-pinned action references and scoped permissions) that corroborate the skill's claims, which are mostly traceable in-source. Deducted: static review confirms file presence and consistency only, not independent reproduction; some referenced paths (references, docs) were not provided in this evidence set.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →