Dev & Engineering

cmux Test Bisect

When a Swift package suite is red on main, this skill bisects old commits on CI and decides per test whether the test went stale or the code regressed.

51/ 100
Use with care

Useful, but reliability, evidence or controls still have material gaps.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 28k
Last updated
1d ago
License
NOASSERTION
git-bisecttest-debuggingciswift-package-manager
+4github-actionsgh-cliregression-triageios-testing

What does this skill do, and when should you use it?

cmux-test-bisect is one of 25 skills bundled in the manaflow-ai/cmux repository, located at skills/cmux-test-bisect/SKILL.md. It targets the pain point that PR CI only runs filtered tests, so full suites can drift red on main for weeks unnoticed: it drives scripts/ci/package_bisect.py to run tests at probe commits on CI and produce a failure matrix. It is repo-specific — it covers only the SwiftPM suites under Packages/iOS and Packages/Shared, and all probes run on CI, never locally. Adopting it only makes sense if you can dispatch CI on manaflow-ai/cmux.

  • Reads the failing test set from the red CI run and separates flakes from real regressions (rerun the full suite twice)
  • Bisects a commit range on CI with package_bisect.py — each probe pushes a bisect/<name>/<sha10> branch with today's iOS CI files laid over the old commit
  • Produces a per-test × per-probe matrix with verdicts like broken by <sha>, flaky, failing at the oldest probe
  • Supports --filter regex narrowing, repeatable --patch to apply known fixes, and --bisect to run a second experiment in parallel
  • Guides reading the culprit PR's diff to decide stale test vs. regression and fix accordingly
  • Verifies with focused test-ios.yml runs and cleans up bisect/ branches afterwards
Good fit
  • A cmux maintainer sees a package suite under Packages/iOS red on main and wants the commit behind each failure
  • PR CI ran only filtered tests and the full suite has drifted for weeks — locate every failure in one pass
  • Someone asks which PR broke a test, and each test needs attribution as stale test or code regression
  • Older commits hang or fail to build for a known reason — apply the fix via --patch and bisect only the actual question
Not a fit
  • Developers outside manaflow-ai/cmux: every command is hardwired to that repo's CI workflow, package names, and runner setup, so it does not port directly
  • Users wanting fast local bisection: the skill explicitly forbids full cmux builds on laptops and depends entirely on CI queuing, up to 45 minutes per wait
  • App-host (cmuxTests) failures on main: the skill says those have their own automatic bisect (#14510) and are out of scope

How do you install this skill?

Before you use it
  • The skill is deeply coupled to manaflow-ai/cmux internal infrastructure (package_bisect.py, test-ios.yml, the org remote, Blacksmith) and is not reusable for other repositories.
  • Using it pushes probe branches to the org remote, dispatches CI runs, and temporarily drops the package-lint gate; be aware of these external side effects before confirming.
  • The core script and test-ios.yml were not present in the reviewed evidence and nothing was executed; confidence is low.
  • Repository license metadata is NOASSERTION and the server directory is BUSL-1.1 (commercial use restricted); verify license compliance before adopting the skill.
  • The skill is English-only with no Chinese-language adaptation.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
Install first
  • GitHub CLI (gh)
  • python3
  • git
  • access to manaflow-ai/cmux CI (test-ios.yml)
  • macOS/Linux checkout of cmux with upstream remote

The README documents no commands for installing the skill collection (the skill doc only covers commands used within the skill), so dedicated install steps are undocumented; the skill lives at skills/cmux-test-bisect/ in the repository and is used as part of it.

Generic route: install into Claude Code manually (macOS / Linux)
tmp="$(mktemp -d)"
git clone --depth 1 https://github.com/manaflow-ai/cmux.git "$tmp"
mkdir -p ~/.claude/skills
cp -R "$tmp/skills/cmux-test-bisect" ~/.claude/skills/
rm -rf "$tmp"

Generated from the source repository and skill path; it copies only this skill's folder. If the author's install steps above differ, follow those first. To scope it to one project, replace ~/.claude/skills with that project's .claude/skills.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • The CmuxMobileShell suite under Packages/iOS is red on main — bisect to find which commit broke it
  • Which PR broke this test? Judge per test whether it went stale or the code regressed
  • PR CI only ran filtered tests and the full suite has drifted for weeks — bisect on CI to locate every failure
  • Older commits fail to build for a known reason — apply --patch and keep bisecting the failing tests

The skill triggers when a suite is red on main, when PR CI only ran filtered tests, or when asked which PR broke a test. Core flow:

  1. Get the failing set via gh run view (grep '✘ Test .* failed after'), run the full suite twice to rule out flakes, and search for an existing fix PR.
  2. In your own worktree from upstream/main, run the bisect:
T=scripts/ci/package_bisect.py
python3 $T --package CmuxMobileShell start --points 6 <old-sha>..upstream/main
python3 $T --package CmuxMobileShell status --wait     # up to 45 min
python3 $T --package CmuxMobileShell next --dispatch
python3 $T --package CmuxMobileShell cleanup
  1. Read the matrix from status (X failed / . passed / - never ran / ? pending / E no test ran due to compile or runner failure). A '-' after an INCOMPLETE probe is never a pass — rerun that window in a second bisect with the hang fix applied via --patch.
  2. Read the culprit PR and decide per test: stale test (fix the test to express its original intent under the new behavior; never delete the assertion) or regression (fix the code, keep the test). When unsure, prefer the reading that keeps the test's name true and say which you chose in the PR.

Options worth knowing: --filter narrows the suite, --paths changes midpoint candidate commits (default: iOS and Shared packages), --patch (repeatable) applies a fix commit to every probe, --bisect <name> runs a second experiment beside the first (--package and --bisect go before the subcommand). State is shared by all worktrees in <git-common-dir>/package-bisect/<name>., with per-attempt log caching so status --refetch only downloads unseen logs.

What are this skill's strengths and limitations?

Pros
  • All probes run on CI — no full cmux builds on your laptop
  • The failure matrix cleanly separates broken by / fixed by / flaky / out-of-range verdicts per test
  • Accounts for real-world traps: old CI scripts on old commits, fake break windows caused by hangs, flake-vs-regression distinction, '-' after INCOMPLETE probes
  • Built-in discipline: cleanup removes bisect/ branches and PRs credit the PR that changed the behavior
Limitations
  • Hard-bound to manaflow-ai/cmux: package names, test-ios.yml, runner config (auto/Blacksmith), and the mf org remote are all repo-specific
  • Each probe wait can be up to 45 minutes; throughput depends on CI queuing
  • GitHub license is NOASSERTION; the repo is GPL-3.0-or-later with parts under BUSL-1.1 (web/ etc.), so confirm the licensing scope for the skill files before adopting

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
cmux Test Bisect this page 51 · Use with care ★ 28k 1d ago NOASSERTION
CI Component Screenshot Baseline Investigator ✓ Microsoft · Official 37 · Not recommended ★ 193k 3d ago MIT
VS Code Smoke Test Assistant ✓ Microsoft · Official 53 · Use with care ★ 193k 3d ago MIT
Hard-Bug Diagnosis Loop 47 · Use with care ★ 281k 3d ago MIT
Testing and CI Conventions for SkillHub 47 · Use with care ★ 5.2k 3d ago Apache-2.0

The skill distinguishes test-ios.yml runs (which have a package job) from ci.yml runs (which do not), and notes app-host (cmuxTests) failures use a separate automatic bisect (#14510); the repo-level tmux comparison concerns the terminal product, not this skill.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Use with care
51/ 100 5-point scale 2.6 / 5
1Trust14 / 25 · 2.8/5

Operations are transparent: probes push `bisect/<name>/<sha>` branches via CI and `cleanup` deletes them, giving rollback; the skill discloses that probes lay current CI files over old commits and drop the package-lint gate. Deductions: pushing branches to the org remote and dispatching paid CI runs happen without an explicit user-confirmation step, and the side effect of disabling a safety (lint) gate is mentioned only in passing; least-privilege and isolation reasoning is incomplete.

2Reliability9 / 20 · 2.3/5

Instructions are self-consistent with unusually good failure-mode coverage (INCOMPLETE probes, `-` never a pass, shallow-checkout fix, adopt/--refetch, watchdog handling). But the core script scripts/ci/package_bisect.py is not present in the evidence, so key paths cannot be reproduced in a static read; capped below 10 and deducted 1 for the unverifiable script.

3Adaptability10 / 15 · 3.3/5

Trigger conditions are precise (suite red on main, filtered PR CI, 'which PR broke a test') and non-fit boundaries are explicit (app-host cmuxTests has its own automatic bisect #14510; only six named SwiftPM packages). Deductions: fully bound to this repository's CI infrastructure (test-ios.yml, Blacksmith runners, org remote); no Chinese-language support; essentially unusable outside the declared repo environment.

4Convention8 / 15 · 2.7/5

Well-layered document (prepare → bisect → read matrix → verdict → land), stable naming, clear verdict table. Deductions: no skill-level license (repo metadata NOASSERTION), no version/changelog or stated maintenance owner, and unversioned references such as #14510.

5Effectiveness6 / 15 · 2.0/5

Addresses a real, high-value problem (filtered PR CI leaving main red for weeks) with a complete path from locating the culprit commit to landing the fix. Deductions: static review cannot verify that the matrix output and landing workflow are directly usable; effect depends on an unseen script and CI behavior; capped at 7 by calibration.

6Verifiability4 / 10 · 2.0/5

Auditable primary material exists in the repo (CI workflows, agents/openai.yaml, detailed process doc), but the referenced package_bisect.py and test-ios.yml are not in the provided evidence and there is no independent third-party execution evidence; capped at 5 by static calibration and deducted 1.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 3dd1fe78859f Review evidence[1][2][3][4][5][6][7][8][9][10][11]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Can I use this skill on my own repository?
Not directly. It hardcodes manaflow-ai/cmux package names, the test-ios.yml workflow, and scripts/ci/package_bisect.py; porting it elsewhere requires substantial rework.
Why do probes push bisect/ branches instead of checking out locally?
Old commits carry old CI scripts, so a plain -f ref=<sha> dispatch fails early; the skill lays today's iOS CI files over the old commit, drops the package-lint gate, and runs on CI.
What does '-' mean in the matrix?
The test never ran at that probe — it is never a pass. If a probe is INCOMPLETE (hung-test watchdog or timeout), every test after it shows '-'; re-probe those commits in a second bisect with the hang fix applied.
How do I tell a stale test from a regression?
Read the culprit PR: if it changed behavior on purpose (PR body, comments, other tests updated the same way), fix the test to express its original intent; if the test's intent still holds, fix the code. When unsure, prefer the reading that keeps the test's name true and state it in the PR.

More skills from this repository

All from manaflow-ai/cmux

Dev & Engineering

cmux Testing Skill

Pick the right scoped local/CI verification for the cmux repo, add behavior-grounded tests, and keep Swift test targets correctly wired in the Xcode project.

★ 28k FS 60 Recommended 1d ago
Dev & Engineering

cmux-capture: Screenshot and Record cmux Windows

Capture screenshots or recordings of a cmux window from the CLI to supply real evidence for PRs, bug reports, and visual verification — with no screen recording permission required.

★ 28k FS 55 Use with care 1d ago
Dev & Engineering

cmux Shared Behavior Rules

Codifies one shared implementation and verification path for cmux behaviors exposed through multiple entrypoints, so fixes never land on one surface and leave others stale.

★ 28k FS 48 Use with care 1d ago
Dev & Engineering

cmux Release Skill

A release-workflow skill for cmux maintainers covering version bumps, changelog assembly, pretag guard, tagging, and release asset verification.

★ 28k FS 53 Use with care 1d ago
Dev & Engineering

cmux Workspace Skill

Lets an AI coding agent work safely inside the cmux workspace that invoked it, without disturbing the workspace or window the user is actually looking at.

★ 28k FS 64 Recommended 1d ago
Dev & Engineering

cmux Architecture Skill

Load cmux's package architecture, layering, dependency inversion, and Swift 6 concurrency rules before adding or heavily rewriting Swift files, packages, coordinators, services, or public APIs in the cmux repository.

★ 28k FS 63 Recommended 1d ago
Dev & Engineering

cmux-browser: Browser Automation Skill for cmux

Open sites, inspect browser surfaces, wait for page state, and extract data through the cmux CLI without stealing focus — built for parallel AI coding agent workflows.

★ 28k FS 60 Recommended 1d ago
Dev & Engineering

cmux Socket Policy Skill

Gives AI coding agents the threading and focus rules for cmux socket/CLI work, so telemetry stays off the main thread and automation never steals your app focus.

★ 28k FS 58 Recommended 1d ago
Automation & Ops

cmux Cloud VM Skill

Lets an AI agent operate cmux Cloud machines through the cmux CLI: run durable remote commands and coding agents on persistent terminals, then present verified results to the user.

★ 28k FS 56 Use with care 1d ago
Dev & Engineering

cmux Debugging Skill

Gives AI agents the project-specific debugging conventions for cmux — a Ghostty-based macOS terminal — so probes, profiling, and UI changes never break typing latency or live agent sessions.

★ 28k FS 56 Use with care 1d ago
Dev & Engineering

cmux Review — Pre-merge Agent Code Review

Reviews agent-written code changes before merge, using an independent review subagent to surface correctness defects and evidence-based verification, while suppressing style noise.

★ 28k FS 55 Use with care 1d ago
Dev & Engineering

cmux Ghostty Submodule Workflow Skill

Standardizes Ghostty submodule commits, GhosttyKit.xcframework rebuilds, and parent pointer updates for cmux contributors, preventing dangling submodule pointers that break CI and checkouts.

★ 28k FS 54 Use with care 1d ago
Dev & Engineering

cmux Dev Workflow Skill

A standardized contributor workflow for cmux: tagged native builds, Xcode project normalization, and sidebar extensions — without disturbing a running cmux instance.

★ 28k FS 53 Use with care 1d ago
Dev & Engineering

cmux Diagnostics

Collects support-safe, read-only diagnostics for cmux so you can pinpoint failing hooks, notifications, session restore, and CLI control without leaking secrets.

★ 28k FS 51 Use with care 1d ago
Dev & Engineering

cmux Core Control

Deterministically control cmux terminal topology — windows, workspaces, panes, surfaces, focus, and attention cues — via the cmux CLI, built for AI coding agent automation.

★ 28k FS 47 Use with care 1d ago
Dev & Engineering

cmux Keyboard Shortcuts

Turns your key preferences into working cmux shortcut bindings, with templates, snapshots, and rollback instead of blind JSON edits.

★ 28k FS 58 Recommended 1d ago
Dev & Engineering

cmux Settings Management Skill (cmux-settings)

Safely view, edit, and roll back cmux's cmux. configuration with hot reload — changes apply on save, no app restart.

★ 28k FS 58 Recommended 1d ago
Dev & Engineering

cmux Localization Rules & Audit Skill

Enforces localization rules and a verifiable audit workflow for every cmux user-facing string change, so no hardcoded English text ever lands.

★ 28k FS 56 Use with care 1d ago
Dev & Engineering

cmux Customization

A skill for safely customizing the cmux terminal: edit cmux., Dock config and Ghostty preferences to tailor actions, layouts, shortcuts and notifications without breaking existing config.

★ 28k FS 55 Use with care 1d ago
Productivity & Collaboration

cmux Markdown Viewer

Opens .md files in a formatted panel beside your cmux terminal with live reload, keeping plans, docs, and notes visible as they change.

★ 28k FS 52 Use with care 1d ago

Related skills