What does this skill do, and when should you use it?
cmux-review is one skill in the skills collection bundled in manaflow-ai/cmux, focused on reviewing changes written by AI coding agents. Its core principle is reducing developer attention: surface concrete correctness defects, suppress style nits and speculative improvements unless requested. The default flow sends task intent, base/head SHAs and the exact diff to a review subagent that discovers issues independently of the author's reasoning, verifies findings with evidence, and gates the merge. For high-risk changes — security, persistence, concurrency, data loss — it escalates to a full adversarial review protocol with receipts. The skill is Markdown instructions and reference documents only; it ships no executable scripts.
- Reads task intent, base/head SHAs and the exact diff, then hands them to a review subagent in the current runtime to find correctness issues and repository-rule violations independently
- Verifies concrete findings with red-before/green-after test evidence, fixes and pushes only when authorized, then re-runs the original discriminator test
- Obtains a fresh subagent review of the repair delta when fixes were non-trivial
- Merges only when the checks that judge the change pass and the approval rule is met, reporting findings, verification performed and coverage gaps
- Runs the adversarial protocol for high-risk changes: source identity, independent reviewers, counterevidence, executable checks, and a persisted review receipt and report
- Engineers running Claude Code or Codex agents who want a final correctness gate before merging agent changes into main
- After an agent makes substantial edits, requiring a review that is independent of the author's reasoning before accepting them
- Re-reviewing a claimed bug fix by the agent instead of trusting a green test at face value
- Changes touching security, persistence, concurrency or data-loss risk that need the full adversarial deep review
- Teams that want a merge report with evidence, verification steps and coverage gaps for traceability
- Teams wanting style checks, lint or formatting consistency — the skill explicitly suppresses style and nits
- Ordinary human PR review without an agent workflow — the process is built around review subagents and agent-authored changes
- Situations requiring a second model or external review service as a gate — the skill explicitly mandates subagents in the current runtime instead
How do you install this skill?
- Static review only; no skill instructions were executed. Scores are low-confidence source inference.
- Heavily depends on cmux-proprietary runtime (Vault, workspace state, comments CLI) and subagent support; most advanced stages are unusable outside cmux.
- Side-effectful repair and push operations are permitted after authorization; confirm scope, and never commit review receipts or scratch artifacts as the skill requires storing them outside source control.
- Audience is English-speaking developers on GitHub PR flow; no Chinese documentation, and reachability of dogfood/Vercel/CI steps from mainland China is unverified.
- Repository license metadata is NOASSERTION and some directories are BUSL-1.1; verify licensing boundaries before use.
- The protocol is heavyweight; for small low-risk changes the cost/benefit may not justify it — reserve for high-risk or explicitly requested deep reviews.
- Shell / CLI
- Local filesystem
git
The source provides no install commands for this individual skill. The README describes the repo as a collection of 25 skills, with this one at skills/cmux-review/SKILL.md plus reference files under references/. With any Agent Skills-compatible host, the generic approach is to copy the skill folder into the host's skills directory; exact steps are undocumented in the source.
tmp="$(mktemp -d)"
git clone --depth 1 https://github.com/manaflow-ai/cmux.git "$tmp"
mkdir -p ~/.claude/skills
cp -R "$tmp/skills/cmux-review" ~/.claude/skills/
rm -rf "$tmp"Generated from the source repository and skill path; it copies only this skill's folder. If the author's install steps above differ, follow those first. To scope it to one project, replace ~/.claude/skills with that project's .claude/skills.
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- Review this agent-written change before merge: base abc1234, head def5678, correctness issues only
- The agent just heavily refactored the payment module — independently review it before I accept
- The agent claims the concurrency bug is fixed; re-review the repair delta and give me red-before/green-after evidence
- This change touches user-data persistence, run the adversarial deep review
The skill is model-triggered via its description: pre-merge review of agent changes, after substantial edits, or when re-reviewing a repair. The default flow has four steps: give a review subagent the intent, SHAs and diff; verify findings, fix when authorized, and re-run the original discriminator; get a fresh subagent review of non-trivial repair deltas; merge when checks and the approval rule pass. Review stays read-only until an authorized repair, and review receipts, scratch reproductions and generated logs must not be committed. For implementation handoff or merge, follow references/dogfood-and-merge.md for per-change verification; app/runtime/UI merges require explicit user approval after dogfood or a direct merge directive (merge, merge it, auto-merge — not finish, lgtm, or ship it). main is nightly: stack fixes, do not revert. For security, persistence, concurrency, data-loss risks or requested deep reviews, switch to the adversarial protocol in references/ (discovery and triage, challenge and verification, receipt persistence).
What are this skill's strengths and limitations?
- Independent subagent review keeps discovery free of the author's reasoning, and the red-before/green-after requirement catches hollow green tests
- Clear scope discipline: correctness defects only, style suppressed — minimizes reviewer attention cost
- Adversarial protocol covers security, concurrency and data-loss dimensions with persisted receipts and pre/post-repair source coordinates
- Merge approval rule is concrete and executable, spelling out which directives constitute merge authorization
- No automated tests or scripts ship with the skill; effectiveness depends entirely on the host model following the instructions
- The subagent mechanism varies across Agent Skills hosts, so portability has real friction
- No install commands, demos or benchmarks in the source make the practical payoff hard to assess
- The repo license is NOASSERTION (README mentions GPL-3.0-or-later plus Business Source License 1.1 for server components); confirm the skill files' specific licensing before commercial use
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| cmux Review — Pre-merge Agent Code Review this page | 55 · Use with care | ★ 28k | 1d ago | NOASSERTION |
| Subagent-Driven Development | 50 · Use with care | ★ 297k | 16d ago | MIT |
| PR Babysitter: Watch Pull Requests Until Merge | 47 · Use with care | ★ 98k | 3d ago | Apache-2.0 |
| Imprint — Your Working Imprint for AI | 46 · Use with care | ★ 103 | 19d ago | MIT |
| PR Submission Conventions | 40 · Not recommended | ★ 5.2k | 3d ago | Apache-2.0 |
SKILL.md distinguishes itself from common alternatives: it refuses second models or external review services as gates, insisting on subagents in the current runtime; and unlike lint or CI tooling it does not check style, focusing on correctness defects and evidence-based verification.
How did FollowSkills review this skill?
Default read-only review, repair only when authorized, prohibition on committing receipts/scratch artifacts, receipts stored outside source control, no external service as a gate, and a schema-backed receipt mechanism give good data-flow and external-effect disclosure; deducted because side-effectful operations (git push, repair commits, media upload) lack explicit confirmation gates (only merge requires approval), authorization boundaries are left to the agent, and the skill itself declares no permission scope.
Well self-consistent: the discovery-challenge-verify-repair-receipt loop is closed, with P0-P3 severity, verification-result enums and evidence kinds; deducted below the static cap because it is a pure instruction skill with no executable tests or scripts shipped, key paths are not statically reproduced, cross-references (cmux-testing, Vault, cmux CLI) are unverified in the given files, and degraded behavior on environments without cmux/subagents is only partially specified.
Triggers are clear (pre-merge, after substantial edits, repair re-review, requested deep review) and the description is semantically precise; deducted because capability boundaries are incompletely declared — strong dependence on cmux runtime (workspace state, Vault, comments CLI), subagent support and GitHub PR flow, only partial guidance for non-cmux environments; no Chinese documentation, and its dogfood/Vercel/GitHub workflow has uncertain reachability from mainland China.
Good layered documentation (main file + references + JSON schema), complete and rigorous receipt schema, stable naming; deducted because the skill itself has no version/changelog, no FAQ or known-limitations section, repository license metadata is NOASSERTION with a mixed GPL/BUSL scheme users must disentangle, and maintenance responsibility is only implicit in repository ownership.
Goals are clear (reduce human review burden via independent discovery, adversarial challenge, executable verification) and the methodology beats generic prompting; deducted because static review cannot confirm output usability, the process is heavyweight (receipts, dual discovery, fresh-context re-review) with questionable cost/benefit for small low-risk changes, and part of the value depends on cmux-proprietary primitives.
A review-receipt JSON schema provides an auditable structured-evidence framework and the PROVEN/DERIVED provenance vocabulary is rigorous; deducted below the static cap because there is no CI evidence or third-party execution covering the skill's key paths, several referenced files are not supplied, and conclusions cannot be independently reproduced.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →