Dev & Engineering

cmux Review — Pre-merge Agent Code Review

Reviews agent-written code changes before merge, using an independent review subagent to surface correctness defects and evidence-based verification, while suppressing style noise.

55/ 100
Use with care

Useful, but reliability, evidence or controls still have material gaps.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 28k
Last updated
1d ago
License
NOASSERTION
code-reviewpre-merge-reviewadversarial-protocolgit-workflow
+3subagent-reviewmerge-approvalverification-evidence

What does this skill do, and when should you use it?

cmux-review is one skill in the skills collection bundled in manaflow-ai/cmux, focused on reviewing changes written by AI coding agents. Its core principle is reducing developer attention: surface concrete correctness defects, suppress style nits and speculative improvements unless requested. The default flow sends task intent, base/head SHAs and the exact diff to a review subagent that discovers issues independently of the author's reasoning, verifies findings with evidence, and gates the merge. For high-risk changes — security, persistence, concurrency, data loss — it escalates to a full adversarial review protocol with receipts. The skill is Markdown instructions and reference documents only; it ships no executable scripts.

  • Reads task intent, base/head SHAs and the exact diff, then hands them to a review subagent in the current runtime to find correctness issues and repository-rule violations independently
  • Verifies concrete findings with red-before/green-after test evidence, fixes and pushes only when authorized, then re-runs the original discriminator test
  • Obtains a fresh subagent review of the repair delta when fixes were non-trivial
  • Merges only when the checks that judge the change pass and the approval rule is met, reporting findings, verification performed and coverage gaps
  • Runs the adversarial protocol for high-risk changes: source identity, independent reviewers, counterevidence, executable checks, and a persisted review receipt and report
Good fit
  • Engineers running Claude Code or Codex agents who want a final correctness gate before merging agent changes into main
  • After an agent makes substantial edits, requiring a review that is independent of the author's reasoning before accepting them
  • Re-reviewing a claimed bug fix by the agent instead of trusting a green test at face value
  • Changes touching security, persistence, concurrency or data-loss risk that need the full adversarial deep review
  • Teams that want a merge report with evidence, verification steps and coverage gaps for traceability
Not a fit
  • Teams wanting style checks, lint or formatting consistency — the skill explicitly suppresses style and nits
  • Ordinary human PR review without an agent workflow — the process is built around review subagents and agent-authored changes
  • Situations requiring a second model or external review service as a gate — the skill explicitly mandates subagents in the current runtime instead

How do you install this skill?

Before you use it
  • Static review only; no skill instructions were executed. Scores are low-confidence source inference.
  • Heavily depends on cmux-proprietary runtime (Vault, workspace state, comments CLI) and subagent support; most advanced stages are unusable outside cmux.
  • Side-effectful repair and push operations are permitted after authorization; confirm scope, and never commit review receipts or scratch artifacts as the skill requires storing them outside source control.
  • Audience is English-speaking developers on GitHub PR flow; no Chinese documentation, and reachability of dogfood/Vercel/CI steps from mainland China is unverified.
  • Repository license metadata is NOASSERTION and some directories are BUSL-1.1; verify licensing boundaries before use.
  • The protocol is heavyweight; for small low-risk changes the cost/benefit may not justify it — reserve for high-risk or explicitly requested deep reviews.
Before you start
Your agent needs
  • Shell / CLI
  • Local filesystem
Install first
  • git

The source provides no install commands for this individual skill. The README describes the repo as a collection of 25 skills, with this one at skills/cmux-review/SKILL.md plus reference files under references/. With any Agent Skills-compatible host, the generic approach is to copy the skill folder into the host's skills directory; exact steps are undocumented in the source.

Generic route: install into Claude Code manually (macOS / Linux)
tmp="$(mktemp -d)"
git clone --depth 1 https://github.com/manaflow-ai/cmux.git "$tmp"
mkdir -p ~/.claude/skills
cp -R "$tmp/skills/cmux-review" ~/.claude/skills/
rm -rf "$tmp"

Generated from the source repository and skill path; it copies only this skill's folder. If the author's install steps above differ, follow those first. To scope it to one project, replace ~/.claude/skills with that project's .claude/skills.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • Review this agent-written change before merge: base abc1234, head def5678, correctness issues only
  • The agent just heavily refactored the payment module — independently review it before I accept
  • The agent claims the concurrency bug is fixed; re-review the repair delta and give me red-before/green-after evidence
  • This change touches user-data persistence, run the adversarial deep review

The skill is model-triggered via its description: pre-merge review of agent changes, after substantial edits, or when re-reviewing a repair. The default flow has four steps: give a review subagent the intent, SHAs and diff; verify findings, fix when authorized, and re-run the original discriminator; get a fresh subagent review of non-trivial repair deltas; merge when checks and the approval rule pass. Review stays read-only until an authorized repair, and review receipts, scratch reproductions and generated logs must not be committed. For implementation handoff or merge, follow references/dogfood-and-merge.md for per-change verification; app/runtime/UI merges require explicit user approval after dogfood or a direct merge directive (merge, merge it, auto-merge — not finish, lgtm, or ship it). main is nightly: stack fixes, do not revert. For security, persistence, concurrency, data-loss risks or requested deep reviews, switch to the adversarial protocol in references/ (discovery and triage, challenge and verification, receipt persistence).

What are this skill's strengths and limitations?

Pros
  • Independent subagent review keeps discovery free of the author's reasoning, and the red-before/green-after requirement catches hollow green tests
  • Clear scope discipline: correctness defects only, style suppressed — minimizes reviewer attention cost
  • Adversarial protocol covers security, concurrency and data-loss dimensions with persisted receipts and pre/post-repair source coordinates
  • Merge approval rule is concrete and executable, spelling out which directives constitute merge authorization
Limitations
  • No automated tests or scripts ship with the skill; effectiveness depends entirely on the host model following the instructions
  • The subagent mechanism varies across Agent Skills hosts, so portability has real friction
  • No install commands, demos or benchmarks in the source make the practical payoff hard to assess
  • The repo license is NOASSERTION (README mentions GPL-3.0-or-later plus Business Source License 1.1 for server components); confirm the skill files' specific licensing before commercial use

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
cmux Review — Pre-merge Agent Code Review this page 55 · Use with care ★ 28k 1d ago NOASSERTION
Subagent-Driven Development 50 · Use with care ★ 297k 16d ago MIT
PR Babysitter: Watch Pull Requests Until Merge 47 · Use with care ★ 98k 3d ago Apache-2.0
Imprint — Your Working Imprint for AI 46 · Use with care ★ 103 19d ago MIT
PR Submission Conventions 40 · Not recommended ★ 5.2k 3d ago Apache-2.0

SKILL.md distinguishes itself from common alternatives: it refuses second models or external review services as gates, insisting on subagents in the current runtime; and unlike lint or CI tooling it does not check style, focusing on correctness defects and evidence-based verification.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Use with care
55/ 100 5-point scale 2.8 / 5
1Trust16 / 25 · 3.2/5

Default read-only review, repair only when authorized, prohibition on committing receipts/scratch artifacts, receipts stored outside source control, no external service as a gate, and a schema-backed receipt mechanism give good data-flow and external-effect disclosure; deducted because side-effectful operations (git push, repair commits, media upload) lack explicit confirmation gates (only merge requires approval), authorization boundaries are left to the agent, and the skill itself declares no permission scope.

2Reliability11 / 20 · 2.8/5

Well self-consistent: the discovery-challenge-verify-repair-receipt loop is closed, with P0-P3 severity, verification-result enums and evidence kinds; deducted below the static cap because it is a pure instruction skill with no executable tests or scripts shipped, key paths are not statically reproduced, cross-references (cmux-testing, Vault, cmux CLI) are unverified in the given files, and degraded behavior on environments without cmux/subagents is only partially specified.

3Adaptability9 / 15 · 3.0/5

Triggers are clear (pre-merge, after substantial edits, repair re-review, requested deep review) and the description is semantically precise; deducted because capability boundaries are incompletely declared — strong dependence on cmux runtime (workspace state, Vault, comments CLI), subagent support and GitHub PR flow, only partial guidance for non-cmux environments; no Chinese documentation, and its dogfood/Vercel/GitHub workflow has uncertain reachability from mainland China.

4Convention9 / 15 · 3.0/5

Good layered documentation (main file + references + JSON schema), complete and rigorous receipt schema, stable naming; deducted because the skill itself has no version/changelog, no FAQ or known-limitations section, repository license metadata is NOASSERTION with a mixed GPL/BUSL scheme users must disentangle, and maintenance responsibility is only implicit in repository ownership.

5Effectiveness6 / 15 · 2.0/5

Goals are clear (reduce human review burden via independent discovery, adversarial challenge, executable verification) and the methodology beats generic prompting; deducted because static review cannot confirm output usability, the process is heavyweight (receipts, dual discovery, fresh-context re-review) with questionable cost/benefit for small low-risk changes, and part of the value depends on cmux-proprietary primitives.

6Verifiability4 / 10 · 2.0/5

A review-receipt JSON schema provides an auditable structured-evidence framework and the PROVEN/DERIVED provenance vocabulary is rigorous; deducted below the static cap because there is no CI evidence or third-party execution covering the skill's key paths, several referenced files are not supplied, and conclusions cannot be independently reproduced.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 3dd1fe78859f Review evidence[1][2][3][4][5][6][7][8][9][10][11]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Does using this skill cost anything?
The skill files are free in the cmux repo; the README states the cmux app is free and open source. However the repo license is NOASSERTION — the code is GPL-3.0-or-later while server components use Business Source License 1.1 — so confirm the licensing scope before commercial use.
Will it modify my code on its own?
Review is read-only by default; fixes are made only when authorized, must be followed by re-running the original discriminator, and non-trivial repairs get a fresh subagent review of the repair delta.
Does a green test suite mean the change is fine?
No. The skill requires that a green test counts only if it exercises the asserted failure, preferring red-before/green-after evidence.
Does it work outside Claude Code?
The SKILL.md is plain Markdown with standard YAML frontmatter, so it should port to other Agent Skills hosts; however, phrases like 'use subagents in the current runtime' depend on host subagent support and may need light adaptation.

More skills from this repository

All from manaflow-ai/cmux

Dev & Engineering

cmux Architecture Skill

Load cmux's package architecture, layering, dependency inversion, and Swift 6 concurrency rules before adding or heavily rewriting Swift files, packages, coordinators, services, or public APIs in the cmux repository.

★ 28k FS 63 Recommended 1d ago
Dev & Engineering

cmux Ghostty Submodule Workflow Skill

Standardizes Ghostty submodule commits, GhosttyKit.xcframework rebuilds, and parent pointer updates for cmux contributors, preventing dangling submodule pointers that break CI and checkouts.

★ 28k FS 54 Use with care 1d ago
Dev & Engineering

cmux Testing Skill

Pick the right scoped local/CI verification for the cmux repo, add behavior-grounded tests, and keep Swift test targets correctly wired in the Xcode project.

★ 28k FS 60 Recommended 1d ago
Dev & Engineering

cmux-capture: Screenshot and Record cmux Windows

Capture screenshots or recordings of a cmux window from the CLI to supply real evidence for PRs, bug reports, and visual verification — with no screen recording permission required.

★ 28k FS 55 Use with care 1d ago
Dev & Engineering

cmux Dev Workflow Skill

A standardized contributor workflow for cmux: tagged native builds, Xcode project normalization, and sidebar extensions — without disturbing a running cmux instance.

★ 28k FS 53 Use with care 1d ago
Dev & Engineering

cmux Release Skill

A release-workflow skill for cmux maintainers covering version bumps, changelog assembly, pretag guard, tagging, and release asset verification.

★ 28k FS 53 Use with care 1d ago
Dev & Engineering

cmux Test Bisect

When a Swift package suite is red on main, this skill bisects old commits on CI and decides per test whether the test went stale or the code regressed.

★ 28k FS 51 Use with care 1d ago
Dev & Engineering

cmux Workspace Skill

Lets an AI coding agent work safely inside the cmux workspace that invoked it, without disturbing the workspace or window the user is actually looking at.

★ 28k FS 64 Recommended 1d ago
Dev & Engineering

cmux-browser: Browser Automation Skill for cmux

Open sites, inspect browser surfaces, wait for page state, and extract data through the cmux CLI without stealing focus — built for parallel AI coding agent workflows.

★ 28k FS 60 Recommended 1d ago
Dev & Engineering

cmux Keyboard Shortcuts

Turns your key preferences into working cmux shortcut bindings, with templates, snapshots, and rollback instead of blind JSON edits.

★ 28k FS 58 Recommended 1d ago
Dev & Engineering

cmux Settings Management Skill (cmux-settings)

Safely view, edit, and roll back cmux's cmux. configuration with hot reload — changes apply on save, no app restart.

★ 28k FS 58 Recommended 1d ago
Dev & Engineering

cmux Socket Policy Skill

Gives AI coding agents the threading and focus rules for cmux socket/CLI work, so telemetry stays off the main thread and automation never steals your app focus.

★ 28k FS 58 Recommended 1d ago
Automation & Ops

cmux Cloud VM Skill

Lets an AI agent operate cmux Cloud machines through the cmux CLI: run durable remote commands and coding agents on persistent terminals, then present verified results to the user.

★ 28k FS 56 Use with care 1d ago
Dev & Engineering

cmux Debugging Skill

Gives AI agents the project-specific debugging conventions for cmux — a Ghostty-based macOS terminal — so probes, profiling, and UI changes never break typing latency or live agent sessions.

★ 28k FS 56 Use with care 1d ago
Dev & Engineering

cmux Localization Rules & Audit Skill

Enforces localization rules and a verifiable audit workflow for every cmux user-facing string change, so no hardcoded English text ever lands.

★ 28k FS 56 Use with care 1d ago
Dev & Engineering

cmux Customization

A skill for safely customizing the cmux terminal: edit cmux., Dock config and Ghostty preferences to tailor actions, layouts, shortcuts and notifications without breaking existing config.

★ 28k FS 55 Use with care 1d ago
Productivity & Collaboration

cmux Markdown Viewer

Opens .md files in a formatted panel beside your cmux terminal with live reload, keeping plans, docs, and notes visible as they change.

★ 28k FS 52 Use with care 1d ago
Dev & Engineering

cmux Backend Development Rules Skill

Provides consistent backend TypeScript and Cloud VM development rules for the cmux repository, keeping Effect boundaries, Postgres migrations, and secret handling on track.

★ 28k FS 51 Use with care 1d ago
Dev & Engineering

cmux Diagnostics

Collects support-safe, read-only diagnostics for cmux so you can pinpoint failing hooks, notifications, session restore, and CLI control without leaking secrets.

★ 28k FS 51 Use with care 1d ago
Dev & Engineering

cmux Billing Runbook Skill

Gives AI coding agents an executable architecture map and ops runbook for Stripe billing, pricing, subscriptions, webhooks, and Pro entitlements in the cmux codebase.

★ 28k FS 50 Use with care 1d ago

Related skills