What does this skill do, and when should you use it?
ship-pr is a skill bundled in the e2e repository (tester-army/e2e) that codifies the path from "implementation done" to "a PR ready for human review." It requires running the same checks CI runs, proving the change with the verify skill, then having a reviewer with no memory of the coding conversation self-review the diff cold. It then commits with Conventional Commits, opens a ready (not draft) PR with verification evidence attached, and hands off to the babysit skill until CI is green, every bot thread is handled, and the Ready for Human Review label is applied. Its core premise: a human should only ever spend time on judgment, not on mechanical PR wrangling.
- Runs the authoring-docs shortening pass on touched docs/**/*.mdx pages first, then pnpm check, pnpm test, testbed tests, and (when relevant) the web benchmark
- Enforces repo rules pnpm check cannot see: .changeset/ entries for user-visible changes, docs updates in the same change, .d.ts and sdk-types.ts review for API changes, filters.yml for new build inputs
- Proves the change with the verify skill on the touched surface (built CLI on the testbed or a benchmark, MCP server, scratch init project, or docs site), keeping output and video as PR evidence
- Spawns a memory-less reviewer (subagent or new session) with references/review-prompt.md against the base branch, fixing every critical/important finding it agrees with
- Branches off up-to-date origin/main, commits with Conventional Commits, writes the title/body with the writing-pr skill (including a ## Verified section), and opens a ready PR via gh pr create with screenshots and videos attached
- Hands the PR to the babysit skill and runs it to green CI, handled threads, and the Ready for Human Review label, babysitting multi-PR stacks bottom-up
- A contributor in the e2e repo who has finished an implementation and is asked to open, create, ship, or submit a PR
- A core developer fixing a bug who needs to attach main-vs-branch verification evidence in the PR body
- A maintainer who changed the runner, web engine, or benchmarks and must also run the web benchmark
- A package author making a public API change who must review the emitted .d.ts and update packages/e2e/tests/types/sdk-types.ts
- A contributor delivering a multi-layer stack where each layer gets its own PR, each based on the one below, babysat until independently ready
- Anyone who wants every doc anchor, changeset entry, and bot thread resolved before a human reviewer spends time
- Other codebases — the skill hardcodes repo-specific conventions (pnpm check, @e2e-dev/testbed, filters.yml, the Ready for Human Review label) and would need substantial rewriting elsewhere
- Quick personal changes that skip verification and self-review — the flow mandates evidence, fresh-context review, and babysitting with no skip path
- Platforms without subagent support where the self-review fallback is unacceptable — the core design relies on a memory-less reviewer
How do you install this skill?
- This is an internal workflow skill (internal: true) tied to tester-army/e2e's CI and directory layout; do not apply it to other repositories as-is.
- The skill hard-depends on sibling skills (verify, babysit, writing-pr, authoring-docs) and gh 2.99+; a static review cannot confirm all references resolve at this revision.
- The publisher TesterArmy is not verified by the FollowSkills curated registry, so attribution evidence is limited.
- No commands were executed in this static review; all scores are inferred from source files and confidence is low.
- Shell / CLI
- Network access
- Local filesystem
pnpmGitHub CLI (gh 2.99+)gitNode.js/pnpm workspace for the e2e repo
The source documents no standalone install commands. The skill lives at .dev/skills/ship-pr/SKILL.md inside the e2e repository and ships with the repo as an internal workflow skill (metadata marks it internal: true).
git clone https://github.com/tester-army/e2e
cd e2eHow do you use this skill?
Once installed, send your agent any of these to trigger it:
- Implementation is done — open a PR for this fix to the e2e repo's main branch
- Ship this change following the ship-pr flow, including verification evidence and self-review
- This feature is finished and the next step is review — run the full PR flow up to the Ready for Human Review label
- Submit this three-layer stack: one PR per layer, babysit each until independently ready
Trigger: whenever you are asked to open, create, ship, or submit a PR, or when implementation is done and the next step is review. The flow is a fixed five steps: 1) run the docs shortening pass first, then pnpm check / pnpm test / testbed tests plus repo-rule checks; 2) verify the change with the verify skill on the touched surface and keep the output and video; 3) spawn a memory-less reviewer with references/review-prompt.md (or cold-read the diff yourself and say so in the handoff); 4) branch off up-to-date origin/main, commit with Conventional Commits, write the PR with the writing-pr skill, open it ready with gh pr create --base main --title ... --body-file <file> --attach ...; 5) hand the PR to the babysit skill until green, threads handled, and the Ready for Human Review label. Any step that cannot run must be named and explained in the handoff — never skipped quietly. For stacks, babysit bottom-up; a layer is labeled only when it is ready on its own.
pnpm check
pnpm test
pnpm --filter @e2e-dev/testbed test
gh pr create --base main --title ... --body-file <file> --attach <each screenshot and video>What are this skill's strengths and limitations?
- Clear finish line: not opening a PR but reaching the Ready for Human Review label with every bot thread handled
- Mandates verification evidence: the PR body must carry a ## Verified section with output, screenshots, and video, so humans can re-check quickly
- Fresh-context self-review: an explicit 'the agent that wrote the change does not get to judge it' rule, with a documented fallback
- Covers repo rules CI misses: changesets, doc sync, .d.ts/schema consistency are listed explicitly where pnpm check cannot catch them
- Failure modes governed: steps that cannot run must be named and explained in the handoff, never silently skipped
- Deeply bound to one repo: pnpm scripts, @e2e-dev/testbed, filters.yml, AGENTS.md, and the Ready for Human Review label are all e2e-specific
- Depends on a suite of sibling skills (authoring-docs, verify, writing-pr, babysit); using ship-pr alone is an incomplete flow
- Metadata marks it internal: true — it is designed as a repo-internal workflow, not a general-purpose public skill
- The ideal path (memory-less subagent review) depends on platform subagent capability; without it, you degrade to cold-reading the diff yourself
- Relies on newer tooling such as gh 2.99+ --attach; older environments will fail
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| Ship a PR (e2e repo PR delivery workflow) this page | 59 · Recommended | ★ 8.7k | 1d ago | Apache-2.0 |
| writing-pr: PR Title & Body Standards | 58 · Recommended | ★ 8.7k | 1d ago | Apache-2.0 |
| RenderCV PR Reviewer | 55 · Use with care | ★ 18k | 6mo ago | MIT |
| Codex Pull Request Editor ✓ OpenAI · Official | 50 · Use with care | ★ 128k | 3d ago | Apache-2.0 |
| Yeet GitHub PR Flow ✓ OpenAI · Official | 41 · Not recommended | ★ 28k | 3mo ago | — |
The repository provides no direct comparison against other PR tooling, so no reliable comparison can be drawn from the source.
How did FollowSkills review this skill?
The skill only describes in-repo PR workflow: running pnpm checks, opening a PR with gh, spawning a fresh-context reviewer. No destructive defaults, no credential access, no data exfiltration; the repo additionally carries a detailed SECURITY.md and telemetry disclosure, and CI explicitly refuses pull_request_target for secrets. Deductions: the skill itself does not declare least-privilege boundaries for gh calls and subagent side effects step by step, no rollback/abort path is written, and unverified publisher leaves attribution incomplete, so not full marks.
Instructions are internally consistent, steps numbered, abnormal cases have an explicit reporting rule ('say which and why in the handoff'), and a self-review fallback exists when no reviewer can be spawned. Deductions: heavy dependence on four sibling skills not present in the provided source (verify, babysit, writing-pr, authoring-docs) plus gh 2.99+ and pnpm; a static read cannot confirm the references resolve, and concrete failure-feedback shapes are not shown, so capped at the static maximum of 10.
Frontmatter gives clear triggers (asked to open/create/ship/submit a PR), internal: true marks the non-public scope, and the audience (this repo's maintainers) is clear. Deductions: capability boundaries are tied to tester-army/e2e's CI structure (filters.yml, testbed) with no declared non-fit range for other repositories; Chinese-language support and mainland-China reachability are unaddressed (GitHub/npm generally reachable but unconfirmed), so 11.
Good layering: SKILL.md main flow plus references/review-prompt.md plus links to related skills; repo-level Apache-2.0 license, changesets versioning and well-commented CI governance. Deductions: the skill itself has no version or changelog, maintenance responsibility and update path for an internal skill are unstated, and key assumptions (changeset entry format, AGENTS.md contents) are not supplied inline, giving 11.
The workflow shows clear marginal value (fresh-context self-review, verification evidence in the PR body, babysitting to a label) with an explicit output format (the Verified section). Deductions: a static read cannot verify the output is directly usable, and effectiveness depends on unseen sibling skills and real CI behavior, so capped at 7 and conservatively 6.
The repo contains auditable CI workflows (benchmark, agent) and extensive real test files, corroborating the check commands the skill cites (pnpm check/test, testbed). Deductions: the skill's own key paths (opening PRs, review, babysitting) have no committed execution evidence and no independent reproduction, so the static cap of 5 yields 4.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →