Dev & Engineering

Ship a PR (e2e repo PR delivery workflow)

Turn finished work in the tester-army/e2e repo into a PR a human can review without fighting CI or bots: checks, verification, a fresh-context self-review, the PR itself, then babysitting to the Ready for Human Review label.

59/ 100
Recommended

Generally reliable with disclosed limitations; trial as directed and keep a rollback path.

See how it was scored ↓
Works as-is in
Claude Code · Codex(Partial support)
Stars
★ 8.7k
Last updated
1d ago
License
Apache-2.0
pull-requestcode-reviewciconventional-commits
+3github-clidev-workflowchangesets

What does this skill do, and when should you use it?

ship-pr is a skill bundled in the e2e repository (tester-army/e2e) that codifies the path from "implementation done" to "a PR ready for human review." It requires running the same checks CI runs, proving the change with the verify skill, then having a reviewer with no memory of the coding conversation self-review the diff cold. It then commits with Conventional Commits, opens a ready (not draft) PR with verification evidence attached, and hands off to the babysit skill until CI is green, every bot thread is handled, and the Ready for Human Review label is applied. Its core premise: a human should only ever spend time on judgment, not on mechanical PR wrangling.

  • Runs the authoring-docs shortening pass on touched docs/**/*.mdx pages first, then pnpm check, pnpm test, testbed tests, and (when relevant) the web benchmark
  • Enforces repo rules pnpm check cannot see: .changeset/ entries for user-visible changes, docs updates in the same change, .d.ts and sdk-types.ts review for API changes, filters.yml for new build inputs
  • Proves the change with the verify skill on the touched surface (built CLI on the testbed or a benchmark, MCP server, scratch init project, or docs site), keeping output and video as PR evidence
  • Spawns a memory-less reviewer (subagent or new session) with references/review-prompt.md against the base branch, fixing every critical/important finding it agrees with
  • Branches off up-to-date origin/main, commits with Conventional Commits, writes the title/body with the writing-pr skill (including a ## Verified section), and opens a ready PR via gh pr create with screenshots and videos attached
  • Hands the PR to the babysit skill and runs it to green CI, handled threads, and the Ready for Human Review label, babysitting multi-PR stacks bottom-up
Good fit
  • A contributor in the e2e repo who has finished an implementation and is asked to open, create, ship, or submit a PR
  • A core developer fixing a bug who needs to attach main-vs-branch verification evidence in the PR body
  • A maintainer who changed the runner, web engine, or benchmarks and must also run the web benchmark
  • A package author making a public API change who must review the emitted .d.ts and update packages/e2e/tests/types/sdk-types.ts
  • A contributor delivering a multi-layer stack where each layer gets its own PR, each based on the one below, babysat until independently ready
  • Anyone who wants every doc anchor, changeset entry, and bot thread resolved before a human reviewer spends time
Not a fit
  • Other codebases — the skill hardcodes repo-specific conventions (pnpm check, @e2e-dev/testbed, filters.yml, the Ready for Human Review label) and would need substantial rewriting elsewhere
  • Quick personal changes that skip verification and self-review — the flow mandates evidence, fresh-context review, and babysitting with no skip path
  • Platforms without subagent support where the self-review fallback is unacceptable — the core design relies on a memory-less reviewer

How do you install this skill?

Before you use it
  • This is an internal workflow skill (internal: true) tied to tester-army/e2e's CI and directory layout; do not apply it to other repositories as-is.
  • The skill hard-depends on sibling skills (verify, babysit, writing-pr, authoring-docs) and gh 2.99+; a static review cannot confirm all references resolve at this revision.
  • The publisher TesterArmy is not verified by the FollowSkills curated registry, so attribution evidence is limited.
  • No commands were executed in this static review; all scores are inferred from source files and confidence is low.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
Install first
  • pnpm
  • GitHub CLI (gh 2.99+)
  • git
  • Node.js/pnpm workspace for the e2e repo

The source documents no standalone install commands. The skill lives at .dev/skills/ship-pr/SKILL.md inside the e2e repository and ships with the repo as an internal workflow skill (metadata marks it internal: true).

git clone https://github.com/tester-army/e2e
cd e2e

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • Implementation is done — open a PR for this fix to the e2e repo's main branch
  • Ship this change following the ship-pr flow, including verification evidence and self-review
  • This feature is finished and the next step is review — run the full PR flow up to the Ready for Human Review label
  • Submit this three-layer stack: one PR per layer, babysit each until independently ready

Trigger: whenever you are asked to open, create, ship, or submit a PR, or when implementation is done and the next step is review. The flow is a fixed five steps: 1) run the docs shortening pass first, then pnpm check / pnpm test / testbed tests plus repo-rule checks; 2) verify the change with the verify skill on the touched surface and keep the output and video; 3) spawn a memory-less reviewer with references/review-prompt.md (or cold-read the diff yourself and say so in the handoff); 4) branch off up-to-date origin/main, commit with Conventional Commits, write the PR with the writing-pr skill, open it ready with gh pr create --base main --title ... --body-file <file> --attach ...; 5) hand the PR to the babysit skill until green, threads handled, and the Ready for Human Review label. Any step that cannot run must be named and explained in the handoff — never skipped quietly. For stacks, babysit bottom-up; a layer is labeled only when it is ready on its own.

pnpm check
pnpm test
pnpm --filter @e2e-dev/testbed test
gh pr create --base main --title ... --body-file <file> --attach <each screenshot and video>

What are this skill's strengths and limitations?

Pros
  • Clear finish line: not opening a PR but reaching the Ready for Human Review label with every bot thread handled
  • Mandates verification evidence: the PR body must carry a ## Verified section with output, screenshots, and video, so humans can re-check quickly
  • Fresh-context self-review: an explicit 'the agent that wrote the change does not get to judge it' rule, with a documented fallback
  • Covers repo rules CI misses: changesets, doc sync, .d.ts/schema consistency are listed explicitly where pnpm check cannot catch them
  • Failure modes governed: steps that cannot run must be named and explained in the handoff, never silently skipped
Limitations
  • Deeply bound to one repo: pnpm scripts, @e2e-dev/testbed, filters.yml, AGENTS.md, and the Ready for Human Review label are all e2e-specific
  • Depends on a suite of sibling skills (authoring-docs, verify, writing-pr, babysit); using ship-pr alone is an incomplete flow
  • Metadata marks it internal: true — it is designed as a repo-internal workflow, not a general-purpose public skill
  • The ideal path (memory-less subagent review) depends on platform subagent capability; without it, you degrade to cold-reading the diff yourself
  • Relies on newer tooling such as gh 2.99+ --attach; older environments will fail

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
Ship a PR (e2e repo PR delivery workflow) this page 59 · Recommended ★ 8.7k 1d ago Apache-2.0
writing-pr: PR Title & Body Standards 58 · Recommended ★ 8.7k 1d ago Apache-2.0
RenderCV PR Reviewer 55 · Use with care ★ 18k 6mo ago MIT
Codex Pull Request Editor ✓ OpenAI · Official 50 · Use with care ★ 128k 3d ago Apache-2.0
Yeet GitHub PR Flow ✓ OpenAI · Official 41 · Not recommended ★ 28k 3mo ago —

The repository provides no direct comparison against other PR tooling, so no reliable comparison can be drawn from the source.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Recommended
59/ 100 5-point scale 3.0 / 5
1Trust17 / 25 · 3.4/5

The skill only describes in-repo PR workflow: running pnpm checks, opening a PR with gh, spawning a fresh-context reviewer. No destructive defaults, no credential access, no data exfiltration; the repo additionally carries a detailed SECURITY.md and telemetry disclosure, and CI explicitly refuses pull_request_target for secrets. Deductions: the skill itself does not declare least-privilege boundaries for gh calls and subagent side effects step by step, no rollback/abort path is written, and unverified publisher leaves attribution incomplete, so not full marks.

2Reliability10 / 20 · 2.5/5

Instructions are internally consistent, steps numbered, abnormal cases have an explicit reporting rule ('say which and why in the handoff'), and a self-review fallback exists when no reviewer can be spawned. Deductions: heavy dependence on four sibling skills not present in the provided source (verify, babysit, writing-pr, authoring-docs) plus gh 2.99+ and pnpm; a static read cannot confirm the references resolve, and concrete failure-feedback shapes are not shown, so capped at the static maximum of 10.

3Adaptability11 / 15 · 3.7/5

Frontmatter gives clear triggers (asked to open/create/ship/submit a PR), internal: true marks the non-public scope, and the audience (this repo's maintainers) is clear. Deductions: capability boundaries are tied to tester-army/e2e's CI structure (filters.yml, testbed) with no declared non-fit range for other repositories; Chinese-language support and mainland-China reachability are unaddressed (GitHub/npm generally reachable but unconfirmed), so 11.

4Convention11 / 15 · 3.7/5

Good layering: SKILL.md main flow plus references/review-prompt.md plus links to related skills; repo-level Apache-2.0 license, changesets versioning and well-commented CI governance. Deductions: the skill itself has no version or changelog, maintenance responsibility and update path for an internal skill are unstated, and key assumptions (changeset entry format, AGENTS.md contents) are not supplied inline, giving 11.

5Effectiveness6 / 15 · 2.0/5

The workflow shows clear marginal value (fresh-context self-review, verification evidence in the PR body, babysitting to a label) with an explicit output format (the Verified section). Deductions: a static read cannot verify the output is directly usable, and effectiveness depends on unseen sibling skills and real CI behavior, so capped at 7 and conservatively 6.

6Verifiability4 / 10 · 2.0/5

The repo contains auditable CI workflows (benchmark, agent) and extensive real test files, corroborating the check commands the skill cites (pnpm check/test, testbed). Deductions: the skill's own key paths (opening PRs, review, babysitting) have no committed execution evidence and no independent reproduction, so the static cap of 5 yields 4.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 449fa93670ee Review evidence[1][2][3][4][5][6][7][8][9][10][11]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Can I use this skill on my own repository?
Not as-is. It hardcodes the e2e repo's pnpm scripts, testbed, changeset conventions, and Ready for Human Review label, but the five-step flow (checks → verify → self-review → open → babysit) can serve as a template to port.
What if my environment has no subagent feature?
The skill has a built-in fallback: cold-read the whole diff yourself against the same review prompt, and state in the handoff that this is how the review was done.
What happens if I skip a step?
Skipping quietly is forbidden: any step that cannot run must be named and explained in the handoff, so the human reviewer knows which checks are missing.
Why must the PR be opened ready, not draft?
Because the babysit phase drives CI to green and handles bot threads on the PR; a draft state muddles the semantics for bots and reviewers.

More skills from this repository

All from tester-army/e2e

Dev & Engineering

writing-pr: PR Title & Body Standards

A PR-writing standard that puts the shape of a change on the first screen: Conventional Commits titles, evidence-driven bodies, and a mandatory local-verification section.

★ 8.7k FS 58 Recommended 1d ago
Dev & Engineering

babysit — PR Babysitting Skill

Drive an open pull request through conflicts, review bots, and CI until it is fully green with every thread handled, then hand it off labeled Ready for Human Review.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

e2e Verify — End-to-End Change Verification Skill

Prove every change with the real CLI against the testbed and benchmark apps — visible evidence, not "it compiles".

★ 8.7k FS 64 Recommended 1d ago
Dev & Engineering

unbox-ai — AI Agent Trace Analysis CLI

Analyze AI agent trace files from the CLI without reading megabytes of raw JSON, and find out why an agent run was slow or expensive.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

e2e Playground Verification Skill

Launch, drive, and prove the apps/testbed playground works — with traces, screenshots, and a bug bash — before you claim a change works.

★ 8.7k FS 59 Recommended 1d ago
Writing & Content

e2e Docs Authoring Skill

Applies a skimmer-first writing and shortening standard whenever you write, edit, or restructure guide pages on the e2e docs site.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

Create Verification Skill (for e2e)

Generates a project-local verify-<app> skill so any coding agent can launch your app, drive it like a user, keep evidence, and bug bash it.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

e2e: Agentic End-to-End Testing

Drive browser and mobile UI tests with natural-language agent goals, paired with exact locator assertions and a replay cache that keeps model costs down.

★ 8.7k FS 52 Use with care 1d ago
Writing & Content

Unslop

Edits AI tells out of any text and injects a human voice, so model-drafted writing stops reading like model-drafted writing. Its description says it must always apply.

★ 8.7k FS 49 Use with care 1d ago

Related skills