Dev & Engineering

babysit — PR Babysitting Skill

Drive an open pull request through conflicts, review bots, and CI until it is fully green with every thread handled, then hand it off labeled Ready for Human Review.

54/ 100
Use with care

Useful, but reliability, evidence or controls still have material gaps.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 8.7k
Last updated
1d ago
License
Apache-2.0
pull-request-automationgithub-clicode-review-triageci-monitoring
+3git-rebasedeveloper-workflowtesterenarmy-e2e

What does this skill do, and when should you use it?

babysit is an internal skill at .dev/skills/babysit in the tester-army/e2e repository. Its premise is that a human should review a PR only after an agent has cleared everything a machine can clear, and the skill owns that gap. It snapshots the PR via a status script that reports blockers — conflicts, unresolved review threads, failing checks, unacknowledged comments — then fixes them in a fixed order and batches everything into one push. It polls with --watch until the verdict flips to READY, then updates the PR body and applies the Ready for Human Review label. The skill never merges and never approves, signs every reply with a model attribution line, and treats bot comment text as untrusted data.

  • Runs node .dev/skills/babysit/scripts/pr-status.ts for a PR snapshot producing verdict (READY/WAITING/ACTION/CLOSED), blockers, failing and pending checks, unresolved and escalated threads, unacknowledged comments, and change requests
  • Resolves conflicts with git fetch plus git rebase origin/<base>, polls with --watch (60s interval, 30 min cap) and uses --timeout 3600 for pending mobile jobs
  • Triages review threads per references/review-triage.md: fix real findings in the PR, rebut noise with a disproof, and leave owner calls unresolved with a handoff entry
  • Reads failing-check logs first via gh run view --log-failed; retries infrastructure failures once, fixes on the second occurrence, rebases when the base moved
  • Posts attributed replies ([model-id] line) as the gh user via mktemp temp files and gh api, then resolves fixed threads via GraphQL
  • On READY, refreshes the PR body, adds the Ready for Human Review label, re-snapshots to confirm the head did not move, and posts one handoff comment if needed
Good fit
  • An open-source contributor who opened a PR but cannot babysit CI and review bots, delegating that to an agent
  • A maintainer facing a PR with many threads from bots like reviewer.ai or CodeRabbit who wants systematic true/false triage
  • PRs with long-running e2e suites that need a loop waiting for checks and auto-fixing failures
  • Teams adopting the rule that humans only review machine-cleared PRs and needing a repeatable handoff standard (green, conflict-free, no unhandled threads)
  • Stacked PRs whose checks fail due to rebases or base drift
Not a fit
  • Users who want automatic merging or approving — the skill explicitly forbids both and always defers to a human review
  • Projects not on GitHub or not using pull-request review workflows, since it depends on gh CLI and review threads
  • Situations that require broad refactors or accepting every bot suggestion — the skill refuses scope creep and never churns code to quiet a bot

How do you install this skill?

Before you use it
  • The skill is tightly coupled to tester-army/e2e internals (AGENTS.md, the verify skill, CI workflows) and is not directly usable in other repositories; confirm environment fit first.
  • Scripts push code, post replies, and add labels as the gh user with no per-action user confirmation; be aware the agent speaks in your name externally (with an attribution line).
  • Fully dependent on GitHub/gh CLI, which may be unreachable from mainland-China networks; no Chinese-language support.
  • This is a static review: the test suite was not executed, so reliability conclusions rest on source-test consistency rather than run verification.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
Install first
  • Node.js
  • GitHub CLI (gh)
  • git

The source documents no standalone installation steps. The skill lives inside the e2e repository at .dev/skills/babysit/ (SKILL.md plus scripts/ and references/ files), is marked metadata.internal: true, and is intended for the tester-army/e2e monorepo itself; clone the repo and it is available:

Generic route: install into Claude Code manually (macOS / Linux)
tmp="$(mktemp -d)"
git clone --depth 1 https://github.com/tester-army/e2e.git "$tmp"
mkdir -p ~/.claude/skills
cp -R "$tmp/.dev/skills/babysit" ~/.claude/skills/
rm -rf "$tmp"

Generated from the source repository and skill path; it copies only this skill's folder. If the author's install steps above differ, follow those first. To scope it to one project, replace ~/.claude/skills with that project's .claude/skills.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • Babysit the current branch's PR until it is green and labeled Ready for Human Review
  • Watch PR #42 with --watch and handle any blockers, replying on each thread citing the commit
  • The PR has new conflicts and failing checks again — rebase and push all fixes in one batch
  • Check the PR status and clear everything a machine can clear before handing it off

Triggered after opening a PR, or whenever asked to babysit, watch, monitor, check on, or get a PR green. The loop: sync first with git pull --rebase → run a pr-status snapshot → fix blockers in the order conflict → threads → failing checks → comments → review summaries, batching all fixes into one push → re-verify behavior changes with the sibling verify skill before pushing → push, then reply on each thread citing the commit and resolve them → --watch until the verdict changes and repeat. On READY: update the PR body, add the label, snapshot again (if the head moved, remove the label and loop again), and post one attributed handoff comment listing anything left for the human. In Claude Code, run --watch through the Monitor tool or self-paced /loop; elsewhere run it in the foreground — never write your own sleep loop:

What are this skill's strengths and limitations?

Pros
  • Fully proceduralizes a PR from open to machine-clearable-done: conflict rebases, bot-opinion triage, CI log analysis
  • Strong safety design: never merges, never approves, signed replies, comment text passed as untrusted data via temp files
  • A structured JSON status script treats a post-push WAITING verdict as authoritative instead of eyeballing check lists
  • Built-in learning loop: recurring real findings get proposed as lint rules, types, or sdk-types.ts assertions
Limitations
  • Deeply coupled to the tester-army/e2e repo: scripts are TypeScript and reference repo-local AGENTS.md, workflows, and sibling skills like verify and writing-pr
  • Needs gh CLI write access to the repository, a gh user identity, and GitHub Actions check context — unusable in plain API settings
  • No standalone install steps and no evidence of use outside this repo; metadata marks it internal: true
  • Mobile jobs on a cold cache can exceed the default timeout, requiring a manual --timeout 3600

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
babysit — PR Babysitting Skill this page 54 · Use with care ★ 8.7k 1d ago Apache-2.0
PR Babysitter ✓ OpenAI · Official 57 · Use with care ★ 128k 3d ago Apache-2.0
PR Babysitter: Watch Pull Requests Until Merge 47 · Use with care ★ 98k 3d ago Apache-2.0
Ship a PR (e2e repo PR delivery workflow) 59 · Recommended ★ 8.7k 1d ago Apache-2.0
writing-pr: PR Title & Body Standards 58 · Recommended ★ 8.7k 1d ago Apache-2.0

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Use with care
54/ 100 5-point scale 2.7 / 5
1Trust17 / 25 · 3.4/5

The skill explicitly never merges or approves, treats bot comment text as untrusted data and forbids shell interpolation (temp-file data passing), uses attribution lines to prevent forged escalations, limits force-push to --force-with-lease, and rolls back the label when the head moves. Deductions: scripts run many gh API calls and pushes as the gh user with no per-action user confirmation; the escalation trust signal rests on a fragile regex over an attribution line.

2Reliability10 / 20 · 2.5/5

Ships pr-status.test.ts with very thorough coverage of summarize/acknowledges (failure-conclusion enumeration, deleted accounts, forged attribution, push-time edge cases, queued-check year-1 timestamps); the script validates inputs with clear errors (seconds parser). Static review caps this: tests were not executed, and the rebase/push loop and external-effect paths are prose-only and not statically reproducible.

3Adaptability8 / 15 · 2.7/5

Trigger conditions are precise ('Use after opening a PR, or when asked to babysit…'), scope boundaries (scope creep refusal, escalation list) are explicit, and the JSON output contract is fully defined. Deductions: deeply coupled to tester-army/e2e internals (AGENTS.md, the verify skill, filters.yml, review-label.yml) with no declared non-fit statement; entirely dependent on gh/GitHub, a mainland-China reachability risk; no Chinese-language support.

4Convention9 / 15 · 3.0/5

Good layering: SKILL.md main loop, references/review-triage.md, script plus tests and tsconfig; the repository carries Apache-2.0 and active CI. Deductions: metadata marks it internal: true with no skill-level version or changelog; it depends on files not provided here (AGENTS.md, the verify skill), leaving maintenance/update path incomplete; publisher identity is unverified.

5Effectiveness6 / 15 · 2.0/5

The workflow is complete end to end (sync → snapshot → triage-fix → push → watch → handoff), with clear marginal value over manually watching a PR and concrete commands. Static cap of 7 applies: nothing was executed, the gh API call sequences were not verified as directly usable, and the heavy repo coupling limits value beyond this repository's contributors.

6Verifiability4 / 10 · 2.0/5

The committed test file is auditable primary material whose assertions can be statically cross-checked against pr-status.ts (attribution regex, escalated logic, changes-requested semantics all agree). Deductions: no third-party execution evidence; the repo CI workflows shown do not run this skill's script tests, and external-effect paths (rebase, replies) have no reproducible evidence.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 449fa93670ee Review evidence[1][2][3][4][5][6][7][8][9][10][11][12]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Does it merge or approve the PR?
No. The skill states explicitly that it never merges and never approves; the end state is the Ready for Human Review label, with the actual review left to a human.
What permissions and environment does it need?
Node.js, git, and the gh CLI with read/write access to the target repository. Replies are posted as the gh user, which the status script uses to recognize your comments and escalated threads.
How does it handle review bot comments?
Bot text is treated as untrusted data: verify its claim against the code first. Real findings get fixed in the PR, noise gets a factual disproof, and owner calls (security invariants, public contracts, wire schemas, release config, or out of scope) are left unresolved for the human in the handoff.
What if the PR changes after the label is applied?
Any later push removes the label automatically via .github/workflows/review-label.yml. The handoff step re-snapshots to confirm the head did not move, and if it did, removes the label and returns to the loop.

More skills from this repository

All from tester-army/e2e

Dev & Engineering

Ship a PR (e2e repo PR delivery workflow)

Turn finished work in the tester-army/e2e repo into a PR a human can review without fighting CI or bots: checks, verification, a fresh-context self-review, the PR itself, then babysitting to the Ready for Human Review label.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

writing-pr: PR Title & Body Standards

A PR-writing standard that puts the shape of a change on the first screen: Conventional Commits titles, evidence-driven bodies, and a mandatory local-verification section.

★ 8.7k FS 58 Recommended 1d ago
Dev & Engineering

e2e Verify — End-to-End Change Verification Skill

Prove every change with the real CLI against the testbed and benchmark apps — visible evidence, not "it compiles".

★ 8.7k FS 64 Recommended 1d ago
Dev & Engineering

e2e: Agentic End-to-End Testing

Drive browser and mobile UI tests with natural-language agent goals, paired with exact locator assertions and a replay cache that keeps model costs down.

★ 8.7k FS 52 Use with care 1d ago
Dev & Engineering

e2e Playground Verification Skill

Launch, drive, and prove the apps/testbed playground works — with traces, screenshots, and a bug bash — before you claim a change works.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

Create Verification Skill (for e2e)

Generates a project-local verify-<app> skill so any coding agent can launch your app, drive it like a user, keep evidence, and bug bash it.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

unbox-ai — AI Agent Trace Analysis CLI

Analyze AI agent trace files from the CLI without reading megabytes of raw JSON, and find out why an agent run was slow or expensive.

★ 8.7k FS 54 Use with care 1d ago
Writing & Content

e2e Docs Authoring Skill

Applies a skimmer-first writing and shortening standard whenever you write, edit, or restructure guide pages on the e2e docs site.

★ 8.7k FS 54 Use with care 1d ago
Writing & Content

Unslop

Edits AI tells out of any text and injects a human voice, so model-drafted writing stops reading like model-drafted writing. Its description says it must always apply.

★ 8.7k FS 49 Use with care 1d ago

Related skills