Dev & Engineering

e2e Verify — End-to-End Change Verification Skill

Prove every change with the real CLI against the testbed and benchmark apps — visible evidence, not "it compiles".

64/ 100
Recommended

Generally reliable with disclosed limitations; trial as directed and keep a rollback path.

See how it was scored ↓
Works as-is in
Codex · Claude Code · ChatGPT(Partial support) · Claude.ai(Partial support)
Stars
★ 8.7k
Last updated
1d ago
License
Apache-2.0
e2e-testingplaywrightcli-verificationmcp-server
+3bug-reproductionregression-testingmobile-testing

What does this skill do, and when should you use it?

verify is an internal skill of the tester-army/e2e repository (a next-generation e2e testing framework for web and mobile apps), located at .dev/skills/verify/SKILL.md. Its core principle: passing unit tests is not evidence that the CLI prints the right line, a locator finds a node on a real page, or an agent step replays. Any change to the runner, engine, CLI, reporters, MCP server, or docs must be verified with the built CLI against the testbed and benchmark applications. The skill specifies how to pick the verification surface (runner, reporters, web/mobile engines, agent, MCP, packaging, docs each have a distinct path), how to collect evidence (terminal output, video, stills, main-vs-branch comparisons), and how to clean up afterward. It targets contributors iterating inside this repository and opening PRs — not end users looking for a general testing tool.

  • Picks the verification path by change surface: runner/CLI changes run the testbed suite, reporter changes also run test:stress, engine changes run real Chromium benchmark scenarios, mobile changes run the mobile benchmark on a simulator or emulator
  • Rebuilds dist after every change (pnpm build) and runs deterministic tests through the built CLI directly, reading terminal output and the reports/traces under .e2e/
  • Drives pages like a coding agent via the e2e mcp server (open_session, observe, locate, screenshot, action verbs) and records video of what it drove
  • For bug fixes, runs the regression test on main first to capture the failure, then builds main in a git worktree and runs the same command on both sides
  • Records every attempt with --video / --headed (webm/mp4), cuts stills with ffmpeg, and uploads media to the PR via gh pr create|edit --attach
  • Checks side effects: artifacts land where the docs say, filled secrets stay out of every file under .e2e/ (grep -r), and MCP sessions and /tmp worktrees are cleaned up
Good fit
  • A contributor to the e2e repo changes a CLI flag or locator logic and must verify real behavior end to end on the testbed app before opening a PR
  • A reporter bug is fixed (list//markdown/junit) and output must be verified including hostile titles, in both printed and written form
  • Agent logic (src/agent/) or prompts are changed: run the tests-agent/ replay first, then --no-cache --ai-trace to confirm the live agent still reaches the goal
  • A user-reported bug needs reproducing: trigger the failure path on a real page and read the exact message a user gets
  • A docs page changed: run pnpm docs:dev, inspect the rendered page, and screenshot it as PR evidence
  • The @e2e-dev/mobile engine changed: verify on iOS/Android devices one run per device, following apps/mobile-benchmark/README.md
Not a fit
  • Not for end users of the e2e framework: this is a repo-internal contributor skill (frontmatter marks metadata.internal: true) that verifies this repository's own code, not a testing skill for your own projects
  • Not for environments that cannot meet the heavy prerequisites: pnpm builds, Playwright Chromium, simulator/emulator devices for mobile, and an AI_GATEWAY_API_KEY for live agent verification
  • Not for changes that cannot be run against a real app: the skill explicitly rejects concluding from unit tests or "it compiles" alone — pure static code review belongs elsewhere

How do you install this skill?

Before you use it
  • Static review with low confidence: nothing was executed; key files referenced by the skill (AGENTS.md, .mcp., benchmark test directories) were not provided in the evidence.
  • The skill is marked internal:true and targets this repository's contributors; it is not suitable for distribution or invocation as a general-purpose external skill.
  • Agentic verification depends on AI_GATEWAY_API_KEY and the Vercel AI Gateway (overseas service); live agent steps may be unreachable from mainland-China networks, though recorded replay is unaffected.
  • Telemetry is on by default (PostHog); despite full disclosure and opt-out flags, confirm E2E_TELEMETRY_DISABLED before integrating.
  • e2e is in active pre-1.0 development; APIs and config can change between minor releases, so skill instructions may drift.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
  • MCP Server
Install first
  • Node.js
  • pnpm
  • Playwright (chromium via playwright-core)
  • GitHub CLI (gh 2.99+ for PR media upload)
  • ffmpeg (for stills)
  • AI_GATEWAY_API_KEY for live agent runs

The skill ships inside the e2e repository as one of 10 skills in the monorepo, at .dev/skills/verify/SKILL.md. There is no separate install command — it is in place once you clone the repo:

git clone https://github.com/tester-army/e2e

Repository-level setup (the Setup block from SKILL.md):

pnpm install --frozen-lockfile
pnpm --filter @e2e-dev/web exec playwright-core install chromium
pnpm build

For MCP-based verification in Claude Code, .mcp. at the repo root registers the e2e mcp server. Install steps for other hosts (e.g. Codex) are not documented in the source.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • I changed the --reporter logic — run the testbed plus test:stress and show me the real list and markdown output
  • Reproduce this bug report: the locator can't find the node on a real page; run the regression test on main first, then compare on my branch
  • Verify my agent change: run the web benchmark's tests-agent/ replay, then go live with --no-cache --ai-trace and record a video
  • Do the full verification before I open this PR — screenshot the docs page and attach the video and still to the PR

Triggered while iterating and before opening a PR, when asked to run, test, check, or screenshot something, or to reproduce a bug report. Flow: change code → pnpm build (a stale dist is the most common false result) → pick the run from the surface table (most runs call the CLI directly from the app directory, e.g. cd apps/testbed && node node_modules/e2e/dist/cli/bin.js run tests/todos.e2e.ts) → collect evidence. The full gate, as CI runs it:

pnpm test:testbed
pnpm test:web-benchmark

Video recording and stills:

cd apps/web-benchmark && node node_modules/e2e/dist/cli/bin.js run --video --headed tests/<scenario>.e2e.ts
ffmpeg -ss <seconds> -i <video> -frames:v 1 after.png

Uploading to the PR:

gh pr create --body-file body.md --attach ./after.png --attach ./video.webm

Options to know: live agent runs need AI_GATEWAY_API_KEY; benchmark agent suites replay committed recordings by default and only go live with --no-cache --ai-trace when the change is to the agent itself; the MCP server runs the dist it started with, so restart it after a rebuild; APP_ALREADY_RUNNING means another checkout holds the app port — never stop a process you did not start.

What are this skill's strengths and limitations?

Pros
  • Holds verification to the standard of what a user actually sees, eliminating "it compiles" false evidence
  • A clear surface-to-run mapping covering runner, reporters, web/mobile engines, agent, MCP, packaging, and docs end to end
  • Demands visible PR evidence (terminal output, video, stills) and supports main-vs-branch regression comparisons
  • Takes failure paths and side effects seriously: trigger error codes, confirm secrets stay out of .e2e/, check artifact locations
Limitations
  • Only applies to development inside the e2e repository; other projects cannot reuse it directly
  • High setup cost: pnpm builds, Playwright Chromium, and simulator/emulator devices for mobile
  • Live agent verification requires AI_GATEWAY_API_KEY and incurs model-call costs
  • The source documents no steps for installing this skill on hosts other than its home repo and Claude Code's MCP setup

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
e2e Verify — End-to-End Change Verification Skill this page 64 · Recommended ★ 8.7k 1d ago Apache-2.0
Create Verification Skill (for e2e) 54 · Use with care ★ 8.7k 1d ago Apache-2.0
e2e: Agentic End-to-End Testing 52 · Use with care ★ 8.7k 1d ago Apache-2.0
e2e Playground Verification Skill 59 · Recommended ★ 8.7k 1d ago Apache-2.0
playwright-cli Browser Automation Skill 56 · Use with care ★ 11k 13d ago MIT

The skill cross-references two sibling skills in the same repo: skills/e2e (the consumer-facing skill for writing tests) and .claude/skills/unbox-ai (for reading agent traces). They complement each other: verify covers contributor-side verification, e2e covers user-side test authoring, unbox-ai covers trace analysis. The source does not name external framework comparisons.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Recommended
64/ 100 5-point scale 3.2 / 5
1Trust18 / 25 · 3.6/5

The skill itself is a least-privilege developer verification workflow: no sensitive permissions requested, explicitly forbids stopping processes it did not start, requires cleanup (worktrees, MCP sessions, apps), and proactively demands grep-verifying that filled secrets never land under .e2e/. The repo's SECURITY.md shows a clear trust model (code-trust vs untrusted model input, secrets never reaching model input, opt-out telemetry with every field disclosed). Deducted: several referenced files (AGENTS.md, fixtures, .mcp.) were not provided so claims cannot be statically confirmed; publisher is unverified.

2Reliability12 / 20 · 3.0/5

Instructions are self-consistent and engineered: the most common false result (stale dist) is named with a rebuild rule; agent recording replay cache, --strict-cache defenses, real CI workflows (benchmark.yml, agent.yml) plus committed deterministic test suites (testbed, web/mobile benchmarks) cover the skill's key paths, meeting the evidence bar to exceed 10. Deducted: multiple referenced files (AGENTS.md, skills/e2e, unbox-ai, benchmark test directories) are absent from the evidence, abnormal-input failure feedback is only partially confirmable, and nothing was executed.

3Adaptability10 / 15 · 3.3/5

Audience is clear (contributors of this repo), trigger description is detailed (when changing runner/engine/CLI/reporter/MCP/docs, while iterating and before PRs), with a surface-selection table giving precise semantic triggers. Deducted: as an internal:true contributor skill its scope is inherently limited to this repo's ecosystem; boundaries and environment assumptions (mobile simulators, AI_GATEWAY_API_KEY, gh version) depend on files not provided; no Chinese-language support declared, and the agent depends on an overseas model gateway (Vercel AI Gateway) with unassessed mainland-China reachability.

4Convention12 / 15 · 4.0/5

Documentation layers well: setup → pick the surface → drive like a user → evidence → cleanup, with good progressive disclosure; the repo has Apache-2.0, changesets versioning, explicit CI gates and a security policy, and clear maintenance responsibility (contributors, Discord). Deducted: the skill itself has no version/changelog, key referenced companions (AGENTS.md, consumer skill) are missing from evidence, leaving hidden assumptions; no committed update path for the internal skill as an external artifact.

5Effectiveness7 / 15 · 2.3/5

The skill solves a real and often skipped need ("it compiles" is not verification): run built artifacts as a user would, exercise failure paths, and produce directly usable evidence (terminal output, video, stills); marginal value clearly exceeds ad-hoc manual verification, and CI plus benchmark suites show these paths actually run. Deducted: static review cannot confirm representative outputs are directly usable, so the value claim is only partially verified.

6Verifiability5 / 10 · 2.5/5

Multiple auditable primary materials exist: real CI workflows (benchmark/agent) and committed deterministic/agentic test suites cover the skill's key paths, recordings are committed per PR, and security/telemetry fields are exhaustively disclosed. Deducted: a static read is not independent reproduction; several referenced companion files and recordings are absent from the evidence, so cross-source corroboration is incomplete — capped at 5.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 449fa93670ee Review evidence[1][2][3][4][5][6][7][8][9][10]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Can I use this skill to test my own project?
Not directly. It is an internal skill of the e2e repository (metadata.internal: true), bound to this repo's testbed, benchmark apps, and pnpm workspace. To write e2e tests for your own project, use the published e2e CLI and npm packages.
Why must I run pnpm build after every change?
The skill states it plainly: a stale dist is the most common false result. The testbed and benchmarks consume the built packages, so no run is trustworthy until dist is rebuilt — and the same applies after switching branches; the MCP server also needs a restart to pick up a new dist.
When do model calls (and their costs) happen?
By default the benchmark agent suites replay committed recordings with no model calls; a step calls the model only when it has no recording. Live runs with --no-cache --ai-trace are reserved for changes to the agent itself.
What if something cannot be verified?
The skill requires you to say so and name the blocker — a confident claim without evidence is worse than "inconclusive".

More skills from this repository

All from tester-army/e2e

Dev & Engineering

Create Verification Skill (for e2e)

Generates a project-local verify-<app> skill so any coding agent can launch your app, drive it like a user, keep evidence, and bug bash it.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

e2e: Agentic End-to-End Testing

Drive browser and mobile UI tests with natural-language agent goals, paired with exact locator assertions and a replay cache that keeps model costs down.

★ 8.7k FS 52 Use with care 1d ago
Dev & Engineering

e2e Playground Verification Skill

Launch, drive, and prove the apps/testbed playground works — with traces, screenshots, and a bug bash — before you claim a change works.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

babysit — PR Babysitting Skill

Drive an open pull request through conflicts, review bots, and CI until it is fully green with every thread handled, then hand it off labeled Ready for Human Review.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

Ship a PR (e2e repo PR delivery workflow)

Turn finished work in the tester-army/e2e repo into a PR a human can review without fighting CI or bots: checks, verification, a fresh-context self-review, the PR itself, then babysitting to the Ready for Human Review label.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

writing-pr: PR Title & Body Standards

A PR-writing standard that puts the shape of a change on the first screen: Conventional Commits titles, evidence-driven bodies, and a mandatory local-verification section.

★ 8.7k FS 58 Recommended 1d ago
Dev & Engineering

unbox-ai — AI Agent Trace Analysis CLI

Analyze AI agent trace files from the CLI without reading megabytes of raw JSON, and find out why an agent run was slow or expensive.

★ 8.7k FS 54 Use with care 1d ago
Writing & Content

e2e Docs Authoring Skill

Applies a skimmer-first writing and shortening standard whenever you write, edit, or restructure guide pages on the e2e docs site.

★ 8.7k FS 54 Use with care 1d ago
Writing & Content

Unslop

Edits AI tells out of any text and injects a human voice, so model-drafted writing stops reading like model-drafted writing. Its description says it must always apply.

★ 8.7k FS 49 Use with care 1d ago

Related skills