What does this skill do, and when should you use it?
This skill lives in TesterArmy's open-source e2e end-to-end testing framework repository (a monorepo bundling 10 skills). It interviews your codebase and generates a project-local verification skill at `.agents/skills/verify-<app>/` for one primary surface — web, iOS, Android, or an API — covering launch, doctor checks, driving, evidence, and bug bash. It also seeds a feature map with one explore charter per feature, then runs the generated skill end to end once to prove it. Output is written for the next agent, not a human, and assumes the checkout builds and starts.
- Interviews the repo: surface, run command, e2e state, side-effect stores, isolation, naming — asking the user only what code cannot answer.
- Generates
.agents/skills/verify-<app>/SKILL.mdwith concrete sections: Read first, Launch, Doctor, Drive, Evidence, Bug bash, Cleanup, Helpers — no placeholders. - Seeds a feature map (
features/README.mdplus one file per top feature) with routes, test files, proof criteria, and explore charters. - Proves the generated skill once:
npx e2e list+ smoke test, drives one mapped feature, bug bashes one charter vianpx e2e explore, then confirms evidence survives cleanup. - Hands over a summary of paths, mapped features, charters, and evidence locations, without committing on the user's behalf.
- A coding agent changed web app code and the project has no scripted way to prove UI behavior
- A repo with an existing Playwright or Cypress suite that wants to reuse its locators, seeds, and sign-in steps in an e2e-driven verification flow
- An app grew a feature the map lacks; rerun this skill to add it and mark untested features
- Teams wanting every change proven via real user paths, verified side effects, and .e2e/ evidence
- Running charted bug bashes against a local app, where a finding counts only if a bugbash test fails with ASSERTION_FAILED
- Native desktop apps: SKILL.md explicitly says e2e has no desktop engine yet — say so and stop
- Checkouts that don't build or start: the skill requires fixing or precisely reporting that first, since a skill written against a broken base teaches wrong steps
- Humans wanting readable docs: output is explicitly written cold for the next agent mid-task, not for people
How do you install this skill?
- The skill instructs running npx e2e init / npx skills add and registering an MCP server — external changes to the project; run on an isolated branch or backup first and review changes to .agents/skills/ and .mcp..
- e2e is pre-1.0 under active development; APIs and config may change between minor releases, and generated verify-<app> skills may need regeneration after CLI updates.
- Core docs and telemetry endpoints (e2e.tester.army, eu.i.posthog.com) are overseas with no declared mainland-China reachability; telemetry is on by default — disable with E2E_TELEMETRY_DISABLED=1 or npx e2e telemetry disable.
- The 'prove it once' step may spend real model calls and credentials; without a model, the handover must state the bug bash is unproven.
- Publisher TesterArmy is not verified by the FollowSkills curated registry and is treated as unknown; this is a static source review only — nothing was executed.
- Shell / CLI
- Network access
- Local filesystem
- MCP Server
Node.jse2e CLI (`npx e2e init`)e2e MCP server
Initialize e2e in the project first (this also installs the e2e skill and registers the MCP server); the skill lives at skills/create-verification-skill/:
npx e2e initOr install the whole skill collection via skills:
npx skills add tester-army/e2eHow do you use this skill?
Once installed, send your agent any of these to trigger it:
- /create-verification-skill
- make a verification skill for this repo
- This project has no scripted way to prove UI behavior — set up verification for it
Trigger it with /create-verification-skill or a request like "make a verification skill for this repo". The flow: (1) interview the repo — identify the surface, the dev/prod run command, existing e2e or Playwright/Cypress state, side-effect stores, isolation constraints; run npx e2e init if nothing is configured. (2) Generate .agents/skills/verify-<app>/SKILL.md with real facts in every section, and seed the feature map under features/. (3) Prove it end to end: doctor (npx e2e list + smoke test), drive one mapped feature, bug bash one charter:
npx e2e explore '<charter>' --config <bug-bash config> --agent <persona> --max-steps 3 --output .e2e/bugbash/<slug>(add --session <name> when a charter starts signed in; skip if no model and say so in the handover). (4) Clean up, confirm evidence exists, and hand over. Key options: --output .e2e/proof/<name> preserves evidence past the next run's cleanup; port 0 in app.url isolates each run; reuseExisting: true attaches to a running app.
What are this skill's strengths and limitations?
- Interviews the codebase rather than the user, mining Playwright/Cypress suites and configs so generated sections contain real facts, no placeholders
- Strict self-verification: doctor, drive one feature, bug bash one charter, confirm evidence survives cleanup — an unexecuted skill is a draft, not a deliverable
- Complete evidence model: .e2e/report., trace.md, --video, named screenshots, and cleanup rules that never destroy evidence
- Depends on the full e2e framework ecosystem (init, MCP server, CLI); the framework is pre-1.0 and APIs/config can change between minor releases
- The bug bash proof step needs a model; without one it is skipped and reported as unproven
- Covers web, iOS, Android, and API surfaces only; desktop is explicitly unsupported
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| Create Verification Skill (for e2e) this page | 54 · Use with care | ★ 8.7k | 1d ago | Apache-2.0 |
| e2e Playground Verification Skill | 59 · Recommended | ★ 8.7k | 1d ago | Apache-2.0 |
| e2e: Agentic End-to-End Testing | 52 · Use with care | ★ 8.7k | 1d ago | Apache-2.0 |
| e2e Verify — End-to-End Change Verification Skill | 64 · Recommended | ★ 8.7k | 1d ago | Apache-2.0 |
| Playwright Browser Automation Skill | 68 · Recommended | ★ 3.2k | 1mo ago | MIT |
The source names no direct competitor, but SKILL.md treats existing Playwright or Cypress suites as sources of locators, seeds, and sign-in steps — positioning this as a complement to existing scripted test frameworks rather than a replacement.
How did FollowSkills review this skill?
The skill is mostly read-only interviewing plus minimal writes (.agents/skills/, one .gitignore line), explicitly forbids embedding secrets and committing on the user's behalf, stops only processes it started, never kills by process name; the example references credentials by name only (type_secret / credentials.user('admin')). SECURITY.md discloses telemetry with opt-out. Deducted for: external side effects (npx e2e init, skills add, MCP registration) without explicit user confirmation, rewriting existing verify-<app> files governed only by a read-first rule, incomplete rollback guidance, and unverified publisher attribution resting on repo claims alone.
Instructions are self-consistent and cross-checked by the verify-playground example, with clear failure feedback (build failure, APP_ALREADY_RUNNING, LOCATOR_NOT_FOUND, skip bug bash and declare unproven without a model); Doctor-first and npx e2e list preflight are sound. Deducted for: the four-step proof cannot be executed in a static review, key paths of this specific skill are not reproduced here, and the e2e CLI is pre-1.0 with APIs that may change.
Triggers are clear (/create-verification-skill, 'make a verification skill', project lacking a scripted UI proof); inputs/outputs and non-fit boundaries are explicit — native desktop is declared unsupported with an instruction to stop. Deducted for: fuzzy boundary on 'identifiable user-facing features', no Chinese-language support, and a core workflow fully dependent on npm, Playwright, and overseas docs (e2e.tester.army) with no mainland-China reachability note.
Well-layered docs: the main SKILL.md links e2e topics instead of restating them, ships a complete auditable example (SKILL.md plus feature map), discloses known limitations (Gotchas, desktop gap, evidence-clearing rule), Apache-2.0 license, name/description match capability. Deducted for: no version/changelog in the skill itself, maintenance responsibility and update path not stated in the skill files, unverified third-party publisher.
Claims a directly handable verify-<app> skill and feature map, backed by a full example it says was generated and run; marginal value over hand-writing verification scripts is clearly positive. Deducted for: generated-output quality unverifiable statically, the example covers one playground app only, the 'prove it once' requirement depends on model credentials and Playwright, so completeness and comparative-benefit evidence remain limited; capped by the static ceiling.
Skill content is consistent with the in-repo example, README, SECURITY.md, and CI workflows (agent.yml, benchmark.yml covering framework test paths); referenced commands and test structures are auditable. Deducted for: no committed third-party execution evidence for this skill's own key path (generate, doctor, drive, bug bash, cleanup); 'generated and ran' is an author claim only; static confidence is low.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →