Dev & Engineering

Create Verification Skill (for e2e)

Generates a project-local verify-<app> skill so any coding agent can launch your app, drive it like a user, keep evidence, and bug bash it.

54/ 100
Use with care

Useful, but reliability, evidence or controls still have material gaps.

See how it was scored ↓
Works as-is in
Claude Code
Stars
★ 8.7k
Last updated
1d ago
License
Apache-2.0
e2e-testingverificationplaywrightmobile-testing
+3bug-bashtest-generationweb-testing

What does this skill do, and when should you use it?

This skill lives in TesterArmy's open-source e2e end-to-end testing framework repository (a monorepo bundling 10 skills). It interviews your codebase and generates a project-local verification skill at `.agents/skills/verify-<app>/` for one primary surface — web, iOS, Android, or an API — covering launch, doctor checks, driving, evidence, and bug bash. It also seeds a feature map with one explore charter per feature, then runs the generated skill end to end once to prove it. Output is written for the next agent, not a human, and assumes the checkout builds and starts.

  • Interviews the repo: surface, run command, e2e state, side-effect stores, isolation, naming — asking the user only what code cannot answer.
  • Generates .agents/skills/verify-<app>/SKILL.md with concrete sections: Read first, Launch, Doctor, Drive, Evidence, Bug bash, Cleanup, Helpers — no placeholders.
  • Seeds a feature map (features/README.md plus one file per top feature) with routes, test files, proof criteria, and explore charters.
  • Proves the generated skill once: npx e2e list + smoke test, drives one mapped feature, bug bashes one charter via npx e2e explore, then confirms evidence survives cleanup.
  • Hands over a summary of paths, mapped features, charters, and evidence locations, without committing on the user's behalf.
Good fit
  • A coding agent changed web app code and the project has no scripted way to prove UI behavior
  • A repo with an existing Playwright or Cypress suite that wants to reuse its locators, seeds, and sign-in steps in an e2e-driven verification flow
  • An app grew a feature the map lacks; rerun this skill to add it and mark untested features
  • Teams wanting every change proven via real user paths, verified side effects, and .e2e/ evidence
  • Running charted bug bashes against a local app, where a finding counts only if a bugbash test fails with ASSERTION_FAILED
Not a fit
  • Native desktop apps: SKILL.md explicitly says e2e has no desktop engine yet — say so and stop
  • Checkouts that don't build or start: the skill requires fixing or precisely reporting that first, since a skill written against a broken base teaches wrong steps
  • Humans wanting readable docs: output is explicitly written cold for the next agent mid-task, not for people

How do you install this skill?

Before you use it
  • The skill instructs running npx e2e init / npx skills add and registering an MCP server — external changes to the project; run on an isolated branch or backup first and review changes to .agents/skills/ and .mcp..
  • e2e is pre-1.0 under active development; APIs and config may change between minor releases, and generated verify-<app> skills may need regeneration after CLI updates.
  • Core docs and telemetry endpoints (e2e.tester.army, eu.i.posthog.com) are overseas with no declared mainland-China reachability; telemetry is on by default — disable with E2E_TELEMETRY_DISABLED=1 or npx e2e telemetry disable.
  • The 'prove it once' step may spend real model calls and credentials; without a model, the handover must state the bug bash is unproven.
  • Publisher TesterArmy is not verified by the FollowSkills curated registry and is treated as unknown; this is a static source review only — nothing was executed.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
  • MCP Server
Install first
  • Node.js
  • e2e CLI (`npx e2e init`)
  • e2e MCP server

Initialize e2e in the project first (this also installs the e2e skill and registers the MCP server); the skill lives at skills/create-verification-skill/:

npx e2e init

Or install the whole skill collection via skills:

npx skills add tester-army/e2e

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • /create-verification-skill
  • make a verification skill for this repo
  • This project has no scripted way to prove UI behavior — set up verification for it

Trigger it with /create-verification-skill or a request like "make a verification skill for this repo". The flow: (1) interview the repo — identify the surface, the dev/prod run command, existing e2e or Playwright/Cypress state, side-effect stores, isolation constraints; run npx e2e init if nothing is configured. (2) Generate .agents/skills/verify-<app>/SKILL.md with real facts in every section, and seed the feature map under features/. (3) Prove it end to end: doctor (npx e2e list + smoke test), drive one mapped feature, bug bash one charter:

npx e2e explore '<charter>' --config <bug-bash config> --agent <persona> --max-steps 3 --output .e2e/bugbash/<slug>

(add --session <name> when a charter starts signed in; skip if no model and say so in the handover). (4) Clean up, confirm evidence exists, and hand over. Key options: --output .e2e/proof/<name> preserves evidence past the next run's cleanup; port 0 in app.url isolates each run; reuseExisting: true attaches to a running app.

What are this skill's strengths and limitations?

Pros
  • Interviews the codebase rather than the user, mining Playwright/Cypress suites and configs so generated sections contain real facts, no placeholders
  • Strict self-verification: doctor, drive one feature, bug bash one charter, confirm evidence survives cleanup — an unexecuted skill is a draft, not a deliverable
  • Complete evidence model: .e2e/report., trace.md, --video, named screenshots, and cleanup rules that never destroy evidence
Limitations
  • Depends on the full e2e framework ecosystem (init, MCP server, CLI); the framework is pre-1.0 and APIs/config can change between minor releases
  • The bug bash proof step needs a model; without one it is skipped and reported as unproven
  • Covers web, iOS, Android, and API surfaces only; desktop is explicitly unsupported

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
Create Verification Skill (for e2e) this page 54 · Use with care ★ 8.7k 1d ago Apache-2.0
e2e Playground Verification Skill 59 · Recommended ★ 8.7k 1d ago Apache-2.0
e2e: Agentic End-to-End Testing 52 · Use with care ★ 8.7k 1d ago Apache-2.0
e2e Verify — End-to-End Change Verification Skill 64 · Recommended ★ 8.7k 1d ago Apache-2.0
Playwright Browser Automation Skill 68 · Recommended ★ 3.2k 1mo ago MIT

The source names no direct competitor, but SKILL.md treats existing Playwright or Cypress suites as sources of locators, seeds, and sign-in steps — positioning this as a complement to existing scripted test frameworks rather than a replacement.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Use with care
54/ 100 5-point scale 2.7 / 5
1Trust16 / 25 · 3.2/5

The skill is mostly read-only interviewing plus minimal writes (.agents/skills/, one .gitignore line), explicitly forbids embedding secrets and committing on the user's behalf, stops only processes it started, never kills by process name; the example references credentials by name only (type_secret / credentials.user('admin')). SECURITY.md discloses telemetry with opt-out. Deducted for: external side effects (npx e2e init, skills add, MCP registration) without explicit user confirmation, rewriting existing verify-<app> files governed only by a read-first rule, incomplete rollback guidance, and unverified publisher attribution resting on repo claims alone.

2Reliability9 / 20 · 2.3/5

Instructions are self-consistent and cross-checked by the verify-playground example, with clear failure feedback (build failure, APP_ALREADY_RUNNING, LOCATOR_NOT_FOUND, skip bug bash and declare unproven without a model); Doctor-first and npx e2e list preflight are sound. Deducted for: the four-step proof cannot be executed in a static review, key paths of this specific skill are not reproduced here, and the e2e CLI is pre-1.0 with APIs that may change.

3Adaptability9 / 15 · 3.0/5

Triggers are clear (/create-verification-skill, 'make a verification skill', project lacking a scripted UI proof); inputs/outputs and non-fit boundaries are explicit — native desktop is declared unsupported with an instruction to stop. Deducted for: fuzzy boundary on 'identifiable user-facing features', no Chinese-language support, and a core workflow fully dependent on npm, Playwright, and overseas docs (e2e.tester.army) with no mainland-China reachability note.

4Convention10 / 15 · 3.3/5

Well-layered docs: the main SKILL.md links e2e topics instead of restating them, ships a complete auditable example (SKILL.md plus feature map), discloses known limitations (Gotchas, desktop gap, evidence-clearing rule), Apache-2.0 license, name/description match capability. Deducted for: no version/changelog in the skill itself, maintenance responsibility and update path not stated in the skill files, unverified third-party publisher.

5Effectiveness6 / 15 · 2.0/5

Claims a directly handable verify-<app> skill and feature map, backed by a full example it says was generated and run; marginal value over hand-writing verification scripts is clearly positive. Deducted for: generated-output quality unverifiable statically, the example covers one playground app only, the 'prove it once' requirement depends on model credentials and Playwright, so completeness and comparative-benefit evidence remain limited; capped by the static ceiling.

6Verifiability4 / 10 · 2.0/5

Skill content is consistent with the in-repo example, README, SECURITY.md, and CI workflows (agent.yml, benchmark.yml covering framework test paths); referenced commands and test structures are auditable. Deducted for: no committed third-party execution evidence for this skill's own key path (generate, doctor, drive, bug bash, cleanup); 'generated and ran' is an author claim only; static confidence is low.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 449fa93670ee Review evidence[1][2][3][4][5][6][7][8][9][10][11][12][13][14]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

What does the generated skill require?
A checkout that builds and starts; a prior `npx e2e init` or `npx skills add tester-army/e2e` to install the e2e skill and register the MCP server; and a model configured for the bug bash proof step.
Where do the generated artifacts go?
`.agents/skills/verify-<app>/SKILL.md` (with YAML frontmatter, otherwise it never registers), the feature map under `features/`, and — if the project has `.claude/skills/` — a relative symlink `.claude/skills/verify-<app>`.
What if a verify-<app>/ already exists?
It is read before anything is written; its SKILL.md and feature files are kept, missing features are added, and existing files change only where the app contradicts them.
When does a bug bash finding count as a bug?
Only when `tests/bugbash/<slug>.e2e.ts` fails with ASSERTION_FAILED on the assertion that encodes it.

More skills from this repository

All from tester-army/e2e

Dev & Engineering

e2e Playground Verification Skill

Launch, drive, and prove the apps/testbed playground works — with traces, screenshots, and a bug bash — before you claim a change works.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

e2e: Agentic End-to-End Testing

Drive browser and mobile UI tests with natural-language agent goals, paired with exact locator assertions and a replay cache that keeps model costs down.

★ 8.7k FS 52 Use with care 1d ago
Dev & Engineering

e2e Verify — End-to-End Change Verification Skill

Prove every change with the real CLI against the testbed and benchmark apps — visible evidence, not "it compiles".

★ 8.7k FS 64 Recommended 1d ago
Dev & Engineering

babysit — PR Babysitting Skill

Drive an open pull request through conflicts, review bots, and CI until it is fully green with every thread handled, then hand it off labeled Ready for Human Review.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

Ship a PR (e2e repo PR delivery workflow)

Turn finished work in the tester-army/e2e repo into a PR a human can review without fighting CI or bots: checks, verification, a fresh-context self-review, the PR itself, then babysitting to the Ready for Human Review label.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

writing-pr: PR Title & Body Standards

A PR-writing standard that puts the shape of a change on the first screen: Conventional Commits titles, evidence-driven bodies, and a mandatory local-verification section.

★ 8.7k FS 58 Recommended 1d ago
Writing & Content

e2e Docs Authoring Skill

Applies a skimmer-first writing and shortening standard whenever you write, edit, or restructure guide pages on the e2e docs site.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

unbox-ai — AI Agent Trace Analysis CLI

Analyze AI agent trace files from the CLI without reading megabytes of raw JSON, and find out why an agent run was slow or expensive.

★ 8.7k FS 54 Use with care 1d ago
Writing & Content

Unslop

Edits AI tells out of any text and injects a human voice, so model-drafted writing stops reading like model-drafted writing. Its description says it must always apply.

★ 8.7k FS 49 Use with care 1d ago

Related skills