What does this skill do, and when should you use it?
verify-playground is one skill in the 10-skill collection bundled in the tester-army/e2e repository, dedicated to verifying the testbed playground web app (apps/testbed). It instructs an agent to launch the target app, run e2e CLI diagnostics (Doctor), drive pages live via an MCP session or scripted tests, and produce durable evidence (trace, video, screenshots, report.). It also supports a bug bash with skeptic and fuzzer personas, where a finding only counts as a bug when a test asserting it fails with ASSERTION_FAILED. The app has no database or env setup — todos live in localStorage and sessions in a cookie — and the server keeps no state.
- Reads e2e.config.ts (target web, tests tests/**/*.e2e.ts) and auto-starts the app via node app/server.mjs on PORT=4271, stopping it after the run
- Runs Doctor diagnostics: npx e2e list to validate the config, npx e2e run tests/smoke.e2e.ts to confirm the app opens, and tests/auth.setup.e2e.ts to save admin / admin-cookie sessions
- Drives pages live through the registered e2e mcp server (open_session, observe, locate, verbs) or writes scripted tests under tests/ (screen, agent.act)
- Writes .e2e/report. every run; supports --trace on, --video, app.screenshot(), and --output .e2e/proof/<name> for proofs that outlive the next run
- Runs a bug bash via e2e.bugbash.config.ts (untracked), adding tests/bugbash/** and skeptic/fuzzer personas, with charters in features/ in slug|target|agent|charter line format
- Cleans up: stops started apps and sessions, removes .e2e/bugbash/<slug>/ for rejected charters, keeps report-cited artifacts
- A maintainer who changed playground app code and needs full verification with screenshot/trace evidence before committing or announcing the fix
- An agent or developer asked to verify or screenshot a specific page, needing reproducible, durable run proof
- A QA engineer who wants a systematic bug bash of the playground using skeptic/fuzzer personas and the ready-made charters in features/
- A contributor debugging the test environment, using Doctor to distinguish config problems (list fails) from app problems (LOCATOR_NOT_FOUND)
- Anyone reproducing a bug bash finding locally by encoding it as tests/bugbash/<slug>.e2e.ts and confirming ASSERTION_FAILED
- Not for verifying your own business apps — this skill is scoped strictly to the apps/testbed playground; use the e2e framework itself for other projects
- Not for mobile targets — the skill's config target is web and everything is browser-based (localStorage, cookie)
- Not for bare API callers with no host shell — it requires shell, pnpm, the e2e CLI, and MCP session driving
How do you install this skill?
- This skill is bound to this repository's apps/testbed application and does not apply to other projects; do not follow its instructions blindly outside the repo.
- The skill depends on an external e2e CLI fetched via npx; the CLI sends anonymous telemetry by default (disable with E2E_TELEMETRY_DISABLED=1), and reachability of npm and the PostHog endpoint from mainland-China networks is unverified.
- The e2e framework is in active pre-1.0 development and APIs/config may change; the bug bash config is untracked and must be created by the user.
- This was a static source review with no commands or tests executed; conclusions rest on documentation consistency and in-repo evidence, with low confidence.
- Shell / CLI
- Network access
- Local filesystem
- MCP Server
Node.jse2e CLI (npm package)Playwright-based web enginepnpm (for manual app serving)
The skill ships inside the tester-army/e2e repo's skill collection (10 skills) at skills/create-verification-skill/references/example/verify-playground/. The README documents no standalone install command for this skill; you need the e2e toolchain for it to run:
npm install e2e
npx e2e initCopy the skill folder into your Agent Skills-compatible host's skills directory to enable it:
git clone https://github.com/tester-army/e2e
cp -r e2e/skills/create-verification-skill/references/example/verify-playground <your-host-skills-dir>/How do you use this skill?
Once installed, send your agent any of these to trigger it:
- I changed the playground nav link — run verification before I commit to confirm nothing broke
- Open the playground and screenshot the settings page for me
- Bug bash the playground's todo feature and compile the findings into a report
- The smoke test is failing — figure out whether it's my config or the app itself
The skill triggers when an agent needs to verify a playground change, is asked to verify or screenshot a page, or is asked to bug bash it. It starts by requiring you to read .agents/skills/e2e/SKILL.md. Typical flow:
1) Run Doctor first whenever anything looks off:
npx e2e list tests/smoke.e2e.ts
npx e2e run tests/smoke.e2e.ts2) For signed-in checks, run setup to save the admin / admin-cookie sessions:
npx e2e run tests/auth.setup.e2e.ts3) Drive live via the MCP server (open_session → observe → locate → verbs) or write scripted tests under tests/; always locate before writing any locator.
4) For proofs that must survive the next run, use an isolated output directory:
npx e2e run --trace on --output .e2e/proof/<name>5) Bug bash uses e2e.bugbash.config.ts; parallel explorers need pnpm app started first and reuseExisting: true, since the port is fixed. Note APP_ALREADY_RUNNING appears if you start the app by hand and then run tests.
What are this skill's strengths and limitations?
- Fully automatic app lifecycle: the runner starts, health-waits, and stops the target — no manual startup, no seeds or database
- Complete evidence chain: report., trace.md, per-attempt videos, named screenshots, plus --output isolation for durable proofs
- Strict bug confirmation: only a test encoding the finding failing with ASSERTION_FAILED counts; the explorer's claim alone does not
- Doctor gives crisp failure attribution: list failure = config/dependency, APP_ALREADY_RUNNING = port conflict, LOCATOR_NOT_FOUND = the app itself
- Scoped strictly to the apps/testbed playground — not directly reusable on other projects
- Port 4271 is fixed and not reused; parallel explorers require extra setup (pnpm app + reuseExisting)
- Depends on the untracked e2e.bugbash.config.ts and local features/ charters, which must exist in your environment
- No documented standalone install for this skill; host must support MCP and shell execution
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| e2e Playground Verification Skill this page | 59 · Recommended | ★ 8.7k | 1d ago | Apache-2.0 |
| Create Verification Skill (for e2e) | 54 · Use with care | ★ 8.7k | 1d ago | Apache-2.0 |
| e2e: Agentic End-to-End Testing | 52 · Use with care | ★ 8.7k | 1d ago | Apache-2.0 |
| Playwright Browser Automation Skill | 68 · Recommended | ★ 3.2k | 1mo ago | MIT |
| playwright-cli Browser Automation Skill | 56 · Use with care | ★ 11k | 13d ago | MIT |
The README positions e2e as a next-generation end-to-end testing framework that drives browsers via Playwright (@e2e-dev/web) and combines natural-language agent steps with locators/assertions in the same test, but the source names no specific competitor, so no comparison is made.
How did FollowSkills review this skill?
The skill describes least-privilege local operation: a fixed port, runner-managed app start/stop, explicit cleanup rules (no pkill node, per-session stops, gitignored proof artifacts), credentials via type_secret/credentials.user with screenshot withholding for password fills, and full data-flow disclosure (localStorage, cookie, stateless server). Deductions: actual execution depends on an external e2e CLI with telemetry (disclosed in README but on by default), no isolation or user-confirmation mechanism in the skill itself, and unverified publisher identity.
Documentation is self-consistent and thorough: a Doctor-first diagnostic, defined error codes (APP_ALREADY_RUNNING, LOCATOR_NOT_FOUND, ASSERTION_FAILED) with failure attribution, port-conflict handling, and bug-bash config caveats. Deductions: static review cannot execute; key paths (launch, test, bug bash) were not reproduced; the bugbash config is untracked and must be recreated by the reader; some assumptions rely on other topic docs.
Trigger conditions are explicit in the description (verify a playground change, screenshot a page, bug bash), boundaries are clear (scoped to apps/testbed, explicitly stating it only describes this app), and the feature map maps routes to tests. Deductions: extremely narrow scope (tied to this repo's testbed), no Chinese-language support statement, and mainland-China reachability of npm/npx dependency is unassessed.
Information architecture is well layered: SKILL.md defers to the underlying e2e skill first, and features/ expands per feature with sub-features, proof, charters, and gotchas; Apache-2.0 license is clear. Deductions: as an example skill it has no independent version, changelog, or maintenance-ownership statement; the untracked bugbash config shifts setup work to the user; known-limitations disclosure is incomplete.
The goal and value proposition are clear: verifiable evidence for playground changes (report, trace, screenshots, proof directory) and a principled distinction between explorer claims and test-confirmed bugs, exceeding manual verification in marginal value. Deductions: static review could not confirm representative outputs are directly usable; the value depends on the e2e framework behaving as claimed, which this review did not establish.
The repository contains real CI workflows (agent.yml, benchmark.yml) and committed test suites (mobile-benchmark tests exercising the testbed ecosystem), providing partial auditable evidence at the framework level. Deductions: the assessed skill's own key paths (playground verification flow) lack independent third-party execution evidence, and the correspondence between committed tests and the skill's claims was not verified item by item.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →