Dev & Engineering

e2e Playground Verification Skill

Launch, drive, and prove the apps/testbed playground works — with traces, screenshots, and a bug bash — before you claim a change works.

59/ 100
Recommended

Generally reliable with disclosed limitations; trial as directed and keep a rollback path.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 8.7k
Last updated
1d ago
License
Apache-2.0
e2e-testingplaywrightweb-testingbug-bash
+3verificationscreenshot-evidencemcp

What does this skill do, and when should you use it?

verify-playground is one skill in the 10-skill collection bundled in the tester-army/e2e repository, dedicated to verifying the testbed playground web app (apps/testbed). It instructs an agent to launch the target app, run e2e CLI diagnostics (Doctor), drive pages live via an MCP session or scripted tests, and produce durable evidence (trace, video, screenshots, report.). It also supports a bug bash with skeptic and fuzzer personas, where a finding only counts as a bug when a test asserting it fails with ASSERTION_FAILED. The app has no database or env setup — todos live in localStorage and sessions in a cookie — and the server keeps no state.

  • Reads e2e.config.ts (target web, tests tests/**/*.e2e.ts) and auto-starts the app via node app/server.mjs on PORT=4271, stopping it after the run
  • Runs Doctor diagnostics: npx e2e list to validate the config, npx e2e run tests/smoke.e2e.ts to confirm the app opens, and tests/auth.setup.e2e.ts to save admin / admin-cookie sessions
  • Drives pages live through the registered e2e mcp server (open_session, observe, locate, verbs) or writes scripted tests under tests/ (screen, agent.act)
  • Writes .e2e/report. every run; supports --trace on, --video, app.screenshot(), and --output .e2e/proof/<name> for proofs that outlive the next run
  • Runs a bug bash via e2e.bugbash.config.ts (untracked), adding tests/bugbash/** and skeptic/fuzzer personas, with charters in features/ in slug|target|agent|charter line format
  • Cleans up: stops started apps and sessions, removes .e2e/bugbash/<slug>/ for rejected charters, keeps report-cited artifacts
Good fit
  • A maintainer who changed playground app code and needs full verification with screenshot/trace evidence before committing or announcing the fix
  • An agent or developer asked to verify or screenshot a specific page, needing reproducible, durable run proof
  • A QA engineer who wants a systematic bug bash of the playground using skeptic/fuzzer personas and the ready-made charters in features/
  • A contributor debugging the test environment, using Doctor to distinguish config problems (list fails) from app problems (LOCATOR_NOT_FOUND)
  • Anyone reproducing a bug bash finding locally by encoding it as tests/bugbash/<slug>.e2e.ts and confirming ASSERTION_FAILED
Not a fit
  • Not for verifying your own business apps — this skill is scoped strictly to the apps/testbed playground; use the e2e framework itself for other projects
  • Not for mobile targets — the skill's config target is web and everything is browser-based (localStorage, cookie)
  • Not for bare API callers with no host shell — it requires shell, pnpm, the e2e CLI, and MCP session driving

How do you install this skill?

Before you use it
  • This skill is bound to this repository's apps/testbed application and does not apply to other projects; do not follow its instructions blindly outside the repo.
  • The skill depends on an external e2e CLI fetched via npx; the CLI sends anonymous telemetry by default (disable with E2E_TELEMETRY_DISABLED=1), and reachability of npm and the PostHog endpoint from mainland-China networks is unverified.
  • The e2e framework is in active pre-1.0 development and APIs/config may change; the bug bash config is untracked and must be created by the user.
  • This was a static source review with no commands or tests executed; conclusions rest on documentation consistency and in-repo evidence, with low confidence.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
  • MCP Server
Install first
  • Node.js
  • e2e CLI (npm package)
  • Playwright-based web engine
  • pnpm (for manual app serving)

The skill ships inside the tester-army/e2e repo's skill collection (10 skills) at skills/create-verification-skill/references/example/verify-playground/. The README documents no standalone install command for this skill; you need the e2e toolchain for it to run:

npm install e2e
npx e2e init

Copy the skill folder into your Agent Skills-compatible host's skills directory to enable it:

git clone https://github.com/tester-army/e2e
cp -r e2e/skills/create-verification-skill/references/example/verify-playground <your-host-skills-dir>/

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • I changed the playground nav link — run verification before I commit to confirm nothing broke
  • Open the playground and screenshot the settings page for me
  • Bug bash the playground's todo feature and compile the findings into a report
  • The smoke test is failing — figure out whether it's my config or the app itself

The skill triggers when an agent needs to verify a playground change, is asked to verify or screenshot a page, or is asked to bug bash it. It starts by requiring you to read .agents/skills/e2e/SKILL.md. Typical flow:
1) Run Doctor first whenever anything looks off:

npx e2e list tests/smoke.e2e.ts
npx e2e run tests/smoke.e2e.ts

2) For signed-in checks, run setup to save the admin / admin-cookie sessions:

npx e2e run tests/auth.setup.e2e.ts

3) Drive live via the MCP server (open_session → observe → locate → verbs) or write scripted tests under tests/; always locate before writing any locator.
4) For proofs that must survive the next run, use an isolated output directory:

npx e2e run --trace on --output .e2e/proof/<name>

5) Bug bash uses e2e.bugbash.config.ts; parallel explorers need pnpm app started first and reuseExisting: true, since the port is fixed. Note APP_ALREADY_RUNNING appears if you start the app by hand and then run tests.

What are this skill's strengths and limitations?

Pros
  • Fully automatic app lifecycle: the runner starts, health-waits, and stops the target — no manual startup, no seeds or database
  • Complete evidence chain: report., trace.md, per-attempt videos, named screenshots, plus --output isolation for durable proofs
  • Strict bug confirmation: only a test encoding the finding failing with ASSERTION_FAILED counts; the explorer's claim alone does not
  • Doctor gives crisp failure attribution: list failure = config/dependency, APP_ALREADY_RUNNING = port conflict, LOCATOR_NOT_FOUND = the app itself
Limitations
  • Scoped strictly to the apps/testbed playground — not directly reusable on other projects
  • Port 4271 is fixed and not reused; parallel explorers require extra setup (pnpm app + reuseExisting)
  • Depends on the untracked e2e.bugbash.config.ts and local features/ charters, which must exist in your environment
  • No documented standalone install for this skill; host must support MCP and shell execution

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
e2e Playground Verification Skill this page 59 · Recommended ★ 8.7k 1d ago Apache-2.0
Create Verification Skill (for e2e) 54 · Use with care ★ 8.7k 1d ago Apache-2.0
e2e: Agentic End-to-End Testing 52 · Use with care ★ 8.7k 1d ago Apache-2.0
Playwright Browser Automation Skill 68 · Recommended ★ 3.2k 1mo ago MIT
playwright-cli Browser Automation Skill 56 · Use with care ★ 11k 13d ago MIT

The README positions e2e as a next-generation end-to-end testing framework that drives browsers via Playwright (@e2e-dev/web) and combines natural-language agent steps with locators/assertions in the same test, but the source names no specific competitor, so no comparison is made.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Recommended
59/ 100 5-point scale 3.0 / 5
1Trust18 / 25 · 3.6/5

The skill describes least-privilege local operation: a fixed port, runner-managed app start/stop, explicit cleanup rules (no pkill node, per-session stops, gitignored proof artifacts), credentials via type_secret/credentials.user with screenshot withholding for password fills, and full data-flow disclosure (localStorage, cookie, stateless server). Deductions: actual execution depends on an external e2e CLI with telemetry (disclosed in README but on by default), no isolation or user-confirmation mechanism in the skill itself, and unverified publisher identity.

2Reliability10 / 20 · 2.5/5

Documentation is self-consistent and thorough: a Doctor-first diagnostic, defined error codes (APP_ALREADY_RUNNING, LOCATOR_NOT_FOUND, ASSERTION_FAILED) with failure attribution, port-conflict handling, and bug-bash config caveats. Deductions: static review cannot execute; key paths (launch, test, bug bash) were not reproduced; the bugbash config is untracked and must be recreated by the reader; some assumptions rely on other topic docs.

3Adaptability10 / 15 · 3.3/5

Trigger conditions are explicit in the description (verify a playground change, screenshot a page, bug bash), boundaries are clear (scoped to apps/testbed, explicitly stating it only describes this app), and the feature map maps routes to tests. Deductions: extremely narrow scope (tied to this repo's testbed), no Chinese-language support statement, and mainland-China reachability of npm/npx dependency is unassessed.

4Convention11 / 15 · 3.7/5

Information architecture is well layered: SKILL.md defers to the underlying e2e skill first, and features/ expands per feature with sub-features, proof, charters, and gotchas; Apache-2.0 license is clear. Deductions: as an example skill it has no independent version, changelog, or maintenance-ownership statement; the untracked bugbash config shifts setup work to the user; known-limitations disclosure is incomplete.

5Effectiveness6 / 15 · 2.0/5

The goal and value proposition are clear: verifiable evidence for playground changes (report, trace, screenshots, proof directory) and a principled distinction between explorer claims and test-confirmed bugs, exceeding manual verification in marginal value. Deductions: static review could not confirm representative outputs are directly usable; the value depends on the e2e framework behaving as claimed, which this review did not establish.

6Verifiability4 / 10 · 2.0/5

The repository contains real CI workflows (agent.yml, benchmark.yml) and committed test suites (mobile-benchmark tests exercising the testbed ecosystem), providing partial auditable evidence at the framework level. Deductions: the assessed skill's own key paths (playground verification flow) lack independent third-party execution evidence, and the correspondence between committed tests and the skill's claims was not verified item by item.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 449fa93670ee Review evidence[1][2][3][4][5][6][7][8][9][10][11][12][13][14]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Can this skill verify my own app?
Not directly. The skill states it only describes what is true of this app (apps/testbed), pinned to port 4271 and testbed test files; use the e2e framework itself for other apps.
Does running it require a paid model?
Scripted tests need no model. The bug bash config adds a model plus skeptic/fuzzer personas, which involves model calls; e2e supports bring-your-own subscription, API key, or local model.
Will run artifacts pollute the git repo?
No. .e2e/results/, .e2e/proof/, and .e2e/bugbash/ are all gitignored. Each new run clears .e2e/results/, so proofs that must outlive it must use --output .e2e/proof/<name>.
What does APP_ALREADY_RUNNING mean?
A server you started by hand (e.g. pnpm app) is already on port 4271. The target port is fixed and never reused — stop your server and retry.

More skills from this repository

All from tester-army/e2e

Dev & Engineering

Create Verification Skill (for e2e)

Generates a project-local verify-<app> skill so any coding agent can launch your app, drive it like a user, keep evidence, and bug bash it.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

e2e: Agentic End-to-End Testing

Drive browser and mobile UI tests with natural-language agent goals, paired with exact locator assertions and a replay cache that keeps model costs down.

★ 8.7k FS 52 Use with care 1d ago
Dev & Engineering

e2e Verify — End-to-End Change Verification Skill

Prove every change with the real CLI against the testbed and benchmark apps — visible evidence, not "it compiles".

★ 8.7k FS 64 Recommended 1d ago
Dev & Engineering

babysit — PR Babysitting Skill

Drive an open pull request through conflicts, review bots, and CI until it is fully green with every thread handled, then hand it off labeled Ready for Human Review.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

Ship a PR (e2e repo PR delivery workflow)

Turn finished work in the tester-army/e2e repo into a PR a human can review without fighting CI or bots: checks, verification, a fresh-context self-review, the PR itself, then babysitting to the Ready for Human Review label.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

writing-pr: PR Title & Body Standards

A PR-writing standard that puts the shape of a change on the first screen: Conventional Commits titles, evidence-driven bodies, and a mandatory local-verification section.

★ 8.7k FS 58 Recommended 1d ago
Writing & Content

e2e Docs Authoring Skill

Applies a skimmer-first writing and shortening standard whenever you write, edit, or restructure guide pages on the e2e docs site.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

unbox-ai — AI Agent Trace Analysis CLI

Analyze AI agent trace files from the CLI without reading megabytes of raw JSON, and find out why an agent run was slow or expensive.

★ 8.7k FS 54 Use with care 1d ago
Writing & Content

Unslop

Edits AI tells out of any text and injects a human voice, so model-drafted writing stops reading like model-drafted writing. Its description says it must always apply.

★ 8.7k FS 49 Use with care 1d ago

Related skills