Dev & Engineering

e2e: Agentic End-to-End Testing

Drive browser and mobile UI tests with natural-language agent goals, paired with exact locator assertions and a replay cache that keeps model costs down.

52/ 100
Use with care

Useful, but reliability, evidence or controls still have material gaps.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 8.7k
Last updated
1d ago
License
Apache-2.0
e2e-testingplaywrightmobile-testingagentic-testing
+3typescriptui-automationbug-bash

What does this skill do, and when should you use it?

e2e is an open-source end-to-end testing framework by TesterArmy for web and mobile apps, provided in this repository as an Agent Skill at skills/e2e/SKILL.md. Agents pursue natural-language goals via agent.act while results are judged exactly with locators and assertions (expect, agent.assert, agent.waitFor, agent.extract). The web engine drives Chromium, Firefox, and WebKit through Playwright; the mobile engine drives iOS simulators, Android emulators, and connected phones through agent-device. Verified agent steps are recorded and replayed on later runs without model calls, and tests without agent steps need no model at all. The project is in active development and APIs may still change before 1.0.

  • Scaffolds e2e.config.ts: declares targets (web() or the mobile engine), the app-under-test start command and log path, plus the agents' model and system prompt
  • Writes and runs tests/**/*.e2e.ts: agent.act performs one goal, expect/agent.assert pins each outcome, and screen/app/browser provide exact interactions and checks
  • Runs npx e2e run on a single file, producing .e2e/report. and a trace page per failed test (steps, cache decisions, app logs, agent turns, screen at failure)
  • Replays cached agent steps that passed a recorded check, with no model calls, until the app changes
  • e2e explore explores the app toward a goal without a test file; bug bashes run parallel explore charters, merge findings, and prove each with a repro test
  • e2e mcp serves an MCP server so a coding agent can drive the live app via open_session/observe/locate
Good fit
  • Web developers adding e2e tests to Vite, Next.js, or Astro projects who want agents to handle tedious flows while assertions stay deterministic
  • Mobile teams writing UI tests for iOS simulators, Android emulators, or real devices (example projects exist for Expo, SwiftUI, Jetpack Compose, Flutter, KMP)
  • Engineers asked to bug bash or QA before a release, using parallel explore charters with repro tests proving each finding
  • Testers with Playwright experience who want semantic goals (agent.act) instead of brittle click sequences without giving up exact assertions
  • Coding agents needing to observe and operate a running app through the e2e mcp server's observe-then-write loop
  • Teams running e2e in CI, reading report. and exit codes, with @e2e-dev/github posting results as a PR comment
Not a fit
  • Manual-testing teams or environments without Node.js/npm — the entire workflow is built on the e2e CLI and TypeScript test files
  • Situations where no LLM provider subscription or API key can be configured (except for tests with no agent steps)
  • Projects that cannot tolerate API changes before 1.0 — the README explicitly states the framework is in active development

How do you install this skill?

Before you use it
  • This is a static source review with no execution; confidence is low, and all reliability/effectiveness judgments rest on documented claims only.
  • e2e init modifies .gitignore, .mcp., and .cursor/mcp. and installs dependencies; review these changes manually before accepting.
  • agent.* steps require model calls and incur ongoing costs; validate the environment with tests that have no agent steps first.
  • Model configuration (Vercel AI Gateway, OpenAI/Copilot subscriptions) depends on overseas services; reachability from mainland-China networks is unverified — evaluate proxies or local models yourself.
  • Telemetry is on by default (disable with npx e2e telemetry disable); the feedback command offers redaction and --dry-run, but inspect output with --dry-run before sending.
  • The replay cache and --strict-cache semantics are complex and unverified by execution; pilot on a small scale before production use.
  • Publisher is unverified; example model ids and dependency versions cannot be confirmed — verify actual package versions on upgrade.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
  • MCP Server
Install first
  • Node.js / npm (npx e2e CLI)
  • e2e npm package
  • Playwright browsers via @e2e-dev/web or agent-device via @e2e-dev/mobile
  • an LLM provider credential for agent steps (AI SDK gateway / API key)

The source does not document commands for installing this skill into a client like Claude Code; the skill lives at skills/e2e/SKILL.md in the repo and can be copied into your client's skills directory. Framework setup inside the tested project is documented:

Claude Code / Codex and other agent clients (skill itself, no documented command)

# Copy the skills/e2e/ directory from the tester-army/e2e repo into your client's skills directory

Inside the app under test (framework init, README quick start)

npx e2e init

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • Add end-to-end tests to this Next.js project: initialize the e2e config, then write a test for a member upgrading to the Pro plan
  • Bug bash this app with parallel explore runs and write a repro test for every finding
  • Our e2e test fails in CI — read the trace report, find the cause, and fix the locator
  • Write a mobile test that verifies the billing page shows Pro after login on an Android emulator

Trigger when a project depends on e2e, when asked for end-to-end, browser, mobile, or agentic UI tests, when bug bashing, or when an e2e run fails. Workflow: check for e2e.config.ts and the tests glob (default tests/**/*.e2e.ts); if absent, follow the setup topic. Learn the screens' accessible names from components, with --headed, or via the e2e mcp server. Write tests/<feature>.e2e.ts with one goal per act and an assertion right after; run a single file with npx e2e run tests/<feature>.e2e.ts; on failure read the named trace page at .e2e/results/<test>/trace.md and never add a sleep. Run the CLI as npx e2e (or pnpm exec e2e). Read all eight topics offline with npx e2e guide <topic> (setup, writing-tests, agent, running, explore, debugging, mcp, bug-bash). Secrets are declared under credentials and secrets in the config and never appear in test code.

What are this skill's strengths and limitations?

Pros
  • Combines agent flexibility with exact locator assertions — semantic goals without sacrificing deterministic outcomes
  • Replay cache avoids model calls for verified act steps while the app is unchanged, keeping cost predictable
  • Dual web and mobile engines (Playwright + agent-device) with complete runnable examples for eight stacks
  • Trace pages on failure include cache decisions, app logs, and the screen at failure
  • Built-in secrets/credentials namespaces keep secrets out of test code
Limitations
  • Pre-1.0: APIs and config can change between minor releases
  • Agent steps require LLM provider authentication and carry ongoing cost and budget-management concerns
  • No benchmarks, independent reviews, or stability evidence in the source material — reliability must be verified yourself
  • The CLI sends anonymous telemetry by default (commands, engines, failure locations); opt out with E2E_TELEMETRY_DISABLED=1

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
e2e: Agentic End-to-End Testing this page 52 · Use with care ★ 8.7k 1d ago Apache-2.0
Create Verification Skill (for e2e) 54 · Use with care ★ 8.7k 1d ago Apache-2.0
e2e Verify — End-to-End Change Verification Skill 64 · Recommended ★ 8.7k 1d ago Apache-2.0
e2e Playground Verification Skill 59 · Recommended ★ 8.7k 1d ago Apache-2.0
Playwright Browser Automation Skill 68 · Recommended ★ 3.2k 1mo ago MIT

Both SKILL.md and README state the web engine is built on Playwright (Chromium, Firefox, WebKit) — effectively a layer over the Playwright runtime adding agent goal-driving, a replay cache, and a mobile engine. No other specific competitor is named in the source.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Use with care
52/ 100 5-point scale 2.6 / 5
1Trust14 / 25 · 2.8/5

Positives: dedicated credential/secret namespaces with plaintext never reaching the model (type_secret, Secret masking), pixels masked and withheld after secret fills, feedback command declares redaction and offers --dry-run, upload restricted to project root with .env denied (POLICY_DENIED), telemetry can be disabled, CI example uses least-privilege permissions. Deducted: no disclosed supply-chain detail for runner-installed/downloaded artifacts; init's automatic modification of .gitignore/.mcp./.cursor configs lacks an explicit user-confirmation step; feedback and telemetry data flows are self-declared only, not third-party verifiable; publisher unverified, attribution rests on repository claims alone.

2Reliability10 / 20 · 2.5/5

Positives: internally highly self-consistent docs; complete error-code table (debugging) with fixes; clear exit codes; defined trace.md and report. structures; controlled failure modes for abnormal input (ambiguous locators, invalid model output, context overflow). Deducted: static review cannot execute or reproduce; complex replay-cache semantics rest entirely on documented claims; example model ids (gpt-6-luna-fast etc.) unverifiable; many described behaviors depend on unseen implementation code.

3Adaptability9 / 15 · 3.0/5

Positives: precise trigger description and topic index, clear workflow, browser/mobile/CI/MCP coverage, partial boundary disclosure ('When it does not fit'). Deducted: limited non-fit notes for non-TS/non-ESM projects and older Node environments; heavy dependence on overseas services (Vercel AI Gateway, GitHub Copilot subscriptions, npm) with no mirror or mainland-China alternative path disclosed; no Chinese-language support mentioned.

4Convention9 / 15 · 3.0/5

Positives: layered docs (SKILL.md plus 8 topic references plus built-in guide command plus in-package docs), rich examples, Apache-2.0 license metadata, known-limitation and debugging content. Deducted: publisher identity unverified, ownership and update path unclear; no version number or changelog in the skill docs; SKILL.md is very long and dense; side effects of e2e init on user files are disclosed only in scattered places.

5Effectiveness6 / 15 · 2.0/5

Positives: clear goal (agentic e2e testing), well-defined output formats (.e2e.ts tests, report., trace.md), bug-bash workflow includes verification and rejection triage, plausible marginal value over hand-written Playwright tests. Deducted: static review cannot confirm representative outputs are directly usable; agent steps incur model costs with no quantified cost/benefit data; the claimed repro-test-proven bug workflow cannot be execution-verified.

6Verifiability4 / 10 · 2.0/5

Positives: consistent cross-references between topics, error codes and config keys corroborated across files, CI example pins action versions, bug-bash explicitly separates model claims from verified findings. Deducted: no third-party execution evidence in a static read; key claims (examples repo, online docs, model ids) cannot be checked within the provided files; no test-suite results or independent user validation available.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision 449fa93670ee Review evidence[1][2][3][4][5][6][7][8]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Can I use it without a model?
Yes. Tests without agent steps need no model. Only agent.* steps (agent.act, agent.assert, etc.) require a model in the config plus that provider's authentication (a saved subscription login or API key); a local endpoint may need none.
How is model cost controlled?
The replay cache is the main lever: an act step verified by a later assertion records its actions and is replayed without a model call while the app is unchanged. The agent topic also covers cost and budgets.
How do I debug a failed run?
The terminal names a trace page per failed test (.e2e/results/<test>/trace.md) with every step, the cache's decision, app logs, agent turns, and the screen at failure. There are also --headed, --debug, --ai-trace flags and error-code docs (debugging topic).
Which platforms and frameworks are supported?
Web through Playwright (Chromium, Firefox, WebKit); mobile through agent-device (iOS simulators, Android emulators, connected phones). Examples cover Vite, Next.js, Astro, Expo, SwiftUI, Jetpack Compose, Kotlin Multiplatform, and Flutter.

More skills from this repository

All from tester-army/e2e

Dev & Engineering

Create Verification Skill (for e2e)

Generates a project-local verify-<app> skill so any coding agent can launch your app, drive it like a user, keep evidence, and bug bash it.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

e2e Verify — End-to-End Change Verification Skill

Prove every change with the real CLI against the testbed and benchmark apps — visible evidence, not "it compiles".

★ 8.7k FS 64 Recommended 1d ago
Dev & Engineering

e2e Playground Verification Skill

Launch, drive, and prove the apps/testbed playground works — with traces, screenshots, and a bug bash — before you claim a change works.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

babysit — PR Babysitting Skill

Drive an open pull request through conflicts, review bots, and CI until it is fully green with every thread handled, then hand it off labeled Ready for Human Review.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

Ship a PR (e2e repo PR delivery workflow)

Turn finished work in the tester-army/e2e repo into a PR a human can review without fighting CI or bots: checks, verification, a fresh-context self-review, the PR itself, then babysitting to the Ready for Human Review label.

★ 8.7k FS 59 Recommended 1d ago
Dev & Engineering

writing-pr: PR Title & Body Standards

A PR-writing standard that puts the shape of a change on the first screen: Conventional Commits titles, evidence-driven bodies, and a mandatory local-verification section.

★ 8.7k FS 58 Recommended 1d ago
Writing & Content

e2e Docs Authoring Skill

Applies a skimmer-first writing and shortening standard whenever you write, edit, or restructure guide pages on the e2e docs site.

★ 8.7k FS 54 Use with care 1d ago
Dev & Engineering

unbox-ai — AI Agent Trace Analysis CLI

Analyze AI agent trace files from the CLI without reading megabytes of raw JSON, and find out why an agent run was slow or expensive.

★ 8.7k FS 54 Use with care 1d ago
Writing & Content

Unslop

Edits AI tells out of any text and injects a human voice, so model-drafted writing stops reading like model-drafted writing. Its description says it must always apply.

★ 8.7k FS 49 Use with care 1d ago

Related skills