Automation & Ops

Agent Browser: Deterministic Web Automation

Headless browser automation for AI agents using accessibility tree snapshots and ref-based element selection, enabling fast and reliable web interactions.

49/ 100
Use with care

Useful, but reliability, evidence or controls still have material gaps.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 20k
Last updated
7d ago
License
MIT
browser-automationaccessibility-treeheadless-browserref-based-selection
+2session-isolationstate-persistence

What does this skill do, and when should you use it?

This skill wraps the agent-browser CLI tool to provide fast, deterministic browser automation. It generates element refs from accessibility tree snapshots, avoiding brittle CSS selectors. It supports multi-tab, iframe, network interception, cookie and storage management, session isolation, and state persistence. The skill is ideal for complex single-page applications (SPAs) and multi-step workflows requiring stable, repeatable interactions.

Performs headless browser operations including: opening/navigating/back/forward/reloading/closing pages; generating accessibility tree snapshots (interactive elements, JSON output, scoped to selector); clicking, filling, typing, hovering, checking/unchecking, selecting, pressing keys, scrolling, and dragging via refs; getting text, html, value, attributes, title, URL, and element count; checking element visibility, enabled state, and checked state; waiting for elements, delays, text, URL, network idle, or custom function; creating and managing isolated browser sessions; saving and loading auth state (cookies and storage); taking screenshots and generating PDFs; intercepting and mocking network routes; managing cookies and localStorage; switching tabs and iframes.

Good fit
  • Automating multi-step web form submissions, such as registration or checkout flows, requiring stable and deterministic element selection.
  • Scraping data from websites requiring login, by saving and loading auth state to avoid repeated logins.
  • Testing user flows in complex single-page applications (SPAs) like internal dashboards, where DOM updates frequently.
  • Handling multiple user sessions (e.g., admin and user) simultaneously for parallel browser testing.
  • Mocking API responses via network interception for frontend development or testing without a backend service.

How do you install this skill?

Before you use it
  • Skill depends on external CLI tool agent-browser; its availability and security are not verified in the document.
  • No mention of sensitive data handling, credential storage, or network request security; users must evaluate these themselves.
  • No troubleshooting guide, making it hard to diagnose issues.
  • Documentation is in English; Chinese users may need translation, and mainland China network reachability is not addressed.
  • Static review; not executed, reliability needs further verification.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
Install first
  • Node.js
  • agent-browser CLI

Install the agent-browser CLI globally: npm install -g agent-browser. Then run agent-browser install to download Chromium (add --with-deps on Linux for system dependencies). This skill requires no special installation beyond ensuring the command is available in your PATH.

Generic route: install into Claude Code manually (macOS / Linux)
tmp="$(mktemp -d)"
git clone --depth 1 https://github.com/eosphoros-ai/DB-GPT.git "$tmp"
mkdir -p ~/.claude/skills
cp -R "$tmp/skills/agent-browser" ~/.claude/skills/
rm -rf "$tmp"

Generated from the source repository and skill path; it copies only this skill's folder. If the author's install steps above differ, follow those first. To scope it to one project, replace ~/.claude/skills with that project's .claude/skills.

How do you use this skill?

Invoke agent-browser commands as a tool. Typical workflow: first agent-browser open <url>, then agent-browser snapshot -i --json to get interactive element snapshot and parse refs, then interact with agent-browser click @e2 or agent-browser fill @e3 "text". Re-snapshot after page changes. Best practices: always use -i --json flags, wait for stability (wait --load networkidle), use sessions to isolate contexts, and save auth state with state save/load. Use --headed for debugging.

What are this skill's strengths and limitations?

Pros
  • Ref-based selection is more stable than CSS selectors, resistant to DOM changes
  • Lightweight with only a Node.js CLI and Chromium dependency
  • Supports session isolation and state persistence to skip login flows
  • Network interception and mocking capabilities for versatile testing
  • Accessibility tree snapshots output JSON for easy parsing, AI-agent friendly
Limitations
  • Requires installation and download of Chromium, which may consume significant disk space
  • Chromium-only; does not support other browser engines
  • Documentation does not mention Windows support (installation commands include Linux-specific `--with-deps`)
  • Smaller ecosystem and community compared to other automation tools
  • Headless mode may require additional wait logic for complex pages to ensure stability

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
Agent Browser: Deterministic Web Automation this page 49 · Use with care ★ 20k 7d ago MIT
Cua Driver GUI Automation Skill 61 · Recommended ★ 29k 1d ago MIT
Cua Driver Skill 58 · Recommended ★ 29k 1d ago MIT
cmux Computer Use Skill 50 · Use with care ★ 28k 1d ago NOASSERTION
OpenCLI Sitemap Author 56 · Use with care ★ 30k 17d ago Apache-2.0

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Use with care
49/ 100 5-point scale 2.5 / 5
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
1Trust15 / 25 · 3.0/5

The skill document provides clear usage boundaries (when to use agent-browser vs built-in browser tool) and mentions session isolation and state persistence, but lacks details on least privilege, user confirmation, data-flow transparency, or sensitive-data handling. As a wrapper around a CLI tool, no malicious or overreaching risks found, but security guidance is missing, deducting to 15.

2Reliability6 / 20 · 1.5/5

Document includes installation commands, core workflow, and many command examples, but no tests or error-handling instructions; static review cannot verify actual operation. Deducted to 6.

3Adaptability10 / 15 · 3.3/5

Clearly describes applicable scenarios (complex SPAs, multi-step workflows) and non-applicable scenarios (needing screenshots/PDFs), trigger conditions are clear. But no mention of environment dependencies (e.g., network access) or Chinese language support, and core function depends on an external CLI tool, limiting applicability. Deducted to 10.

4Convention9 / 15 · 3.0/5

Documentation is well-structured with installation notes, command examples, and best practices, but lacks versioning, changelog, troubleshooting, or known limitations; maintenance responsibility unclear due to dependency. Deducted to 9.

5Effectiveness5 / 15 · 1.7/5

Shows two examples (search & extract, multi-session testing) but no actual outputs or verification; static review cannot confirm direct usability. Deducted to 5.

6Verifiability4 / 10 · 2.0/5

Document is purely author description without tests or other verification evidence; static review cannot independently confirm functionality. Deducted to 4.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Aug 07, 2026 Reviewed revision 4211e02c10be Review evidence[1][2][3][4][5][6][7][8]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Does this skill cost anything?
No. agent-browser is open source and Chromium is free. You only need Node.js and npm to install the CLI globally.
Will it work with my website?
It works with any website accessible via Chromium. For sites requiring login, you can use `state save/load` to persist auth; for SPAs, it handles dynamic content well via accessibility tree snapshots.
What if an element has no ref?
Ensure you use `-i` to include only interactive elements. If an element is not captured, try scoping the snapshot with `-s "#main"` or check if it's inside an iframe that needs switching via `frame`.
Where can I run it?
It can run on any OS with Node.js and Chromium installed, but the documentation explicitly mentions Linux and macOS. Windows support may be untested.

More skills from this repository

All from eosphoros-ai/DB-GPT

Related skills