Dev & Engineering

jev-use — A Bounded Computer-Use Loop over Cua Driver

Build a tightly bounded computer-use loop on top of Cua Driver: the driver only observes and acts, TypeSafe Jev picks from application-owned candidate IDs, and the caller verifies every action.

58/ 100
Recommended

Generally reliable with disclosed limitations; trial as directed and keep a rollback path.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 29k
Last updated
1d ago
License
MIT
computer-usecua-driverdesktop-automationaccessibility
+2visual-groundingbounded-decision-loop

What does this skill do, and when should you use it?

jev-use is a recipe defined in skills/jev-use/SKILL.md of the trycua/cua repository (MIT licensed). Its core design keeps the decision layer entirely above Cua Driver: the Driver supplies observations and executes actions, the application constructs complete candidate tables, and TypeSafe Jev returns only one candidate ID — it never invents tool names, coordinates, or arguments. The recipe prescribes strict rules for observation freshness, visual evidence validation, native accessibility candidate construction, and credential handling, and offers a deterministic mock path so development can proceed without a model API key. A runnable reference lives in libs/cua-driver/examples/jev-use/.

  • Obtains fresh Cua Driver observations (browser DOM, native accessibility tree, or visual regions) through one persistent CLI or MCP session.
  • Builds a candidate table capped at 24 action candidates plus reobserve and abstain, each with the complete Driver tool and arguments.
  • Communicates with TypeSafe Jev over stdin/stdout using the cua.jev_choice_request_v1/v2 protocol, sending only goal, compact observation, history, and candidate IDs.
  • Validates the returned choice, rejecting unknown, duplicate, stale, capture-mismatched, or below-confidence results.
  • Executes at most one Driver action (background delivery by default), then reobserves and verifies completion against an independent postcondition.
Good fit
  • An engineer building a desktop automation agent who wants the model restricted to a whitelist of application-owned actions instead of freely generating coordinates and tool calls.
  • A team that wants System 1 style selection decisions over structured interface elements, with execution firmly locked behind driver boundaries.
  • Developers driving native desktop apps on macOS, Windows, or Linux who need candidates built from the accessibility tree (AX/UIA/AT-SPI).
  • Automation scenarios mixing browser DOM with optional visual grounding, where every visual click must be bound to an exact capture ID.
  • Teams that want to develop against the deterministic mock path first, then switch to live Jev without code changes in the recipe.
Not a fit
  • Use cases where the model should freely generate coordinates, refs, or tool calls — the recipe explicitly forbids Jev from inventing any arguments.
  • Users without a running Cua Driver environment (macOS/Windows/Linux desktop) or unwilling to integrate a TypeSafe Jev adapter.
  • Tasks that require actions beyond the default risk policy — delete, send, purchase, and close actions are excluded unless the task spec allows them.

How do you install this skill?

Before you use it
  • This is a static source review; no tests were executed. Key-path reproducibility is inferred from file self-consistency and CI configuration.
  • The skill depends on external Cua Driver and extension contracts (including AGPL-licensed cua-perception assets); verify licensing and trust boundaries before deployment or redistribution.
  • Reachability of core dependencies (cua.ai installers, GitHub, Hugging Face) from mainland China is not declared; Chinese users may need extra network configuration.
  • Foreground delivery and unrestricted permission mode (CUA_DRIVER_PERMISSION_MODE: unrestricted in CI) carry real desktop-action risk; production use should enable approval and authorization flows.
  • Publisher identity is unverified/unknown; the skill file itself lacks a version number and update path.
Before you start
Your agent needs
  • Shell / CLI
  • Network access
  • Local filesystem
  • MCP Server
Install first
  • Cua Driver
  • TypeSafe Jev adapter (mock path runs without TYPESAFE_API_KEY)

The source documents no per-host install command for this skill alone. The repository README notes that the cua installer can install cua skills into AI coding agents:

curl -fsSL https://cua.ai/install.sh | sh
cua auth login

The installer offers to add cua skills and the cua MCP server to Claude Code, Codex, Cursor, and others; the skill file itself is at skills/jev-use/SKILL.md in the repo.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • Set up a bounded computer-use loop in mock mode following the jev-use recipe that clicks a button in Calculator and verifies the result.
  • Build a candidate table for my desktop automation agent so the model can only choose from application-owned candidate IDs, never generate coordinates.
  • Following libs/cua-driver/examples/jev-use/, wire up NativeAccessibilitySource and cua.jev_choice_request_v2 for a native desktop app.
  • How should jev-use rules trigger reobserve or abstain when visual and accessibility evidence disagree?

The skill is triggered as a recipe/guide describing how to structure the jev-use loop: state the goal and get a fresh Driver observation through one persistent CLI or MCP session; prefer fresh accessibility or browser DOM tokens; if visual grounding is needed, discover parse_visual_regions through the current MCP tool inventory and validate its versioned result and capture ID; build a bounded candidate table and send only goal, compact observation, history, and candidate IDs to Jev; resolve one returned ID, execute at most one action, and reobserve. Browser tasks use the v1 protocol; native desktop apps use NativeAccessibilitySource and v2. The runnable example is in libs/cua-driver/examples/jev-use/; use checked-in fixtures for deterministic development. The mock path requires no TYPESAFE_API_KEY; the live key comes only from the process environment or a secure interactive prompt.

What are this skill's strengths and limitations?

Pros
  • Crisp decision boundary: the model can only pick from application-constructed candidate IDs, eliminating hallucinated arguments and unbound coordinate clicks.
  • Deterministic mock path works with no model API key; Python and TypeScript adapters expose equivalent mock and live behavior.
  • Explicit, actionable rules for freshness, visual evidence validation, native role mapping, and credential handling.
  • Ships a runnable example and checked-in fixtures; the host repository is MIT licensed.
Limitations
  • Tightly coupled to the Cua Driver and TypeSafe Jev ecosystem; unusable outside that stack.
  • Candidate table capped at 24 actions and delete/send/purchase/close excluded by default, limiting flexibility.
  • The visual adapter only works when the Driver advertises both parse_visual_regions and the capture-bound click.capture_id input.
  • The source provides no independent test suite or benchmark results proving the recipe's own success rate.

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
jev-use — A Bounded Computer-Use Loop over Cua Driver this page 58 · Recommended ★ 29k 1d ago MIT
Desktop App Control with Computer Use 61 · Recommended ★ 6.4k 3d ago Apache-2.0
squirrelscan Website Audit & Fix Loop 49 · Use with care ★ 97 3d ago MIT
Angular Developer Skill 63 · Recommended ★ 674 3d ago —
SwiftUI WCAG Accessibility Auditor Skill 58 · Recommended ★ 13 7mo ago —

The repo README covers the whole Cua product family (Driver, Spaces, Lume, CUA-S1, Bench); jev-use is one specific recipe for integrating a TypeSafe Jev decision model with Cua Driver, as opposed to letting a general agent freely call Driver tools. The source names no third-party alternatives.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Recommended
58/ 100 5-point scale 2.9 / 5
1Trust18 / 25 · 3.6/5

SKILL.md enforces least privilege and clear safety boundaries: Jev may only select from application-owned candidate IDs, never invent tool names or coordinates; capture_id binding prevents stale coordinate clicks; delete/send/purchase/close actions are excluded by default; keys are read only from environment or secure interactive prompts, never in source, args, or logs; foreground delivery requires explicit escalation/authorization. Deductions: trust relies on external Driver/extension contracts (e.g., cua-perception from signed release assets with AGPL caveats) and the file itself lacks detailed rollback and user-confirmation flow.

2Reliability10 / 20 · 2.5/5

The decision loop, freshness rules, and failure modes (reobserve/abstain, rejecting unknown/duplicate/stale choices) are clearly specified and self-consistent; the mock path works without an API key and a dedicated CI workflow (authorized-live-jev-use.yml) exercises mock checks. Deduction: static review only — the referenced example directory libs/cua-driver/examples/jev-use/ was not provided, so key-path reproducibility and failure-feedback quality cannot be confirmed; capped at 10.

3Adaptability9 / 15 · 3.0/5

Trigger conditions are precise in the description (jev-use recipe or similar Jev integrations; explicitly not for adding model logic or credentials), with declared boundaries between browser v1 and native v2 and stated non-fit ranges. Deductions: narrow developer audience requiring TypeSafe Jev and Cua Driver; no Chinese-language documentation; reachability of core dependencies (cua.ai installers, Hugging Face, GitHub) from mainland China is not addressed.

4Convention10 / 15 · 3.3/5

Well-layered structure (decision loop, freshness, native candidates, credentials/proof), MIT license is clear, and the repo shows version-anchoring and governance signals. Deductions: the skill file itself has no version, changelog, or maintenance-owner statement; install/dependency notes live outside the skill.

5Effectiveness6 / 15 · 2.0/5

A runnable reference example and deterministic mock path are provided, with verification via an independent application oracle rather than model output; the decoupling-of-decision-and-execution value proposition is clear. Deductions: static review cannot confirm outputs are directly usable; example/test details are only partially visible and comparative-benefit evidence is limited.

6Verifiability5 / 10 · 2.5/5

Auditable primary material exists: specific contracts in SKILL.md are cross-corroborated by repository CI workflows including mock assertions, credential-leak checks, and audit scripts. Deduction: nothing was executed, coverage is limited, and independently reproducible conclusions are not achievable from a static read.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Oct 10, 2026 Reviewed revision e32127764436 Review evidence[1][2][3][4][5][6][7][8][9][10][11]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Do I need a model API key?
Not for development: the deterministic mock path runs without TYPESAFE_API_KEY. For live Jev, the key is read only from the process environment or a secure interactive prompt — never from source, logs, or command arguments.
Which platforms are supported?
Cua Driver operates native apps and browsers on macOS, Windows, and Linux. The recipe uses the v2 accessibility protocol for native desktop apps and v1 for browser tasks.
Why must visual clicks bind to a capture_id?
Page refs, screenshot IDs, and visual region IDs are observation-local and invalid once the UI changes. The recipe forbids retrying an expired capture as an unbound coordinate action; conflicting semantic and visual evidence triggers reobserve or abstain.
How are risky actions handled?
Delete, send, purchase, and close actions are excluded from the candidate table by default unless the task spec explicitly allows that risk, and action text comes only from task parameters.

More skills from this repository

All from trycua/cua

Automation & Ops

Cua Driver Skill

Let agents operate real desktop applications: drive native GUIs on macOS, Windows, and Linux via accessibility trees and element tokens, then verify from fresh state.

★ 29k FS 58 Recommended 1d ago
Automation & Ops

Cua GUI Automation Skill

Give AI agents eyes and hands on a real computer: click buttons, fill forms, and run end-to-end visual QA on any application's GUI.

★ 29k FS 51 Use with care 1d ago
Automation & Ops

Cua Driver GUI Automation Skill

Lets an AI agent operate real native app windows on macOS, Windows, and Linux: observe state, act precisely, and verify the outcome.

★ 29k FS 61 Recommended 1d ago
Dev & Engineering

Cua Sandboxes

Spin up disposable Linux or macOS machines locally or in the Cua cloud, so agents can run code, test apps, drive a desktop GUI, and browse the web without ever touching your own computer.

★ 29k FS 52 Use with care 1d ago
Automation & Ops

Cua Volume Skill

Give AI agents one versioned, user-level shared volume that persists files across Spaces, syncs to your Mac in seconds, and shares safely with other agents.

★ 29k FS 52 Use with care 1d ago
Automation & Ops

Cua Spaces (cua-spaces skill)

Through the cua MCP server, lets your agent work inside a watchable remote or local computer — running commands, moving files, spawning coding agents and sharing the host network.

★ 29k FS 46 Use with care 1d ago
Dev & Engineering

Poll GitHub Work

Polls and ranks open GitHub issues, RFCs, and PRs so maintainers know what to work on next — or starts one explicitly selected item.

★ 29k FS 52 Use with care 1d ago

Related skills