What does this skill do, and when should you use it?
unbox-ai is a CLI skill that normalizes AI agent trace files and answers questions about them with bounded output. It supports gateway exports, opencode session exports, and AI SDK devtools databases (.devtools/generations.), all auto-detected. Every command is read-only with bounded output, safe to run freely, and never floods the context window with duplicated context. For apps instrumented with @ai-sdk/devtools, it can analyze the database mid-run while the app writes it, enabling live debugging.
- Reads *.trace. (gateway or opencode exports) and .devtools/generations. (AI SDK devtools databases), auto-detecting the format.
- Reports tokens, latency, cost, and per-tool calls/failures via subcommands: runs, summary, tools, events, event, messages, get, compare.
- Supports -- machine-readable output plus jq pipelines for cost hotspots and retry-loop detection.
- compare diffs two runs: token/cost/time deltas, tool-set changes, system prompt diff, and a content-aligned trajectory table over tool sequences.
- On live devtools databases, each command reads a fresh snapshot scoped with --run, so an in-flight agent can be debugged while it executes.
- Diagnose why an agent run was slow or expensive: scan the latency column in events, aggregate cost per segment with the jq recipe, and find the burning segment.
- Debug a stuck agent: use tools to detect retry loops (same tool + same args repeating), then messages --grep for the error evidence trail.
- Live-debug an app instrumented with @ai-sdk/devtools: repeatedly read .devtools/generations. snapshots while the app runs and watch the active run.
- Evaluate prompt-caching wins: check the caching percentage and re-paid token count to decide whether a single-conversation or cache-friendly structure is worthwhile.
- Compare two agent runs to find the divergence point: compare --trajectory shows a content-aligned action table marking exactly where the runs differ.
- Workflows that need to write, replay, or modify traces — all commands are read-only.
- Trace formats outside the three supported ones (gateway exports, opencode exports, AI SDK devtools databases); unknown files fail with a readable error.
- Deep visual exploration is only offered via the view/devtools localhost viewers, which should be offered to the human, not used by the agent.
How do you install this skill?
- The core dependency npx unbox-ai comes from an unverified identity; first run pulls and executes code from npm — review the package or first use it in a controlled environment.
- This is a static review; no commands were executed, so output formats, error behavior, and recipes are unverified in practice.
- The skill has no version or maintenance/update statement, and the attributed unbox-ai project is not in this repository; track upstream changes yourself.
- Trace files analyzed may contain prompts, secrets, or business data; -- and get can emit unbounded raw content — mind data exposure in shared environments.
- No Chinese-language support is declared, and analysis output is oriented toward English CLI output.
- Shell / CLI
- Network access
- Local filesystem
Node.jsnpx unbox-ai CLIjq (for the analysis recipes)
The skill ships in the tester-army/e2e repo's 10-skill collection at .claude/skills/unbox-ai; the source documents no separate skill-install command. The CLI itself runs via npx:
npx unbox-aiThe source does not document steps for copying the skill folder into Claude Code / Codex; place it under the host's skills directory per that host's convention.
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- Look at runs/trace. and tell me why this agent run burned so many tokens — find the most expensive segment.
- Check .devtools/generations. — is the agent stuck in a retry loop on some tool?
- Compare run0. and run1. and show me where the two runs diverged.
- Analyze this opencode session export's cache hit rate — is switching to a single conversation worth it?
Pass a trace file path to a subcommand, starting wide then drilling: summary for the overview, tools for per-tool stats, events for per-generation metrics, event <idx> to drill in, messages --grep to search, get for exact raw values. For multi-run sources, start with runs and scope with --run <n>; truncated output prints the exact get invocation to fetch the rest — follow those pointers. bash
unbox-ai summary trace.
unbox-ai events trace. -- | jq '...'
unbox-ai compare a. b. --trajectory
For live devtools databases, always start with runs and keep passing --run; indexes are run-local. A run marked [live] is still executing — re-run for updated state.
What are this skill's strengths and limitations?
- All commands are read-only with bounded output, safe to run freely without blowing the context window.
- Auto-detects three common trace formats and fails unknown files with a readable error instead of guessing.
- compare --trajectory aligns tool sequences with LCS and pinpoints exactly where two runs diverged.
- Supports mid-run analysis of a database being written live — rare among trace tools.
- When cost is reported as - (AI SDK devtools traces), no dollar amounts are available; only token-based reasoning.
- The jq analysis recipes require jq in the environment.
- -- output is unbounded; plain forms should be preferred first.
- No separate installation docs or test evidence for the skill in the source.
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| unbox-ai — AI Agent Trace Analysis CLI this page | 54 · Use with care | ★ 8.7k | 1d ago | Apache-2.0 |
| OpenCLI AutoFix — Automatic Adapter Repair | 58 · Recommended | ★ 30k | 17d ago | Apache-2.0 |
| Agents Observe | 54 · Use with care | ★ 695 | 1mo ago | MIT |
| Playwright Trace CLI ✓ Microsoft · Official | 48 · Use with care | ★ 97k | 3d ago | Apache-2.0 |
| BlockWatch | 67 · Recommended | ★ 29 | 3d ago | MIT |
The source names no direct competitor, but it explicitly rejects the alternative of cat/Reading raw trace JSON — traces carry megabytes of duplicated context, and unbox-ai's value is replacing that with bounded CLI output.
How did FollowSkills review this skill?
The skill declares all commands read-only with bounded output and restricts interactive viewers to humans, showing good data-flow transparency; however the core tool is fetched via npx unbox-ai whose source is not in the evidence, so dependency and supply-chain security are unverifiable, and there is no rollback or isolation guidance.
Instructions are self-consistent with a clear workflow and described failure behavior (readable errors on unknown files, regex fallback, pointer mechanism), but static review cannot execute anything; availability of the npx package and actual output formats are unproven, so failure-feedback quality rests on documentation alone.
Trigger conditions in the description are fairly precise (when to invoke, three auto-detected formats), audience and scenarios are clear, and boundaries (unsupported formats fail fast, degraded cost handling) are stated; no Chinese-language support is addressed, and format coverage relies on self-description.
SKILL.md is well layered (overview → commands → metric glossary → recipes → formats → live debugging) with progressive disclosure and stable naming; but the skill itself has no version, changelog, or maintenance statement, and attribution points to an unbox-ai project absent from this repository, with license applicability to the skill unstated.
The goal (analyze agent traces without reading raw JSON) is clear with plausible marginal value, and recipes ship copyable jq pipelines; but output correctness is unverified by execution, and first-run npx adds a download cost, so the benefit ratio is only partly confirmed.
Evidence is limited to the skill's own prose; the repository's CI and test suites (benchmark, agent workflows) all target the e2e framework, not the unbox-ai skill, and no third-party execution evidence covers the skill's key paths, capping this dimension.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →