Data & Analysis ✓ NVIDIA · Official synthetic-datadataset-generationdata-pipelinesnemo-data-designerpython

Data Designer Synthetic Data Skill

Build synthetic datasets and data-generation pipelines from a natural-language specification.

FollowSkills review · FSRS-2.0
Not recommended
52/ 100 5-point scale 2.6 / 5
Trust17 / 25 · 3.4/5

The skill limits workspace exploration, forbids installing dependencies without permission, and gives controls for seed data, personal-data sampling, sandbox retries, licensing, and security reporting. Points are deducted because external model/network data flows, synthetic-person privacy risks, rollback, and general confirmation before preview/create are not fully disclosed.

Reliability7 / 20 · 1.8/5

Interactive and autopilot workflows, missing-CLI handling, network-failure guidance, parameter inspection, and several common-error messages are specified. Points are deducted because there is no executable test suite for this skill, key behavior depends on an external CLI, model configuration, and context files not supplied here, and several failure paths can only be inferred statically; the static cap keeps this at or below 10.

Adaptability9 / 15 · 3.0/5

The audience, dataset-description input, Python-script output, interactive/autopilot routing, and seed-data conditions are reasonably clear. Points are deducted for weak non-fit boundaries, model/hardware/cost/scale guidance, limited Chinese-language support, and uncertain mainland-China reachability where preview or generation may depend on overseas services.

Convention10 / 15 · 3.3/5

The documentation includes goals, rules, workflows, tips, troubleshooting, an output template, and layered references; the skill card states owner, Apache 2.0 licensing, version v0.6.1, use case, and governance-related information. Points are deducted for no formal changelog, no explicit maintenance owner/update path, limited tabular examples and FAQ coverage, and the supplied metadata being marked NOASSERTION; the benchmark also reports missing recommended Instructions, metadata.author/tags, and run_script guidance.

Effectiveness6 / 15 · 2.0/5

The skill clearly targets a directly usable Python configuration file with load_config_builder() and covers samplers, person data, validation, preview, and creation. The benchmark reports positive correctness and effectiveness results for the core task. Points are deducted because the evaluation dataset and reproducible artifacts are unavailable, outputs still require review, and success depends on external CLI, model, and configuration availability; the static cap keeps this at or below 7.

Verifiability3 / 10 · 1.5/5

The pinned source material provides workflows, references, an example script, and a benchmark report, giving limited auditability. Points are deducted because the benchmark source data is unavailable, no committed tests cover this skill’s key paths, and linked third-party documentation is not reproduced or independently corroborated; the static cap keeps this at or below 5.

Evidence confidence:Low Reviewed Jul 20, 2026 Reviewed revision 55f18499943e
Before you use it
  • Preview and generation may send prompts, seed fields, or synthetic data to external models or network services; confirm data flow, compliance, and de-identification before using real personal information.
  • Autopilot proceeds through validation, preview, and optional creation; beyond the stated large-dataset confirmation, explicitly confirm network calls, cost, output location, and record count before execution.
  • Mainland-China reachability, model-alias configuration, and local dependency availability are not verified; provide a local-model or offline fallback where needed.
  • The benchmark report omits its raw evaluation dataset and reproducible artifacts, so it should not be treated as independent execution verification.
See the full review method →

What does this skill do, and when should you use it?

This skill is for creating datasets, generating synthetic data, and building data-generation pipelines. It directs an agent to use the Data Designer library to produce a Python configuration builder from the user’s dataset description. It supports Interactive mode and Autopilot mode, with Interactive as the default. The instructions define when to read workflow and reference files and provide rules for retaining columns, handling seed data, and configuring validation and templated fields.

Selects the Interactive or Autopilot workflow from the user’s request; reads the matching workflow and, when needed, the person-sampling or seed-dataset reference; builds a synthetic-dataset configuration with the Data Designer API; writes a Python file in the current directory containing load_config_builder(); adds PEP 723 inline dependency metadata; and provides prescribed responses for a missing CLI, Python-version requirements, and preview network failures.

  1. A data engineer needs a synthetic dataset generated from a natural-language specification.
  2. An ML team needs a configurable pipeline for producing training or evaluation data.
  3. A user wants an agent to make generation decisions interactively or autonomously.
  4. A developer needs sampler, validation, Jinja2 expression, or LLM-judge columns configured correctly.

What are this skill's strengths and limitations?

Pros
  • Targets synthetic dataset creation and data-generation pipeline construction.
  • Offers both Interactive and Autopilot workflows.
  • Documents rules for column retention, seed data, person sampling, templating, and LLM-judge score access.
  • Defines a concrete Python output interface and dependency format.
Limitations
  • The supplied SKILL.md does not include the contents of either workflow file.
  • Execution depends on the Data Designer library and an appropriate runtime environment.
  • Preview operations may require network access, but no specific network configuration is documented.
  • The `$ARGUMENTS` substitution mechanism requires support from the target agent or light adaptation.

How do you install this skill?

Run npx skills add nvidia/skills --skill data-designer --yes to install the skill in a supported agent environment. The README does not specify the exact local installation path for each agent.

How do you use this skill?

Describe the dataset to the agent, for example: Create a synthetic dataset of customer reviews with ratings and quality scores. Use Autopilot when the user asks the agent to decide or does not want questions; otherwise use the default Interactive mode.

FAQ

Is a seed dataset required?
No. The skill says to use one only when the user explicitly provides seed data or asks to build from existing records.
What runtime is required?
The `data-designer` CLI must be available, and it requires Python >= 3.10. If it is missing, the instructions require asking before creating an environment or installing it.
Can the workflow run without user questions?
Yes. Autopilot is intended for requests such as “you decide” or “make reasonable assumptions”; Interactive is the default otherwise.
Will the skill drop output columns?
No by default. A column may be dropped only when the user requests it or when it is solely a helper column for deriving other columns.

More skills from this repository

All from NVIDIA/skills

Data & Analysis ✓ NVIDIA · Official

NeMo Data Designer Synthetic Data Skill

Build synthetic datasets and declarative data-generation pipelines from a natural-language description.

Dev & Engineering ✓ NVIDIA · Official

Clinical ASR Flywheel: Environment Setup

Validate that a clinical ASR evaluation environment can complete a TTS-to-ASR round trip through NVIDIA-hosted speech services.

Data & Analysis ✓ NVIDIA · Official

NV-Generate-MR

Generate synthetic body MRI volumes through NVIDIA’s rflow-mr workflow.

Data & Analysis ✓ NVIDIA · Official

NVIDIA Physical AI Defect Image Generation

Orchestrate defect-image generation, augmentation, inference, and labeling for AOI datasets on OSMO.

Data & Analysis ✓ NVIDIA · Official

PAIDF AnomalyGen

Fine-tune, generate, evaluate, and refine synthetic anomaly images.

Dev & Engineering ✓ NVIDIA · Official

NeMo Relay Installation Guide

Choose and verify the right NeMo Relay path for CLIs, language packages, and maintained frameworks.

Dev & Engineering ✓ NVIDIA · Official

NeMo Relay Quick Start

Helps first-time NeMo Relay users prove observable execution value through the smallest suitable trial.

Dev & Engineering ✓ NVIDIA · Official

DALI Dynamic Mode Assistant

Helps agents write, review, and migrate NVIDIA DALI imperative dynamic-mode code.

Dev & Engineering ✓ NVIDIA · Official

NeMo Relay Migration Assistant

Safely migrate NeMo Flow projects to NeMo Relay.

Data & Analysis ✓ NVIDIA · Official

Earth2Studio Weather Data Fetch

Fetch validated weather and climate variables from Earth2Studio sources by time and data type.

Dev & Engineering ✓ NVIDIA · Official

AMC Sample Dataset Calibration

Verify a running NVIDIA AutoMagicCalib service end to end with its bundled sample dataset.

Dev & Engineering ✓ NVIDIA · Official

NeMo Relay Typed Wrappers & Codecs

Add typed boundaries to NeMo Relay integrations while preserving predictable JSON middleware semantics and caller-visible behavior.

Data & Analysis ✓ NVIDIA · Official

TAO DAFT Dataset Converter

Guides AI agents through tao-daft conversion between supported NVIDIA TAO DAFT dataset formats.

Data & Analysis ✓ NVIDIA · Official

Synthetic Brain MRI Generator

Generate synthetic brain MRI volumes through NVIDIA’s documented workflow.

Data & Analysis ✓ NVIDIA · Official

RAGAS RAG Quality Evaluator

Benchmark retrieval-augmented generation quality against filesystem datasets.

Dev & Engineering ✓ NVIDIA · Official

Earth2Studio Data Source Builder

Connect remote weather data to Earth2Studio with tested source wrappers.

Data & Analysis ✓ NVIDIA · Official

DICOM Series Preflight

Header-only validation for one DICOM series before conversion or inference.

Dev & Engineering ✓ NVIDIA · Official

Earth2Studio Diagnostic Builder

Build Earth2Studio wrappers for single-step diagnostic data transformations.

Data & Analysis ✓ NVIDIA · Official

NVIDIA AI-Q Deep Research

Run deep research through a reachable local or self-hosted AI-Q Blueprint backend.

Dev & Engineering ✓ NVIDIA · Official

Holoscan SDK Setup Guide

Inspects a Linux host and selects the most suitable Holoscan SDK installation path.

Related skills