Data & Analysis ✓ NVIDIA · Official cudfgpu-dataframespandasdask-cudfetlparquetnullable-semantics

NVIDIA cuDF DataFrame Guide

Helps pandas users write correct, high-performance GPU DataFrame code with NVIDIA cuDF.

FollowSkills review · FSRS-2.0
Not recommended
57/ 100 5-point scale 2.9 / 5
Trust18 / 25 · 3.6/5

The skill is scoped to local cuDF/data-processing guidance and shows no malware, credential theft, covert exfiltration, destructive default, or grossly excessive permission. It makes GPU/CPU boundaries and conversion boundaries visible, but lacks explicit user confirmation, sensitive-data handling, dependency-security guidance, rollback, and clear data-flow controls for external documentation access, so points are deducted.

Reliability9 / 20 · 2.3/5

Path selection, compatibility notes, semantic validation advice, and OOM/unsupported-operation troubleshooting are reasonably detailed. However, the supplied material does not fully reproduce dependencies, referenced files, or key APIs; silent fallback is acknowledged; and there is no statically verifiable CI or test coverage, so the score is conservatively capped.

Adaptability10 / 15 · 3.3/5

The audience, cuDF/cudf.pandas/dask-cuDF scenarios, 100K-row gate, and negative PyTorch case are reasonably clear. The description remains broad, explicit negative triggers are limited, Chinese-language interaction and mainland-China reachability are not addressed, and environment-fit evidence is thin, so points are deducted.

Convention10 / 15 · 3.3/5

The documentation provides compatibility, three implementation paths, memory guidance, troubleshooting, reference files, a release version, and license information; the benchmark also recommends refreshing maintenance evidence. It lacks recommended Instructions, Examples, and Purpose sections, and installation/dependency notes, changelog, maintenance ownership, and update path remain incomplete, so points are deducted.

Effectiveness6 / 15 · 2.0/5

The material directly covers core ETL, joins, aggregations, nulls, reshaping, time series, and multi-GPU tasks, with actionable code patterns. The committed benchmark and evaluation tasks indicate potential utility, but static review cannot verify correctness, performance gains, or runtime usability, so the score stays within the static ceiling.

Verifiability4 / 10 · 2.0/5

A pinned release, external documentation references, committed evaluation tasks, and a benchmark report provide some auditability. However, independently reproducible execution logs, CI workflows, test outputs, and cross-source corroboration are absent, warranting a low score.

Evidence confidence:Low Reviewed Jul 20, 2026 Reviewed revision 55f18499943e
Before you use it
  • Performance, correctness, and security metrics in the benchmark are claims contained in the supplied materials; this review did not execute cuDF, CUDA, dask-cuDF, or any evaluation.
  • Before use, verify GPU architecture, CUDA/driver, Python, cuDF versions, and availability of the referenced files, then compare pandas and cuDF results on representative data.
  • The skill does not specify sensitive-data handling, external WebFetch behavior, offline/network failure handling, or rollback; external documentation may be unreachable in mainland-China or restricted-network environments.
  • cudf.pandas may silently fall back to CPU for unsupported operations, so successful execution alone must not be treated as evidence of GPU acceleration.
See the full review method →

What does this skill do, and when should you use it?

This NVIDIA-authored guide helps implementers work with cuDF and dask-cuDF GPU DataFrames. It covers the cudf.pandas accelerator, explicit cuDF APIs, and dask-cuDF for datasets larger than one GPU's memory. Its scope includes ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, spilling, and pandas parity checks. The tracked cuDF release is 26.04, with stated CUDA, driver, and Python compatibility requirements.

It instructs users to read CSV or Parquet data, filter, join, group, reshape, process strings, and handle time-series workloads through cudf.pandas or explicit cuDF, then move to dask-cuDF for multi-GPU datasets. It provides examples for enabling spill, configuring RMM, profiling cudf.pandas fallbacks, and maintaining a pandas reference path. It recommends comparing shape, labels, null counts, ordering, representative values, and dtypes before claiming semantic parity.

  1. A pandas user wants GPU acceleration with minimal code changes and third-party pandas compatibility.
  2. An implementer is migrating named DataFrame code or optimizing a visible ETL hot path with explicit control.
  3. A workload exceeds single-GPU memory and needs dask-cuDF with spill enabled.
  4. A pipeline contains joins, groupby, reshaping, nullable columns, or time-series logic that requires pandas comparison.
  5. A GPU pipeline encounters out-of-memory errors, allocator fragmentation, or repeated allocation overhead.

What are this skill's strengths and limitations?

Pros
  • Covers minimal-change acceleration, explicit API migration, and multi-GPU processing.
  • Gives concrete guidance on the 100K-row size gate, conversion boundaries, and semantic validation.
  • Addresses null handling, sort stability, floating-point differences, fallbacks, and OOM troubleshooting.
  • Includes copyable examples for cudf.pandas, dask-cuDF, spilling, and RMM.
Limitations
  • Below 100K rows, GPU transfer overhead may outweigh the benefit.
  • It requires compatible NVIDIA hardware, CUDA, drivers, and Python versions.
  • Some pandas operations may fall back to CPU or be unsupported, requiring reference-file workarounds and narrow compatibility boundaries.
  • The source provides no standalone test suite, benchmark results, or evidence of validation on specific platforms.

How do you install this skill?

Install the individual skill with the skills CLI:

npx skills add nvidia/skills --skill accelerated-computing-cudf --yes

The README does not provide commands for installing cuDF, CUDA, drivers, or Python; those runtime dependencies must be prepared through their respective documentation.

How do you use this skill?

After installing the skill, ask an Agent for a concrete task, for example: “Migrate this pandas ETL code to explicit cuDF and compare pandas and GPU results for nulls, ordering, and representative values.” Request cudf.pandas for minimal-change acceleration, explicit cuDF for controlled migrations, or dask-cuDF for data larger than GPU memory. The guide tracks cuDF release 26.04.

How does this skill compare with similar options?

The guide explicitly positions three paths: cudf.pandas for compatibility and minimal changes, explicit cuDF for control and hot-path work, and dask-cuDF for datasets that exceed GPU memory.

FAQ

Does this skill install cuDF automatically?
No. The source documents installation of the skill through the skills CLI, but does not document installation commands for cuDF or CUDA.
When is GPU acceleration a poor fit?
The guide sets 100K rows as a minimum size gate and says smaller workloads are usually dominated by GPU transfer overhead.
What should I do when cudf.pandas falls back to CPU?
Use the cudf.pandas profiler to identify CPU-heavy hot paths, then switch to explicit cuDF and consult api-patterns.md for known gaps and workarounds.
Does it guarantee exact pandas results?
No. It requires validation of nulls, ordering, aggregates, shape, labels, dtypes, and representative values, and notes that floating-point non-associativity can produce differences.

More skills from this repository

All from NVIDIA/skills

Data & Analysis ✓ NVIDIA · Official

cuPyNumeric Parallel Shard Loader

Builds processor-sized parallel loading paths from sharded on-disk data into distributed cuPyNumeric arrays.

Data & Analysis ✓ NVIDIA · Official

TAO VCN Sample Router

Routes visual-defect gap samples to eligible mining and synthetic-anomaly modules.

Data & Analysis ✓ NVIDIA · Official

TAO AOI Image Mining

Embed target and source images with one encoder, then mine deduplicated nearest-neighbour AOI images for augmentation.

Data & Analysis ✓ NVIDIA · Official

TAO VCN Classification Gap Analysis

Find the weakest VCN classification samples and turn them into augmentation targets.

Dev & Engineering ✓ NVIDIA · Official

NeMo MBridge GPU Memory Tuning

Diagnose GPU OOMs in Megatron Bridge training and reduce fragmentation, activation, and PEFT memory usage.

Finance & Investment Banking ✓ NVIDIA · Official

cuFOLIO Portfolio Optimization

GPU-accelerated CVaR portfolio optimization, backtesting, and rebalancing.

Dev & Engineering ✓ NVIDIA · Official

Megatron Bridge Activation Recompute

Trade extra compute for lower GPU activation memory in Megatron Bridge training.

Dev & Engineering ✓ NVIDIA · Official

DALI Dynamic Mode Assistant

Helps agents write, review, and migrate NVIDIA DALI imperative dynamic-mode code.

Dev & Engineering ✓ NVIDIA · Official

cuPyNumeric Migration Readiness

Assess whether NumPy code is ready to scale on GPUs before committing to a substantial cuPyNumeric port.

Automation & Ops ✓ NVIDIA · Official

VSS Video Embedding Deployment

Deploy and operate NVIDIA’s video embedding service for files, text, and live streams.

Dev & Engineering ✓ NVIDIA · Official

cuPyNumeric Installation Guide

Guides safe cuPyNumeric installation and verifies that it actually works.

Dev & Engineering ✓ NVIDIA · Official

Jetson Clock Customizer

Lock or cap Jetson CPU, GPU, and EMC clock behavior before flashing a BSP image.

Automation & Ops ✓ NVIDIA · Official

Jetson Health Snapshot

Read-only diagnostics that consolidate a live Jetson’s hardware, resource, and service state.

Automation & Ops ✓ NVIDIA · Official

Megatron-LM on SLURM

Launch, monitor, and troubleshoot multi-node Megatron-LM training on SLURM GPU clusters.

Automation & Ops ✓ NVIDIA · Official

VSS Multi-Camera 3D Detection and Tracking

Deploy multi-camera 3D perception with DeepStream and BEV Fusion.

Dev & Engineering ✓ NVIDIA · Official

NeMo MBridge CPU Offload

Configure, validate, and troubleshoot CPU offloading in Megatron Bridge to relieve GPU memory pressure.

Dev & Engineering ✓ NVIDIA · Official

NeMo MBridge CUDA Graphs

Configure, validate, and benchmark CUDA Graph capture in Megatron Bridge to reduce host-driver overhead.

Automation & Ops ✓ NVIDIA · Official

TAO on SLURM

Run TAO training, evaluation, and inference on a remote GPU cluster through SSH, SLURM, and Pyxis/Enroot.

Dev & Engineering ✓ NVIDIA · Official

TAO Platform Execution SDK

Submit, monitor, and manage NVIDIA TAO GPU training jobs across supported platforms.

Automation & Ops ✓ NVIDIA · Official

TAO NVIDIA GPU Host Setup

Checks and standardizes NVIDIA drivers, CUDA, and container runtime prerequisites for TAO GPU hosts.

Related skills