Claude Skills for a DevOps / SRE — Page 4
NeMo AutoModel Recipe Development
Build, modify, and validate NeMo AutoModel training and evaluation recipes.
NeMo MBridge CPU Offload
Configure, validate, and troubleshoot CPU offloading in Megatron Bridge to relieve GPU memory pressure.
NeMo MBridge CUDA Graphs
Configure, validate, and benchmark CUDA Graph capture in Megatron Bridge to reduce host-driver overhead.
NeMo Relay Call Instrumentation
Wrap tool and LLM calls with Relay lifecycle events, middleware, and guardrails.
NeMo Relay Context Isolation
Keeps NeMo Relay scope stacks independent across concurrent requests and async workflows while preserving ancestry propagation.
NeMo Relay Typed Wrappers & Codecs
Add typed boundaries to NeMo Relay integrations while preserving predictable JSON middleware semantics and caller-visible behavior.
Omniverse Realtime Viewer
Routes USD viewer requests to the right architecture and guides rendering, interaction, UI, and validation.
NVIDIA NuRec Neural Reconstruction Router
Routes NuRec requests to the right upstream workflow skill.
RAG Performance Benchmark
Benchmark a deployed NVIDIA RAG Blueprint server and expose latency, throughput, and bottleneck behavior from one YAML config.
TAO on SLURM
Run TAO training, evaluation, and inference on a remote GPU cluster through SSH, SLURM, and Pyxis/Enroot.
TAO Platform Execution SDK
Submit, monitor, and manage NVIDIA TAO GPU training jobs across supported platforms.
TAO NVIDIA GPU Host Setup
Checks and standardizes NVIDIA drivers, CUDA, and container runtime prerequisites for TAO GPU hosts.
VSS Video Summarizer
Creates timestamped narrative summaries of recorded videos through LVS, with a VLM fallback.
AutoMagicCalib Calibration Stack Launcher
Deploy the AutoMagicCalib microservice and web UI with Docker Compose for a ready-to-use camera calibration stack.
CUDA-Q Quantum Onboarding
Guides developers from CUDA-Q installation to quantum kernels, GPU simulation, and real QPU execution.
cuOpt Numerical Optimization API
Guide agents through GPU-accelerated LP, MILP, and QP modeling with cuOpt
Clinical ASR Flywheel: Environment Setup
Validate that a clinical ASR evaluation environment can complete a TTS-to-ASR round trip through NVIDIA-hosted speech services.
DOCA Bench Benchmarking Skill
Measure DOCA library throughput, latency, and bandwidth reproducibly on real NVIDIA networking hardware.
DOCA CollectX Telemetry Deployment
Deploy, operate, and debug CollectX telemetry collectors so counters reach downstream exporters.
DOCA DMA Development Guide
Guides hands-on DOCA DMA memory-copy development on BlueField and ConnectX systems.
DOCA Flow Performance Measurement
Guides reproducible, defensible measurements of host-side and DPU-CPU DOCA Flow control-plane rule rates.
DOCA Flow Tune
Guides engineers through snapshotting, analyzing, and optimizing live or captured DOCA Flow pipelines.
DOCA Programming Guide
Guides developers through building, testing, and debugging library-agnostic DOCA applications from shipped samples.
NVIDIA DOCA Public Knowledge Map
Routes agents to authoritative DOCA docs, samples, versions, and install paths.