Automation & Ops ✓ Google · Official gkekubernetescompute-classesnode-autoscalinggpu-workloadsspot-vmsworkload-governance

GKE ComputeClasses Advisor

Configure, optimize, and troubleshoot GKE ComputeClasses for cost, capacity, and workload placement.

FollowSkills review · FSRS-2.0
Not recommended
57/ 100 5-point scale 2.9 / 5
Trust18 / 25 · 3.6/5

The skill requires placeholders, an “EXAMPLE TEMPLATE - DO NOT DEPLOY” label, collection of critical environment context, and refusal to disable node security or embed service-account keys. It also addresses access governance, prompt-injection resistance, Spot eviction, and operational safeguards. It does not fully disclose data flows, require explicit execution confirmation, define rollback procedures, or provide dependency-security review, so points are deducted.

Reliability8 / 20 · 2.0/5

The material contains detailed schema constraints, edge cases, scheduling explanations, and diagnostic errors in the logging script, making the happy path plausible. No executable tests, CI coverage, or complete contents of the referenced files are provided, and several version- and behavior-specific claims cannot be reproduced from static evidence, so the score remains conservative under the static cap.

Adaptability11 / 15 · 3.7/5

The audience, use cases, and non-fit boundary are clear, covering Spot, GPU/TPU, machine families, zones, governance, and troubleshooting, with additional Pod, PV, and Autopilot boundaries. Chinese-language guidance, explicit input/output contracts, and positive/negative trigger examples are missing; reachability of required GKE/GCP services from mainland China is not addressed, so points are deducted.

Convention9 / 15 · 3.0/5

The skill has layered sections, an index, quick actions, example assets, limitation notes, and repository-level installation guidance. The README supplies Apache-2.0 licensing, installation, maintenance issue paths, and basic provenance through the official organization and revision. Skill-level versioning, changelog, named maintenance ownership, dependency-version matrix, FAQs, and the indexed reference files are absent from the supplied evidence, so points are deducted.

Effectiveness7 / 15 · 2.3/5

The content directly targets ComputeClass configuration, optimization, and troubleshooting, with templates for several representative scenarios and an actionable log-monitoring script. Without execution results, real task validation, or comparison with alternatives, direct usability on the target GKE versions and regions is unverified; the static ceiling therefore limits this dimension to 7.

Verifiability4 / 10 · 2.0/5

The supplied files include revision-scoped YAML and shell material, version conditions, error messages, and reference paths, providing some auditability. There is no committed test suite, CI result, third-party execution evidence, or cross-source corroboration; external URLs appear only in comments and this review cannot add evidence, so the score is low.

Evidence confidence:Low Reviewed Jul 20, 2026 Reviewed revision 513a7a51e85f
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
Before you use it
  • Review every YAML against the target GKE version, region, quotas, CUDs/Reservations, and actual CRD schema before use; the documentation explicitly says not to deploy the example templates directly.
  • The logging script reads project logs and can append to a user-selected path; verify gcloud/kubectl identity, least-privilege logging access, sensitivity of log contents, and output-file permissions first.
  • The skill depends on GKE, gcloud, kubectl, jq, and version-specific features; mainland-China reachability, a compatibility matrix, and executable test evidence are not supplied.
See the full review method →

What does this skill do, and when should you use it?

This skill is for engineers managing GKE ComputeClasses, including Spot VM fallback, accelerator and machine-family selection, zone targeting, and node-pool auto-creation. It provides configuration guidance, YAML template rules, scheduling diagnostics, and access-governance advice. It also covers stateful PVs, reservations, Autopilot mode, Karpenter migration, and node-security boundaries. It does not cover cluster-level Node Auto Provisioning configuration or general GKE cluster creation.

Guides an agent in designing ComputeClasses around cost, accelerators, machine families, zones, CUDs/Reservations, and workload constraints; produces constrained YAML examples with placeholders; explains priority ordering, Spot and reservation fallback, node-pool auto-creation, GPU/Spot taints, PV topology, and Pod scheduling; directs users to check the ComputeClass CRD with kubectl and use the repository's reference documents, log script, and YAML assets for troubleshooting and governance.

  1. A GKE platform engineer needs Spot VMs for horizontally scalable stateless workloads with on-demand fallback when capacity is unavailable.
  2. A machine-learning team needs workloads placed on specific accelerators such as L4, H100, or v5p.
  3. An infrastructure team needs a reliable priority fallback chain based on CUDs, Reservations, machine families, and zones.
  4. An operator needs to diagnose Pending or noScaleUp Pods caused by GPU or Spot taints, selector conflicts, zonal PV constraints, or storage topology.
  5. A platform administrator needs to separately control who can modify ComputeClass objects and which workloads can consume them.

What are this skill's strengths and limitations?

Pros
  • Covers ComputeClass configuration, cost optimization, scheduling diagnostics, and access governance.
  • Addresses practical failure modes involving GPU and Spot taints, disk generations, reservation fallback, and scarce large machine shapes.
  • Provides entry points to reference documentation, YAML assets, and a log script.
  • Uses CUDs/Reservations, Pod requests, and zone constraints to shape more reliable fallback strategies.
Limitations
  • Recommendations depend on the GKE version, region, available capacity, workload constraints, and existing reservations; generic templates are not production-ready by themselves.
  • The source provides guidance and examples but does not document a test suite or actual cluster validation results.
  • Users must provide environment context and perform validation and troubleshooting with kubectl or repository assets.

How do you install this skill?

Run npx skills add google/skills, then select GKE ComputeClasses during installation. The README does not document a specific installation directory or a separate single-skill installation command.

How do you use this skill?

After installation, use a prompt such as Help me configure a GKE ComputeClass for GPU workloads with on-demand fallback and troubleshoot pending pods. Supply existing CUDs/Reservations, workload statefulness, cluster pools, target region/zones, and Pod requests for more specific guidance. Do not use it for cluster-level Node Auto Provisioning configuration or general GKE cluster creation.

FAQ

Can this skill create a GKE cluster or configure cluster-level Node Auto Provisioning?
No. It focuses on GKE ComputeClasses; the source explicitly excludes cluster creation and cluster-level Node Auto Provisioning configuration.
Will using it guarantee lower costs?
No. It can guide Spot, On-Demand, DWS, Reservation, and CUD strategies, but actual cost depends on capacity, workload constraints, and committed resources.
Why might a Pod remain Pending after selecting a GPU or Spot ComputeClass?
Common causes include missing GPU or Spot taint tolerations, conflicting hard node selectors, zonal resource shortages, or Pod requests that do not fit the available node shapes.
Who should adopt this skill?
It is best suited to platform and operations teams managing GKE compute placement, autoscaling, GPU/TPU workloads, cost policies, or ComputeClass access governance.

More skills from this repository

All from google/skills

Automation & Ops ✓ Google · Official

GKE Cluster Autoscaler Guide

Configure, optimize, and troubleshoot automatic GKE node scaling.

Automation & Ops ✓ Google · Official

GKE Workload Autoscaling

Configure manual scaling, HPA, and VPA for GKE workloads while improving resource requests from observed usage.

Automation & Ops ✓ Google · Official

GKE Cost Optimization

Reduce GKE spending through rightsizing, quotas, Spot VMs, and capacity planning.

Automation & Ops ✓ Google · Official

GKE AI Inference Deployment Assistant

Deploy, tune, and autoscale GPU/TPU AI inference services on GKE.

Automation & Ops ✓ Google · Official

GKE Batch & HPC Workloads

Configure and run queued, parallel, and high-performance workloads on GKE.

Automation & Ops ✓ Google · Official

AI Workload Migration to GKE Inference

Move existing AI inference workloads to self-hosted inference on Google Kubernetes Engine.

Automation & Ops ✓ Google · Official

GKE Cluster Creation Advisor

Plan, provision, and audit GKE clusters against production best practices.

Automation & Ops ✓ Google · Official

GKE Production Golden Path

Set production-oriented GKE defaults, readiness checks, and decision guardrails for cluster design.

Automation & Ops ✓ Google · Official

GKE Multi-Tenancy Planner

Plan isolation, quotas, access control, and cost attribution for teams sharing a GKE cluster.

Automation & Ops ✓ Google · Official

GKE Production Readiness Review

Assess whether GKE clusters and workloads are ready for production.

Automation & Ops ✓ Google · Official

GKE App Onboarding Assistant

Containerize and deploy an application to Google Kubernetes Engine for the first time.

Automation & Ops ✓ Google · Official

GKE Basics Navigator

A focused entry point for discovering GKE cluster needs and routing each task to the right specialist skill.

Automation & Ops ✓ Google · Official

GKE Observability Configuration

Configure GKE logging, monitoring, and Prometheus metrics.

Automation & Ops ✓ Google · Official

GKE Storage Manager

Configure persistent disks, Filestore, and GCS FUSE storage for GKE workloads.

Automation & Ops ✓ Google · Official

GKE Workload Reliability Guide

Reduce GKE workload disruption with high-availability configuration.

Automation & Ops ✓ Google · Official

GKE Backup & Disaster Recovery

Configure Backup for GKE policies and restore workflows for stateful workloads.

Automation & Ops ✓ Google · Official

GKE GPU/TPU Disruption Diagnosis

Diagnose and mitigate GKE GPU/TPU node disruptions caused by host maintenance.

Automation & Ops ✓ Google · Official

GKE Workload Security

Audit and harden workload-level security controls for GKE applications and namespaces.

Automation & Ops ✓ Google · Official

GKE Platform Security Hardening

Helps platform teams plan and apply cluster-level security hardening for Google Kubernetes Engine.

Automation & Ops ✓ Google · Official

GKE Enterprise RAG Search Architect

Designs and validates enterprise RAG search systems built on GKE and AlloyDB.

Related skills