Google Cloud Reliability Architect
Evaluate and improve Google Cloud workload reliability using the Well-Architected Framework.
The skill only supplies Google Cloud reliability assessment questions, principles, and a checklist; it executes no commands, accesses no credentials, changes no resources, and exposes no data-exfiltration path. It also links the Google Cloud Well-Architected Framework as grounding material. Six points are deducted because sensitive-data handling, user confirmation, data-flow disclosure, rollback boundaries, and dependency security are not specified.
The objective, principles, questions, and validation checklist are broadly consistent, and there are no scripts or runtime dependencies. Eight points are awarded because there is no test coverage, abnormal-input handling, failure feedback, or key-path reproduction evidence, and this review did not execute the skill; the static calibration cap therefore applies.
The intended audience and scenario are clear: reliability assessment for Google Cloud workloads across design, deployment, and operations. Five points are deducted because invocation triggers, non-fit boundaries, input/output formats, Chinese-language support, and mainland-China network reachability are not defined.
The documentation is well organized with metadata, overview, principles, product examples, assessment questions, and a validation checklist. Repository context supplies installation, Apache-2.0 licensing, contribution, and support paths. Five points are deducted for missing skill-specific versioning, changelog, named maintenance responsibility, dependency notes, output examples, and troubleshooting guidance.
The questions and checklist can support an initial reliability review and cover SLOs, redundancy, scaling, observability, graceful degradation, recovery testing, and postmortems. Six points are awarded because the material is generic and does not define output format, prioritization, evidence requirements, or complete architecture-specific results; static review also provides no execution evidence of directly usable outputs.
Each core principle includes a Google Cloud documentation path, providing some primary-source traceability. Four points are awarded because there are no committed tests, CI coverage, representative assessment outputs, or cross-source corroboration; this conclusion is based only on static source review.
- This is guidance based on a framework, not a completed architecture audit or reliability guarantee.
- Before relying on it, provide workload-specific SLOs, RTO/RPO targets, data sensitivity, permission boundaries, and recovery evidence.
- The grounding links point to Google Cloud documentation; confirm that the target users' network environment can reach those sources or provide accessible local copies.
What does this skill do, and when should you use it?
This skill focuses on the Reliability pillar of the Google Cloud Well-Architected Framework. It provides guidance for reliability, resilience, availability, redundancy, fault tolerance, and disaster recovery in Google Cloud workloads. Its coverage includes user-focused SLIs and SLOs, resource redundancy, horizontal scalability, observability, graceful degradation, recovery testing, data-loss recovery, and blameless postmortems. It also uses Google Cloud product examples and a validation checklist to support architecture reviews.
Generates guidance from the embedded reliability principles and recommendations for Google Cloud workloads; asks assessment questions about reliability targets, redundancy, scaling, monitoring, alerting, graceful degradation, failure recovery, data recovery, and postmortems; and applies a validation checklist covering SLIs/SLOs, cross-zone or cross-region redundancy, autoscaling, health checks, backups, circuit breakers, retries, chaos practices, and postmortem processes.
- A cloud architect designing a new Google Cloud workload needs guidance on availability, redundancy, and disaster recovery.
- A platform engineering team reviewing an existing system needs to identify single points of failure, scalability risks, and failover gaps.
- An SRE team defining SLOs, error budgets, monitoring, alerts, and user-experience measures needs a reliability framework.
- An operations team preparing regional failover, release rollback, or data-recovery exercises needs a readiness review.
- An engineering leader conducting an incident review needs a structured approach to root-cause analysis and recurrence prevention.
What are this skill's strengths and limitations?
- Covers reliability design, operations, recovery testing, and organizational learning.
- Includes concrete assessment questions and a validation checklist for architecture reviews.
- Provides examples across compute, networking, storage, databases, operations, and disaster recovery.
- Uses the Reliability pillar of the Google Cloud Well-Architected Framework as its basis.
- Focused on Google Cloud workloads rather than cloud reliability in general.
- The source provides no test suite, automated assessment tool, or implementation scripts.
- The source does not specify a platform compatibility matrix, permission requirements, or operating cost.
- It requires adaptation to the workload’s architecture, business objectives, and organizational constraints; it does not replace live recovery exercises.
How do you install this skill?
Use the command provided in the repository README: npx skills add google/skills. During installation, select skills/cloud/google-cloud-waf-reliability. The README does not specify a more precise destination folder or a separate installation command for this skill.
How do you use this skill?
After installation, ask for a Google Cloud workload reliability assessment, for example: “Evaluate this Google Cloud architecture’s reliability and disaster recovery readiness, and identify improvements using the validation checklist.” Provide architecture details, SLOs, redundancy, observability, backup, and recovery information for more targeted guidance.