Google Cloud Live Multimodal Streaming Architect
Design and deploy Google Cloud solutions for live, bidirectional multimodal streams.
The workflow requires user confirmation after decomposition, architecture, and implementation planning, and it calls for least privilege, TLS, OIDC, Cloud Armor, and human oversight. Points are deducted because retention and deletion of sensitive audio/video, isolation, key management, concrete IAM scopes, and rollback procedures are not specified.
The four-phase workflow is internally coherent and includes dry runs, connectivity checks, security checks, and verification scripts. Static files contain no executable tests, pinned dependencies, abnormal-input handling, or detailed diagnostic failure paths, so the score is limited and deductions apply.
The name, target scenario, discovery questions, and non-fit case of simple text chat are reasonably clear, covering modalities, latency, and network constraints. Chinese-language support is not stated, and the core documentation, platform, and services depend on Google-hosted overseas reachability without a mainland-China fallback.
The material has clear phases, supporting references, product mapping, design guidance, and an output template; the repository provides Apache-2.0 licensing, contribution guidance, and an active-development note. The skill lacks its own installation/dependency notes, compatibility versions, changelog, FAQ, known limitations, and explicit maintenance owner.
It can produce requirements analysis, product mapping, Mermaid architecture, IaC planning, deployment guidance, and validation reporting for the stated real-time multimodal use case, with a defined output template. Points are deducted because outputs depend on external retrieval and model generation, with no verified representative output, runnable IaC, or sufficiently developed alternative comparison.
The files cite specific Google Cloud, ADK, codelab, and template sources and require validation steps, providing some auditability. There is no committed test suite, CI coverage, execution log, or third-party reproduction evidence, so conclusions remain largely unexecuted and static.
- Do not apply generated Terraform, deployment commands, or IAM configuration directly to production; review permissions, regions, data flows, cost, quotas, and rollback first.
- Live audio/video may contain biometric or otherwise sensitive data; add data classification, retention/deletion, encryption-key, access-audit, and cross-border-transfer controls before deployment.
- Confirm that required Google documentation, APIs, models, and control-plane services are reachable from mainland China; otherwise prepare reachable mirrors, alternatives, or offline inputs.
- The skill retrieves multiple external resources without pinning versions; record the versions used and verify product names, APIs, and configurations before implementation.
What does this skill do, and when should you use it?
This skill guides agents through designing a tailored, multi-product Google Cloud solution for live, bidirectional, multimodal streaming workloads. It follows four phases: requirements discovery, solution design, implementation planning, and validation. It analyzes modalities, latency, safety monitoring, knowledge access, device constraints, and network limitations, then produces a technical decomposition, product mapping, Mermaid architecture diagram, deployment plan, and validation procedures. It is intended for real-time streaming agentic systems, not simple text chat applications or workloads without real-time streaming requirements.
Asks about audio, video, or text inputs, target latency, safety monitoring, knowledge bases, client devices, and network constraints; identifies workload components and cross-cloud, hybrid, or on-premises integrations; generates and iterates on a confirmed technical decomposition; uses specified Google Cloud documentation and repository references to map products, create a Mermaid architecture diagram, and draft recommendations; writes the design to solution-architecture-guide.md; generates infrastructure-as-code such as Terraform, deployment instructions, validation scripts or commands, and checks for connectivity, routing, and security policies.
- A cloud architect designing a Google Cloud system for live technical guidance over audio or video streams.
- An engineering team that needs real-time detection of hazards, operational risks, or incorrect steps in a video stream.
- A solution designer building a multi-agent system that accesses product documentation, knowledge bases, or schematic repositories for grounded guidance.
- A Google Cloud team planning deployment components involving Cloud Run, WebSockets, Gemini Live API, or ADK streaming tools.
What are this skill's strengths and limitations?
- Covers requirements discovery, architecture design, IaC planning, and validation in one workflow.
- Targets live bidirectional multimodal streaming and multi-agent cloud solutions rather than generic chat.
- Requires grounding design guidance in specified Google Cloud documentation and implementation resources.
- Includes review, iteration, and final validation checkpoints.
- Depends on access to specified Google Cloud documentation and implementation resources; offline use is not documented.
- Implementation requires projects, billing associations, APIs, and IAM permissions, but the provided text does not list exact permissions.
- The SKILL.md provides no test suite, fixed product inventory, or guarantee of deployment success.
- It generates Terraform, scripts, and commands, but the concrete implementation depends on the user's requirements and retrieved resources.
How do you install this skill?
In an Agent Skills-compatible client, run: npx skills add google/skills. During installation, select skills/cloud/google-cloud-solution-agentic-ai-bidirectional-streaming. The README does not specify a client-specific installation directory or version requirement.
How do you use this skill?
Place the skill folder where the client can discover Agent Skills, then use a prompt such as: "Design a Google Cloud solution for a multi-agent system that provides bidirectional voice guidance and real-time safety monitoring from a video stream, including deployment and validation plans." The workflow begins with requirements discovery and requests confirmation during decomposition, architecture, and implementation planning.