DevOps 与云
部署、CI/CD、云平台和基础设施
Skills 列表

google-cloud-solution-agentic-ai-borderless-data-lakehouse
Guides agents to discover requirements and design a governed, secure borderless open data lakehouse with agentic AI integration. Use when designing a multi-product architecture that connects data silos to AI agents, joining data across clouds, or running federated queries across Google Cloud and external data sources, including on-premises or other cloud providers. Don't use for simple single-cloud data warehouses or non-AI workloads.
google
gke-workload-scaling
Manages scaling for GKE workloads using HPA and VPA. Use when configuring Horizontal Pod Autoscaler (HPA), configuring Vertical Pod Autoscaler (VPA), or applying best practices for GKE workload autoscaling. Do not use for cluster-level autoscaling (Cluster Autoscaler), static cluster sizing, or configuring node-level machine styles directly.
google
cloud-logging-query-generation
Generates Logging Query Language (LQL) queries for Google Cloud Logging from natural language. Use this skill when you need to query log data or when you are debugging issues. You can filter log data by Google Cloud service. Don't use this skill to query other databases, such as SQL or Cloud Spanner.
google
google-cloud-solution-agentic-ai-bidirectional-streaming
Guides agents to interactively discover customer requirements for live, bidirectional multi-agent AI systems that process continuous streams of multimodal data for real-time technical guidance and safety monitoring. Generates a custom Google Cloud solution that uses opinionated best practices and architecture guidance. Use when users need agentic assistance to design and create a multi-product solution in the cloud for live bidirectional multimodal streaming workloads. Don't use for simple text-based chat applications or workloads without real-time streaming requirements.
google
vercel-cli
从命令行部署、管理、检查和排查 Vercel 项目。用于 Vercel 部署、构建失败、项目和团队、环境变量、域名和 DNS、日志、指标、Speed Insights、Core Web Vitals、请求追踪、用量、活动、告警、防火墙规则、缓存、定时任务、部署钩子、Edge Config、功能标志、集成、连接器、Blob 存储、容器注册表 (VCR)、微前端、滚动发布、自定义环境、Sandbox、agent/MCP 设置、OAuth 应用、预览访问、本地开发或 `vercel api` 回退。
vercel
gke-basics
负责管理 GKE 集群的核心创建与资源供应、凭据获取、Autopilot 与 Standard 模式的选择以及工作负载部署。适用于创建 GKE 集群、获取 kubectl 访问凭据、配置 Workload Identity,或在 Autopilot 和 Standard 模式之间做出决策等场景。请勿用于 GKE 专项网络配置(请使用 gke-networking)、高级安全加固(请使用 gke-platform-security 或 gke-workload-security)或集群升级(请使用 gke-upgrades)。
google
agent-platform-alert-configuration
基于 OpenTelemetry (OTel) 指标为 AI Agent 配置最佳实践告警策略。适用于分析、编写或部署告警策略,以监控 Agent 的延迟、错误率、Token 消耗及质量指标。注意:可靠性(Reliability)、成本(Cost)、安全防护(Safety)和系统安全(Security)告警使用通用 OTel 指标,支持跨运行时生效(如 Cloud Run、Vertex AI 等);而质量(Quality)告警依赖 Vertex AI Online Monitors,仅适用于 Vertex AI 部署环境。
google
bigtable-basics
协助完成 Bigtable 的实例与表资源创建、高性能 Schema 设计以及数据查询。适用于设计 Bigtable 行键(row key)、配置列族(column family)、编写 SQL 查询或客户端代码(Java、Go、Python),以及排查性能瓶颈和热点(hotspotting)问题。也适用于使用 gcloud 或 cbt CLI 配置 Bigtable 集群。请勿用于常规 Cloud SQL 管理。
google
agent-platform-prompt-management
用于在 Agent Platform 中管理和编排 Prompt。适用于创建、列出、查询、版本控制或删除 Agent Platform 中的托管 Prompt。请勿用于模型训练、模型部署到 Endpoint,或管理非 Agent Platform 的 Prompt。
google
cloud-run-basics
管理 Cloud Run 服务、任务和工作者池。当您需要部署响应 HTTP 请求的应用程序(服务)、运行事件触发或定时任务(任务),或处理始终在线的基于拉取的背景处理(工作者池)时使用。
google
zod-4
Zod 4 schema validation patterns. Trigger: When creating or updating Zod v4 schemas for validation/parsing (forms, request payloads, adapters), including v3 -> v4 migration patterns.
prowler-cloud
huggingface-llm-trainer
Train or fine-tune language and vision models using TRL (Transformer Reinforcement Learning) or Unsloth with Hugging Face Jobs infrastructure. Covers SFT, DPO, GRPO and reward modeling training methods, plus GGUF conversion for local deployment. Includes guidance on the TRL Jobs package, UV scripts with PEP 723 format, dataset preparation and validation, hardware selection, cost estimation, Trackio monitoring, Hub authentication, model selection/leaderboards and model persistence. Use for tasks involving cloud GPU training, GGUF conversion, or when users mention training on Hugging Face Jobs without local GPU setup.
huggingface
ml-pipeline
设计并实现生产级机器学习管道基础设施:使用 MLflow 或 Weights & Biases 配置实验跟踪,创建 Kubeflow 或 Airflow DAG 进行训练编排,使用 Feast 构建特征存储模式,部署模型注册表,并自动化重新训练和验证工作流。在构建 ML 管道、编排训练工作流、自动化模型生命周期、实现特征存储、管理实验跟踪系统、设置 DVC 进行数据版本控制、调整超参数或配置 MLOps 工具(如 Kubeflow、Airflow、MLflow 或 Prefect)时使用。
jeffallan
chaos-engineer
设计混沌实验、创建故障注入框架,并组织分布式系统的游戏日演练——生成运行手册、实验清单、回滚流程及事后复盘模板。在需要设计混沌实验、实施故障注入框架或开展游戏日演练时使用。适用于混沌实验、韧性测试、爆炸半径控制、游戏日、反脆弱系统、故障注入、Chaos Monkey、Litmus Chaos等场景。
jeffallan
salesforce-developer
编写和调试Apex代码,构建Lightning Web组件,优化SOQL查询,实现触发器、批处理作业、平台事件以及Salesforce平台上的集成。适用于开发Salesforce应用程序、定制CRM工作流、管理调控器限制、批量处理或设置Salesforce DX和CI/CD流水线。
jeffallan
sre-engineer
定义服务等级目标,创建错误预算策略,设计事件响应流程,开发容量模型,并为生产系统生成监控配置和自动化脚本。在定义SLI/SLO、管理错误预算、构建大规模可靠系统、事件管理、混沌工程、减少琐事或容量规划时使用。
jeffallan
cloud-architect
设计云架构,制定迁移计划,生成成本优化建议,并制定跨 AWS、Azure 和 GCP 的灾难恢复策略。适用于设计云架构、规划迁移或优化多云部署。调用场景包括卓越架构框架、成本优化、灾难恢复、着陆区、安全架构、无服务器设计。
jeffallan
microservices-architect
设计分布式系统架构,将单体应用分解为有界上下文服务,推荐通信模式,并生成服务边界图和弹性策略。适用于设计分布式系统、分解单体应用或实现微服务模式时——包括服务边界、领域驱动设计(DDD)、Saga模式、事件溯源、CQRS、服务网格或分布式追踪。
jeffallan
terraform-engineer
在 AWS、Azure 或 GCP 上使用 Terraform 实现基础设施即代码时使用。适用于模块开发(创建可复用模块、管理模块版本)、状态管理(迁移后端、导入现有资源、解决状态冲突)、提供商配置、多环境工作流以及基础设施测试。
jeffallan
devops-engineer
创建 Dockerfile、配置 CI/CD 流水线、编写 Kubernetes 清单、生成 Terraform/Pulumi 基础设施模板。处理部署自动化、GitOps 配置、事件响应手册和内部开发者平台工具。在设置 CI/CD 流水线、容器化应用、管理基础设施即代码、部署到 Kubernetes 集群、配置云平台、自动化发布或响应生产事件时使用。适用于流水线、Docker、Kubernetes、GitOps、Terraform、GitHub Actions、值班或平台工程。
jeffallan
kubernetes-specialist
在部署或管理 Kubernetes 工作负载时使用。调用以创建部署清单、配置 Pod 安全策略、设置服务账户、定义网络隔离规则、调试 Pod 崩溃、分析资源限制、检查容器日志或调整工作负载大小。用于 Helm Chart、RBAC 策略、NetworkPolicy、存储配置、性能优化、GitOps 流水线以及多集群管理。
jeffallan
mcp-builder
**所有MCP服务器工作的必读内容** - mcp-use框架的最佳实践和模式。 **在进行任何MCP服务器工作之前,请先阅读此内容**,包括: - 创建新的MCP服务器 - 修改现有的MCP服务器(添加/更新工具、资源、提示、组件) - 调试MCP服务器问题或错误 - 审查MCP服务器代码的质量、安全性或性能 - 回答关于MCP开发或mcp-use模式的问题 - 对server.tool()、server.resource()、server.prompt()或组件进行任何更改 此技能包含关键的架构决策、安全模式和常见陷阱。 在实现MCP功能之前,请务必查阅相关的参考文件。
mcp-use
chatgpt-app-builder
**所有MCP服务器工作的必读内容**——mcp-use框架最佳实践与模式。 **在进行任何MCP服务器工作之前,请先阅读本文**,包括: - 创建新的MCP服务器 - 修改现有MCP服务器(添加/更新工具、资源、提示、组件) - 调试MCP服务器问题或错误 - 审查MCP服务器代码的质量、安全性或性能 - 回答关于MCP开发或mcp-use模式的问题 - 对server.tool()、server.resource()、server.prompt()或组件进行任何更改 本技能包含关键的架构决策、安全模式和常见陷阱。 在实现MCP功能之前,请务必查阅相关的参考文件。
mcp-use
mcp-apps-builder
**所有MCP服务器工作的必读内容**——mcp-use框架最佳实践与模式。 **在进行任何MCP服务器工作之前,请先阅读本文**,包括: - 创建新的MCP服务器 - 修改现有的MCP服务器(添加/更新工具、资源、提示、组件) - 调试MCP服务器问题或错误 - 审查MCP服务器代码的质量、安全性或性能 - 回答关于MCP开发或mcp-use模式的问题 - 对server.tool()、server.resource()、server.prompt()或组件进行任何更改 本技能包含关键的架构决策、安全模式和常见陷阱。 在实现MCP功能之前,请务必查阅相关的参考文件。
mcp-use