DevOps 与云

部署、CI/CD、云平台和基础设施

358 個 Skills 可用

Skills 列表

tao-run-automl

tao-run-automl

3.1Kdevops-cloud

Run AutoML / hyperparameter optimization (HPO) for NVIDIA TAO networks using AutoMLRunner. Handles algorithm

nvidia avatarnvidia
獲取
tao-train-image-classification

tao-train-image-classification

3.1Kdevops-cloud

PyTorch-based TAO image classification. Supports a wide range of backbones (FAN, EfficientNet, ResNet, etc.)

nvidia avatarnvidia
獲取
tao-train-ocrnet

tao-train-ocrnet

3.1Kdevops-cloud

OCRNet for scene text recognition. Recognizes text content from cropped text-region images and supports CTC

nvidia avatarnvidia
獲取
tao-train-optical-inspection

tao-train-optical-inspection

3.1Kdevops-cloud

Optical Inspection for defect detection using Siamese networks. Compares image pairs to detect manufacturing

nvidia avatarnvidia
獲取
google-agents-cli-scaffold

google-agents-cli-scaffold

3.1Kdevops-cloud

>

google avatargoogle
獲取
google-agents-cli-observability

google-agents-cli-observability

3.1Kdevops-cloud

>

google avatargoogle
獲取
google-agents-cli-deploy

google-agents-cli-deploy

3.1Kdevops-cloud

>

google avatargoogle
獲取
google-agents-cli-publish

google-agents-cli-publish

3.1Kdevops-cloud

>

google avatargoogle
獲取
tao-setup-nvidia-gpu-host

tao-setup-nvidia-gpu-host

3.1Kdevops-cloud

Host setup for TAO GPU backends. Checks and, after user approval, installs NVIDIA driver branch 580, CUDA Toolkit 13.0, and NVIDIA Container Toolkit 1.19.0 for Docker/local-Docker and Kubernetes GPU worker hosts. The `--check-only` path works on any Linux distribution; `--install` automates debian-family (Ubuntu/Debian/Pop!_OS/Mint/Zorin/Raspbian), rhel-family (Fedora/RHEL/Rocky/AlmaLinux), and suse-family (openSUSE/SLES) hosts, and prints actionable manual-install steps for everything else. Use when the user asks to "set up an NVIDIA GPU host", "check TAO Docker GPU runtime", or prepare a Kubernetes GPU worker for TAO.

nvidia avatarnvidia
獲取
holoscan-install-container

holoscan-install-container

3Kdevops-cloud

Install Holoscan SDK via the NGC Docker container. Use for container-based installs; not for native apt/pip/Conda installs.

nvidia avatarnvidia
獲取
mcore-run-on-slurm

mcore-run-on-slurm

3Kdevops-cloud

How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-rank failure diagnosis.

nvidia avatarnvidia
獲取
vss-setup-video-analytics-api

vss-setup-video-analytics-api

3Kdevops-cloud

Use to deploy the vss-video-analytics-api REST service standalone (config-source, data-log bind, Elasticsearch, optional Kafka). Not for full warehouse deploy.

nvidia avatarnvidia
獲取
vss-deploy-video-embedding

vss-deploy-video-embedding

3Kdevops-cloud

Use this skill when deploying, operating, or integrating the VSS 3.2 GA RT-Embed Video Embedding microservice. Covers Docker Compose bring-up, GPU and storage prerequisites, the `/v1` REST API (file uploads, text and video embeddings, live RTSP streams, health and metrics), Redis/Kafka/OTel integration, common failure modes, and teardown.

nvidia avatarnvidia
獲取
vss-query-analytics

vss-query-analytics

3Kdevops-cloud

Use this skill when reading video-analytics metrics, incidents, alerts, and sensor data via the VA-MCP server (port 9901). Not for live VLM or incident-range narrative reports.

nvidia avatarnvidia
獲取
nemo-automodel-launcher-config

nemo-automodel-launcher-config

3Kdevops-cloud

Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.

nvidia avatarnvidia
獲取
dynamo-router-starter

dynamo-router-starter

3Kdevops-cloud

Start or patch Dynamo router modes and run router endpoint smoke checks. Use for round-robin, KV-aware, least-loaded, or device-aware routing setup; use recipe-runner for recipe deployment and troubleshoot for failure diagnosis.

nvidia avatarnvidia
獲取
dynamo-recipe-runner

dynamo-recipe-runner

3Kdevops-cloud

Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes. Use for model/backend/GPU/deployment-mode recipe bring-up; use router-starter for router-only mode work and troubleshoot for broken deployments.

nvidia avatarnvidia
獲取
dynamo-interconnect-check

dynamo-interconnect-check

3Kdevops-cloud

Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. Use after recipe-runner brings a deployment up (especially disagg/multi-node) to confirm the KV transport is correct; use troubleshoot for diagnosing already-failed pods.

nvidia avatarnvidia
獲取
nemotron-speech

nemotron-speech

3Kdevops-cloud

Routes NVIDIA Nemotron Speech (Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.

nvidia avatarnvidia
獲取
physical-ai-infrastructure-setup-and-resilient-scaling

physical-ai-infrastructure-setup-and-resilient-scaling

3Kdevops-cloud

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

nvidia avatarnvidia
獲取
cuopt-server-api-python

cuopt-server-api-python

2.8Kdevops-cloud

cuOpt REST server — start server, endpoints, Python/curl client examples. Use when the user is deploying or calling the REST API.

nvidia avatarnvidia
獲取
cuopt-install

cuopt-install

2.8Kdevops-cloud

Install cuOpt for Python, C, or server via pip, conda, or Docker; verify the install. For building cuOpt from source, see cuopt-developer.

nvidia avatarnvidia
獲取
aiq-deploy

aiq-deploy

2.8Kdevops-cloud

Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.

nvidia avatarnvidia
獲取
ctf-forensics

ctf-forensics

2.7Kdevops-cloud

Provides digital forensics and signal analysis techniques for CTF challenges. Use when analyzing disk images, memory dumps, event logs, network captures, cryptocurrency transactions, steganography, PDF analysis, Windows registry, Volatility, PCAP, Docker images, coredumps, side-channel power traces, DTMF audio spectrograms, packet timing analysis, CD audio disc images, or recovering deleted files and credentials.

ljagiello avatarljagiello
獲取