所有 Skills

找到 8642 个 Skills

Skills 列表

loopy

loopy

3Kagent-workflows

Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication. Use when a user asks to analyze code or coding threads for recurring work, find a published loop, interview them to turn a goal into a bounded loop, review a loop for weak checks or unsafe authority, execute a loop with an evidence receipt, learn from completed runs, save or reuse a project loop, or validate and submit a loop to Loop Library.

forward-future avatarforward-future
获取
launch-nemo-rl

launch-nemo-rl

3Kbackend-api

Playbook for launching, monitoring, stopping, and debugging NeMo-RL recipes on a Kubernetes cluster via the nrl-k8s CLI. Covers ephemeral vs long-lived RayCluster modes, iterating on runs, and debugging hung or failed training jobs.

nvidia avatarnvidia
获取
digital-health-clinical-asr-finetune

digital-health-clinical-asr-finetune

3Kbackend-api

Stage 4 of the Clinical ASR Flywheel. Use when priority KER is above 0.3 to run stock NeMo SFT on Parakeet TDT v2 and offline cycle N+1 re-eval. NOT for generic word boosting (use /finetune-asr).

nvidia avatarnvidia
获取
digital-health-clinical-asr-eval

digital-health-clinical-asr-eval

3Kbackend-api

Stage 3 of Clinical ASR Flywheel. Score a NeMo manifest, produce the five-section KER leaderboard (by-ipa_source diagnostic). Not for ASR auth (/riva-asr).

nvidia avatarnvidia
获取
earth2studio-discover

earth2studio-discover

3Kresearch-knowledge

Find Earth2Studio models, data sources, and examples for a weather/climate use case. Do NOT use for writing inference code, downloading data, or installation.

nvidia avatarnvidia
获取
earth2studio-deterministic-forecast

earth2studio-deterministic-forecast

3Kagent-workflows

Build deterministic forecast scripts with Earth2Studio (model, data source, IO, inference). Do NOT use for ensemble, diagnostics, data-only fetch, or install.

nvidia avatarnvidia
获取
earth2studio-data-fetch

earth2studio-data-fetch

3Kprompting-reasoning

Fetch weather/climate data via Earth2Studio data sources for specific variables and times. Do NOT use for inference pipelines, model discovery, or installation.

nvidia avatarnvidia
获取
nemo-mbridge-recipe-recommender

nemo-mbridge-recipe-recommender

3Ktesting-qa

Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal. Use when selecting a starting recipe, comparing library and benchmark configs, resizing parallelism for a GPU allocation, or distinguishing convergence changes, semantics-preserving execution tuning, and benchmark-only shortcuts.

nvidia avatarnvidia
获取
nemo-mbridge-perf-cpu-offloading

nemo-mbridge-perf-cpu-offloading

3Ktesting-qa

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

nvidia avatarnvidia
获取
vss-generate-video-report

vss-generate-video-report

3Kresearch-knowledge

Use this skill when producing a VSS analysis report — Mode A per-clip VLM, Mode B incident-range via video-analytics. Not for standalone video summarization, real-time alerts or ad-hoc Q&A.

nvidia avatarnvidia
获取
nemo-mbridge-perf-parallelism-strategies

nemo-mbridge-perf-parallelism-strategies

3Kbackend-api

Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.

nvidia avatarnvidia
获取
nemo-mbridge-multi-node-slurm

nemo-mbridge-multi-node-slurm

3Ktesting-qa

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.

nvidia avatarnvidia
获取
nemo-mbridge-perf-moe-vlm-training

nemo-mbridge-perf-moe-vlm-training

3Kagent-workflows

Practical guidance for training MoE VLMs in Megatron Bridge. Compares FSDP and 3D-parallel approaches, using rounded lessons from Qwen3-VL, Qwen3-Next, and other multimodal experiments.

nvidia avatarnvidia
获取
nemo-mbridge-perf-activation-recompute

nemo-mbridge-perf-activation-recompute

3Kbackend-api

Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. Use for activation memory OOMs or regressions involving recompute_granularity, recompute_num_layers, recompute_modules, recompute_method, selective recompute, full recompute, or activation checkpointing.

nvidia avatarnvidia
获取
nemo-mbridge-mlm-bridge-training

nemo-mbridge-mlm-bridge-training

3Ktesting-qa

Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. Covers correlation testing, available recipes, and multi-GPU examples.

nvidia avatarnvidia
获取
tilegym-adding-cutile-kernel

tilegym-adding-cutile-kernel

3Kagent-workflows

Add a new cuTile GPU kernel operator to TileGym. Covers dispatch registration in ops.py, cuTile backend implementation, __init__.py exports, test creation, and benchmark in tests/benchmark. Use when adding, creating, or implementing a new cuTile operator/kernel in TileGym, or when asking how to register a new cuTile op.

nvidia avatarnvidia
获取
nemo-mbridge-perf-moe-long-context

nemo-mbridge-perf-moe-long-context

3Kresearch-knowledge

Long-context MoE training guidance for Megatron Bridge. Covers CP sizing, selective recompute, dispatcher choices, and practical patterns from DSV3, Qwen3, and Qwen3-Next long-context experiments.

nvidia avatarnvidia
获取
nemo-mbridge-perf-expert-parallel-overlap

nemo-mbridge-perf-expert-parallel-overlap

3Ktesting-qa

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as DeepEP and HybridEP.

nvidia avatarnvidia
获取
vss-generate-video-calibration

vss-generate-video-calibration

3Kbackend-api

Use to run AutoMagicCalib on local MP4s, RTSP, or the bundled sample dataset, and to deploy vss-auto-calibration when needed. Do not use for non-AMC calibration or runtime analytics.

nvidia avatarnvidia
获取
dicom-series-preflight

dicom-series-preflight

3Kagent-workflows

Used for header-only preflight of one DICOM series folder before conversion or inference. Not for de-identification or clinical clearance.

nvidia avatarnvidia
获取
nemo-mbridge-perf-moe-optimization-workflow

nemo-mbridge-perf-moe-optimization-workflow

3Kbackend-api

Evidence-gated workflow for MoE performance optimization in Megatron Bridge. Covers measurement contracts, the Three Walls framework, parallel folding, profiling, matched A/B tuning, and final validation.

nvidia avatarnvidia
获取
nemo-mbridge-perf-hierarchical-context-parallel

nemo-mbridge-perf-hierarchical-context-parallel

3Ktesting-qa

Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

nvidia avatarnvidia
获取
nemo-mbridge-perf-sequence-packing

nemo-mbridge-perf-sequence-packing

3Ktesting-qa

Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.

nvidia avatarnvidia
获取
nemo-mbridge-perf-moe-hardware-configs

nemo-mbridge-perf-moe-hardware-configs

3Kresearch-knowledge

Representative, point-in-time MoE training playbooks by hardware and model family. Use them as candidate seeds, then revalidate the exact runtime, semantics, topology, and steady-state throughput.

nvidia avatarnvidia
获取
想按分类查看?试试 /category/writing-content.