CtrlK
BlogDocsLog inGet started
Tessl Logo

physical-ai-infrastructure-setup-and-resilient-scaling

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

69

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is physical-ai-infrastructure-setup-and-resilient-scaling in NVIDIA/skills

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, operationally precise body with strong gated workflows and clear component navigation. The main weaknesses are delegated (non-inline) executable detail and referenced bundle files that do not exist on disk to verify the progressive-disclosure structure.

Suggestions

Ship the referenced bundle (components/*/reference.md, scripts/preflight.sh, scripts/scan_transcript_secrets.py) so the progressive-disclosure references resolve to real files rather than dangling paths.

Inline one or two short executable examples per stage (e.g., a preflight invocation and a verify-hello check) instead of delegating all executable detail to component references.

Trim the 'Evaluation Prompts And Results' section and consolidate repeated gating language to tighten token efficiency.

DimensionReasoningScore

Conciseness

The body is efficient and assumes Claude's competence (no concept explanations of Kubernetes/OSMO), but the Evaluation Prompts section and some repeated gating language could be trimmed, so it sits just below the lean-every-token-earns-its-place anchor.

4 / 5

Actionability

Concrete commands and script names are present ('scripts/preflight.sh', 'osmo workflow query <id>', 'terraform apply', '/v1/health/ready', 'git rev-parse --show-toplevel'), but most executable detail is delegated to component reference files rather than given inline, leaving minor gaps.

4 / 5

Workflow Clarity

The 8-step Setup Flow is explicitly sequenced with mandatory gates ('stop on red preflight', 'Nothing else starts until the cluster gate is green'), a Verification Gates table, and clear feedback loops ('Failed terminal states require events and logs before retry', 'do not resubmit blindly'), matching the explicit-validation-and-feedback-loops anchor.

5 / 5

Progressive Disclosure

Structure is well-signaled with a component matrix table and a one-level-deep load-when model, but no bundle files are actually present on disk to verify the references, and the OSMO CLI component introduces a documented second level (agents/references/tests), so it falls short of the clean one-level anchor.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description with explicit what/when structure, comprehensive trigger keywords, and useful negative triggers that bound the skill's scope. The only minor gap is that some trigger terms are product jargon rather than natural user phrasing.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('set up, scale, validate, or harden') across concrete targets (MicroK8s/AKS clusters, inference endpoints, OSMO, workload readiness, failure recovery), giving comprehensive coverage rather than just naming the domain.

5 / 5

Completeness

Explicitly answers both 'what' (the infra stack and stages) and 'when' with a 'Use when the user wants to...' clause, plus concrete trigger keywords and explicit negative triggers, matching the clear-what-and-when-with-triggers anchor.

5 / 5

Trigger Term Quality

Good keyword coverage with natural domain terms ('physical ai infrastructure', 'resilient scaling', 'microk8s', 'azure aks', 'OSMO deploy'), but several are product names/jargon ('NVCF deployment', 'NIM Operator') rather than phrases a general user would naturally say, so it falls just short of the comprehensive-synonym anchor.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (NVIDIA physical AI infra for SDG) and includes explicit disambiguation ('Don't trigger for: OSMO log summarization or workload-only operations...'), minimizing conflict with adjacent skills.

5 / 5

Total

19

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

12

/

16

Passed

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.