CtrlK
BlogDocsLog inGet started
Tessl Logo

lambda-labs

On-demand GPU cloud instances for ML training.

52

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/lambda-labs/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

76%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, highly actionable reference: concrete commands, exact endpoints, real pricing, and complete request objects that a reader can execute immediately. Its main weaknesses are missing validation before the destructive terminate operation (capping workflow clarity) and a body that carries material the provided reference files already cover, rather than strictly staying an overview.

Suggestions

Add a verification checkpoint before terminating instances (e.g. run list_instances to confirm the target instance ID and confirm no unsaved data on local SSD, since termination is irreversible and there is no auto-stop), and mention checking filesystem attachment before launch since filesystems can only be attached at launch time.

Move the multi-node srun/torchrun details and the Common issues table into references/advanced-usage.md and references/troubleshooting.md respectively, keeping the body's coverage to one-line pointers, to remove duplication with the bundle files.

Trim sections that restate knowledge Claude already has — SSH tunneling basics, ssh-import-id, Jupyter launch, nvidia-smi verification — to free context for the skill's genuinely unique content (1-Click Clusters, filesystem semantics, pricing).

DimensionReasoningScore

Conciseness

The body is dominated by lean tables and code blocks with almost no padded prose, matching 'Efficient; minor instances of over-explanation that could be trimmed' — e.g. "SSH tunneling" basics ("ssh -L 8888:localhost:8888") and "Verify installation" snippets that assume knowledge Claude already has, plus GPU/pricing details restated between the features list and the Available GPUs table. Not a 5 because those known-concept sections and duplicated GPU listings could be cut without losing value.

4 / 5

Actionability

Guidance is copy-paste ready across the board: exact API endpoints ("https://cloud.lambdalabs.com/api/v1/instance-operations/launch"), full curl commands with auth, complete Python request objects (LaunchInstanceRequest with region_name, instance_type_name, ssh_key_names), concrete torchrun/srun invocations, and real prices per GPU — matching 'Fully executable; copy-paste ready code or commands; specific examples cover the common cases'. Not a 4 because the few placeholders (<INSTANCE-IP>, MyModel()) are standard template variables rather than gaps in executable detail.

5 / 5

Workflow Clarity

Sequences are clearly ordered (account setup → launch → SSH → train; numbered Workflow 1/2 with step comments) but validation checkpoints are absent for the destructive terminate-instance operation — the curl and Python terminate examples fire on an instance ID with no verify-before-terminate step — which caps this dimension at 3 per the batch/destructive-operations guideline. Not a 4 because the cap applies; not a 2 because the sequences themselves are coherent and well-sequenced with concrete commands at each step.

3 / 5

Progressive Disclosure

Structure is good: a well-organized overview body with clear section headers and a References section pointing to two real one-level-deep files ([Advanced Usage](references/advanced-usage.md), [Troubleshooting](references/troubleshooting.md)), matching 'Good structure; most content is appropriately placed; references mostly clear'. Not a 5 because the ~550-line body still inlines substantial material that duplicates the reference files (multi-node srun/torchrun content also lives in advanced-usage.md, and a Common issues table overlaps troubleshooting.md) — some of it should be pushed down into the bundles.

4 / 5

Total

16

/

20

Passed

Description

45%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a terse domain label rather than a capability+trigger statement: it identifies what the skill is about but names no actions and gives no 'Use when...' guidance, and it omits the provider name that would distinguish it from sibling GPU-cloud skills. It avoids fluff but undersells the skill's concrete coverage (clusters, filesystems, pricing, API/CLI).

Suggestions

Add an explicit trigger clause, e.g. "Use when the user needs to rent dedicated GPU instances (H100, B200, A100) or launch multi-node training clusters on Lambda Labs, or mentions Lambda, GPU cloud, or per-minute GPU pricing."

State concrete capabilities in third person — "Launch and SSH into on-demand GPU instances, manage persistent filesystems, and run 16-512 GPU Slurm clusters" — instead of the bare noun phrase "On-demand GPU cloud instances for ML training."

Include the provider name "Lambda Labs" in the description so it is distinguishable from alternatives (Modal, RunPod, Vast.ai, SkyPilot) that the skill itself positions against.

DimensionReasoningScore

Specificity

The description is a noun phrase — "On-demand GPU cloud instances for ML training" — that names the domain but contains no verbs or concrete actions at all, matching the anchor 'Names the domain but actions are minimal or generic'. It is not a 1 because the domain (GPU cloud, ML training) is concretely identified, and not a 3 because no capability (launch, connect, train, scale) is actually stated.

2 / 5

Completeness

A clear 'what' is present ("On-demand GPU cloud instances for ML training") but there is no 'when' clause whatsoever, matching 'Has a clear what but when is missing or only weakly implied' and triggering the guideline that a missing 'Use when...' clause caps completeness at 3. Not a 4 because the when is entirely absent rather than merely imprecise.

3 / 5

Trigger Term Quality

"GPU cloud", "instances", and "ML training" are natural phrases users would say, but common variations and synonyms users actually use — rent GPU, H100/A100, fine-tuning, distributed training, serverless GPUs — are absent, matching 'Some relevant keywords but missing common variations or synonyms'. Not a 4 because keyword coverage is thin relative to the skill's breadth (clusters, filesystems, pricing are never mentioned).

3 / 5

Distinctiveness Conflict Risk

"GPU cloud instances for ML training" is somewhat specific but the provider (Lambda Labs) is never named in the description, so it would equally plausibly trigger for competitor skills the body itself lists (Modal, RunPod, Vast.ai, SkyPilot), matching 'Somewhat specific but could still overlap with similar skills'. Not a 4 because without the provider name the overlap risk with other GPU-cloud skills is real, not minor.

3 / 5

Total

11

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (550 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.