CtrlK
BlogDocsLog inGet started
Tessl Logo

nemo-guardrails

NVIDIA's runtime safety framework for LLM applications. Features jailbreak detection, input/output validation, fact-checking, hallucination detection, PII filtering, toxicity detection. Uses Colang 2.0 DSL for programmable rails. Production-ready, runs on T4 GPU.

55

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/nemo-guardrails/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-organized and code-dense with a useful troubleshooting section, but roughly half the workflows reference undefined helper functions, making them non-executable as written, and the single ~280-line file inlines material (hardware specs, extended workflow variants, time-sensitive version/star counts) that belongs in separate reference files. Defining the helper actions or trimming to truly runnable examples plus a reference bundle would lift both actionability and progressive disclosure.

Suggestions

Make workflow examples self-contained: define or import the helper actions ('toxicity_detector', 'extract_facts', 'verify_facts', 'fact_check_action') so the code is executable, or note explicitly that they are placeholders to be implemented.

Remove the empty '## Advanced topics' heading and the time-sensitive details ('⭐ 4,300+' star count, 'v0.12.0 expected') or move them to a versioned reference file.

Split the five detailed workflows, hardware requirements, and integration guides (Presidio, LlamaGuard) into references/ files with one-level-deep links from a concise SKILL.md overview.

DimensionReasoningScore

Conciseness

The body is mostly lean, code-first sections, but it includes unnecessary and time-sensitive padding — 'GitHub ... ⭐ 4,300+', 'Version: v0.9.0+ (v0.12.0 expected)', 'Production: NVIDIA enterprise deployments' — plus an empty '## Advanced topics' heading with no content beneath it. This matches the 3 anchor ('mostly efficient but includes some unnecessary explanation') rather than the 4 anchor's 'minor instances'.

3 / 5

Actionability

The Quick start and jailbreak workflow give concrete, runnable RailsConfig code, but workflows 2-4 hinge on undefined helpers ('toxicity_detector', 'extract_facts', 'verify_facts', 'fact_check_action') and the LlamaGuard import path does not match the real package API, so several blocks function as pseudocode. This lands on the 3 anchor ('pseudocode instead of executable code') rather than the 4 anchor, where gaps would be minor.

3 / 5

Workflow Clarity

The body is clearly sequenced — Quick start, five labeled workflows, a 'When to use vs alternatives' section, and a 'Common issues' section with concrete troubleshooting recipes (threshold adjustment, parallelized checks). Validation checkpoints are implicit rather than explicit, which holds it at the 4 anchor; the destructive/batch cap does not apply since this is configuration guidance, not a destructive operation.

4 / 5

Progressive Disclosure

There are no bundle files at all, and ~280 lines of five detailed workflows, hardware specs, and troubleshooting all live inline in SKILL.md with section headers. It is better organized than the 2 anchor's unstructured inlining, but extended workflow details and hardware requirements clearly belong in separate reference files — matching the 3 anchor ('some structure but could be better organized ... content that should be separate is inline') rather than the 4 anchor's 'most content is appropriately placed'.

3 / 5

Total

13

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, information-dense description with specific capabilities and clear distinctiveness, undermined by the complete absence of any 'Use when...' trigger guidance and a trailing marketing clause ('Production-ready, runs on T4 GPU') that adds fluff without aiding triggering. Adding an explicit trigger clause would move completeness from 3 to 4-5.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user mentions guardrails, jailbreak/prompt-injection attacks, PII leaking, or needs runtime safety checks on LLM input/output.'

Replace the marketing phrase 'Production-ready, runs on T4 GPU' with a concrete capability or fold hardware constraints into the trigger clause.

Include common synonyms users actually say — 'prompt injection', 'content moderation', 'LLM safety' — alongside the existing feature keywords.

DimensionReasoningScore

Specificity

The description lists six concrete capabilities ('jailbreak detection, input/output validation, fact-checking, hallucination detection, PII filtering, toxicity detection') plus the Colang 2.0 DSL, giving broad coverage. It stops short of the 5 anchor because 'Production-ready, runs on T4 GPU' is marketing fluff/over-claim rather than a stated capability, while it clearly exceeds the 3 anchor's '1-2 concrete actions'.

4 / 5

Completeness

The 'what' is explicit and clear (runtime safety framework with named mechanisms), but there is no 'Use when...' clause or equivalent trigger guidance anywhere in the description — the rubric guideline explicitly caps completeness at 3 for this. It is not a 4 because the 'when' is entirely absent rather than merely vague.

3 / 5

Trigger Term Quality

Terms like 'jailbreak detection', 'PII filtering', 'toxicity detection', and 'hallucination detection' are natural phrases users would say when needing this skill. A few common variations are missing (e.g. 'prompt injection', 'content moderation', 'LLM safety'), which keeps it at the 4 anchor rather than the 5 anchor's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

It names a specific vendor framework (NVIDIA's, with Colang 2.0 DSL) and niche safety mechanisms, giving it a clear niche with distinct triggers and minimal conflict risk with other skills — matching the 5 anchor.

5 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.