CtrlK
BlogDocsLog inGet started
Tessl Logo

llamaguard

Meta's 7-8B specialized moderation model for LLM input/output filtering. 6 safety categories - violence/hate, sexual content, weapons, substances, self-harm, criminal planning. 94-95% accuracy. Deploy with vLLM, HuggingFace, Sagemaker. Integrates with NeMo Guardrails.

58

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/llamaguard/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, code-dense skill body that is highly actionable and tightly organized around five practical workflows with useful troubleshooting. The two clear weaknesses are the absence of any progressive disclosure — all deployment and reference material is inlined in one long file — and small executable gaps such as an undefined tokenizer in the vLLM/FastAPI examples.

Suggestions

Move deployment details (vLLM tuning, FastAPI endpoint, NeMo setup, hardware requirements) into references/ files with one-level-deep pointers from SKILL.md, keeping the quick start and core filtering workflows inline.

Define or instantiate `tokenizer` in the vLLM and FastAPI workflows so the examples run as-is, and complete the false-positive threshold snippet (currently references undefined `unsafe_token_id` and a `model(...)` placeholder).

Remove the empty 'Advanced topics' header and deduplicate the throughput/VRAM figures that appear in both the workflow and hardware-requirements sections.

DimensionReasoningScore

Conciseness

The body is lean and code-first with essentially no padding of concepts Claude already knows; each section delivers working code or concrete figures. It misses 5 due to redundancy — throughput and VRAM figures appear twice, the 'Advanced topics' header is empty, and dated version info ('LlamaGuard 3 (2024)') is not isolated in a deprecation/old-patterns section.

4 / 5

Actionability

Mostly executable, copy-paste-ready code covering installation, input/output filtering, vLLM serving, a FastAPI endpoint, and NeMo integration. Minor gaps keep it below 5: `tokenizer` is used but never defined in the vLLM/FastAPI workflows, and the false-positive threshold snippet references an undefined `unsafe_token_id` with a `model(...)` placeholder rather than runnable code.

4 / 5

Workflow Clarity

The five workflows (input filtering, output filtering, vLLM deployment, API endpoint, NeMo integration) are clearly sequenced and well-separated, with a troubleshooting section for access, latency, false-positive, and OOM issues. It is below 5 because no workflow includes an explicit validation checkpoint or feedback loop (e.g. verifying the model output format before parsing category codes), though these are read-only classification calls rather than destructive or batch operations.

4 / 5

Progressive Disclosure

Section headers give the ~320-line body reasonable internal structure, but everything — deployment details, API serving, NeMo configuration, hardware requirements, model-version history — is inlined in SKILL.md with no references/ bundle and no one-level-deep pointers to separate files. This matches the 'some structure, content that should be separate is inline' anchor; it is above 2 because navigation within the file itself is clear.

3 / 5

Total

15

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense, specific description that clearly communicates what the skill covers (model, categories, deployment options, integrations) in third person without fluff. Its main weakness is the complete absence of a 'Use when...' trigger clause, which both caps completeness and leaves the description weaker on natural trigger phrasing.

Suggestions

Append an explicit trigger clause, e.g. 'Use when moderating LLM prompts or responses, filtering unsafe content, or when the user mentions LlamaGuard, content moderation, or safety classification.'

Add natural synonyms users would say — 'content moderation', 'toxicity detection', 'prompt filtering' — to broaden trigger-term coverage.

State the actions the skill itself performs (classify, filter, serve) rather than only the model's properties (accuracy, parameter count).

DimensionReasoningScore

Specificity

The description names the domain and several concrete capabilities — 'input/output filtering', the six enumerated safety categories, 'Deploy with vLLM, HuggingFace, Sagemaker', and 'Integrates with NeMo Guardrails' — matching the anchor for several specific actions with minor gaps. It falls short of 5 because it describes the model's properties more than a comprehensive list of actions the skill performs.

4 / 5

Completeness

The 'what' is clearly answered (specialized moderation model, categories, accuracy, deployment options) but there is no 'Use when...' clause or equivalent explicit trigger guidance, which the judging guidelines cap at 3. It is above 2 because the 'what' is concrete and multi-faceted, not vague.

3 / 5

Trigger Term Quality

Good keyword coverage with natural terms users would say: 'moderation', 'safety', 'input/output filtering', 'guardrails', 'weapons', 'self-harm'. It stays at 4 rather than 5 because common variations like 'content moderation', 'toxicity detection', or 'prompt filtering' are absent.

4 / 5

Distinctiveness Conflict Risk

It names a specific niche (Meta's LlamaGuard moderation model) with model-specific terms like 'LlamaGuard', 'NeMo Guardrails', and the six safety categories, making it mostly distinct. Minor overlap risk remains with adjacent moderation/guardrail skills (OpenAI Moderation API, NeMo Guardrails framework itself), keeping it below 5.

4 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.