CtrlK
BlogDocsLog inGet started
Tessl Logo

guidance

Constrain LLM output with grammars; guarantee valid JSON.

53

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/guidance/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a strong, highly actionable skill body — code throughout is executable and the learning progression (install → concepts → patterns) is sound. The main weaknesses are duplication that inflates token cost and a body/references split that overlaps and even contradicts itself on the Anthropic backend.

Suggestions

Remove the duplicated ReAct agent and generate_person examples — keep one authoritative copy and point to references/examples.md for the rest.

Reconcile the models.Anthropic contradiction: references/backends.md and references/constraints.md use models.Anthropic("claude-sonnet-4-5-20250929") while the body states guidance 0.3.x has no models.Anthropic class.

Trim marketing/padding sections ("GitHub Stars: 18,000+", the Performance Characteristics claims) and move backend configuration detail wholly into references/backends.md.

DimensionReasoningScore

Conciseness

The ~560-line body has real duplication — the ReAct agent appears in full twice ("Stateful Functions" and Pattern 4), generate_person twice, and the chat context-manager example twice — plus padded marketing ("GitHub Stars: 18,000+", "30-50% faster" claims) and benefits lists restating obvious points. Mostly useful content that could be noticeably tightened: anchor 3 ('Mostly efficient but includes some unnecessary explanation or could be tightened').

3 / 5

Actionability

Every section ships executable, copy-paste-ready code: pip install commands, quick start, regex/select/grammar examples, backend configs, and concrete good/bad best-practice pairs. This matches anchor 5 ('Fully executable; copy-paste ready code or commands; specific examples cover the common cases').

5 / 5

Workflow Clarity

The skill follows a clear sequence (installation → quick start → core concepts → backends → patterns → best practices), and the multi-step guidance functions (e.g. the ReAct loop with an explicit done-check and break) are well sequenced. Fits anchor 4 ('Clear sequence with most checkpoints present'); not 5 because there are no error-recovery or output-verification checkpoints (e.g. validating generated JSON parses, or handling backend constraint failures).

4 / 5

Progressive Disclosure

All three referenced files (references/backends.md, constraints.md, examples.md) exist, are one level deep, and are clearly signaled with descriptions in the "See Also" section. Minor gaps: the body duplicates much of the references' content (backend config ↔ backends.md; patterns ↔ examples.md), and the references contradict the body on models.Anthropic (references use it; the body states guidance 0.3.x has no such class). Fits anchor 4 ('Good structure; most content is appropriately placed; references mostly clear; minor organization gaps').

4 / 5

Total

16

/

20

Passed

Description

50%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and technically accurate but under-serves discovery: it delivers a clear two-action "what" with no "when" clause and thin trigger-term coverage. Adding an explicit "Use when..." sentence with natural trigger phrases would raise completeness, trigger quality, and distinctiveness together.

Suggestions

Add an explicit trigger clause, e.g. "Use when you need guaranteed-valid structured output from an LLM, constrained regex/grammar generation, or JSON that must parse."

Include natural user phrasings as trigger terms: "structured output", "constrained generation", "regex-constrained", "JSON schema", "force valid JSON".

Name one or two more concrete capabilities (e.g. select()-based choices, Pydantic-schema JSON) so the "what" is comprehensive rather than minimal.

DimensionReasoningScore

Specificity

"Constrain LLM output with grammars; guarantee valid JSON" names the domain and two concrete actions, but coverage stops there — regex constraints, select(), structured extraction, and multi-step workflows are absent. This matches anchor 3 ('Names domain and 1-2 concrete actions, but not comprehensive') and falls short of anchor 4, which expects several specific actions.

3 / 5

Completeness

There is a clear "what" (constrain output with grammars, guarantee valid JSON) but no "when" — the description has no "Use when..." clause or equivalent trigger guidance, which per the rubric caps completeness at 3. This exactly matches anchor 3 ('Has a clear what but when is missing or only weakly implied').

3 / 5

Trigger Term Quality

The description contains some relevant keywords ("LLM output", "grammars", "JSON") but misses the natural phrases users would say — "structured output", "regex", "constrained generation", "schema", "valid JSON". Anchor 3 ('Some relevant keywords but missing common variations or synonyms') is the closest fit; anchor 4 requires broader keyword coverage that is not present.

3 / 5

Distinctiveness Conflict Risk

"Constrain LLM output with grammars" is a reasonably distinct niche, but "guarantee valid JSON" overlaps with JSON-validation and structured-output skills, and with no "when" guidance the description could trigger for the wrong skill. It sits between anchors 3 and 4; absent any trigger differentiation it lands at anchor 3 ('Somewhat specific but could still overlap with similar skills').

3 / 5

Total

12

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (581 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.