CtrlK
BlogDocsLog inGet started
Tessl Logo

guidance

Control LLM output with regex and grammars, guarantee valid JSON/XML/code generation, enforce structured formats, and build multi-step workflows with Guidance - Microsoft Research's constrained generation framework

55

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/guidance/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers mostly executable, well-sequenced Guidance examples but suffers from significant duplication, marketing padding, and a monolithic structure that inlines material which duplicates the provided reference files. It would score much higher after de-duplicating examples, trimming explanatory fluff, and moving backend/pattern detail into the references with inline pointers.

Suggestions

Remove duplicated examples (generate_person and react_agent each appear twice) and drop the 'Benefits:' bullet lists, 'GitHub Stars' line, and 'Performance Characteristics' claims section that add tokens without new instruction.

Move the Backend Configuration detail into references/backends.md and the extended patterns into references/examples.md, replacing them with inline pointers at the relevant sections instead of a terminal 'See Also' list.

Fix the grammar example to use valid Guidance API (e.g., a real guidance grammar or select/list constructs) so the copy-paste code is actually executable, and add a lightweight validation step (e.g., json.loads on generated output) to the JSON pattern.

DimensionReasoningScore

Conciseness

The ~560-line body is noticeably verbose: the generate_person and react_agent examples each appear twice, every concept is padded with a 'Benefits:' bullet list, and it includes marketing fluff ('GitHub Stars: 18,000+') and a claims-only 'Performance Characteristics' section explaining things Claude already knows. Anchor 2 (several unnecessary/padded sections) fits better than 3 (only occasional tightening needed) given the duplication and padding throughout.

2 / 5

Actionability

Guidance is mostly executable — real install commands, context managers, select(), and @guidance functions — with concrete patterns covering common cases. It is not 5 because the grammar section passes an invented '<gen name regex=...>' template string via grammar=, which is not valid Guidance API and would fail if copied.

4 / 5

Workflow Clarity

There is a clear, coherent progression (Installation → Quick Start → Core Concepts → Backends → Patterns → Best Practices) and the skill is not a destructive/batch operation so the validation cap does not apply. It falls short of 5 because there are no verification checkpoints (e.g., asserting a generated JSON parses) anywhere in the workflows.

4 / 5

Progressive Disclosure

Three substantial reference files exist (references/backends.md, constraints.md, examples.md), but they are only linked in a terminal 'See Also' list rather than signaled inline where the detail lives, while ~200 lines of backend configuration and pattern detail that belongs in those references is inlined in SKILL.md (e.g., the Backend Configuration section duplicates references/backends.md). References present but not clearly signaled with separable content inline matches anchor 3.

3 / 5

Total

13

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong 'what' description with concrete, comprehensive actions and mostly distinct identity, but it entirely lacks a 'Use when...' trigger clause, which both caps completeness and weakens its usefulness for skill routing. Adding explicit usage triggers and a few natural synonyms would raise it substantially.

Suggestions

Add an explicit trigger clause, e.g., 'Use when generating structured LLM outputs, validating JSON/XML, or when the user mentions Guidance, constrained generation, regex-constrained output, or structured output.'

Include natural synonym phrasings users would actually say — 'structured output', 'valid JSON', 'schema-constrained generation' — to broaden trigger coverage.

Sharpen distinctiveness by contrasting with adjacent tools, e.g., 'with Guidance (rather than JSON-schema validation or Pydantic retries)'.

DimensionReasoningScore

Specificity

The description lists four concrete, specific actions ("Control LLM output with regex and grammars", "guarantee valid JSON/XML/code generation", "enforce structured formats", "build multi-step workflows"), giving comprehensive coverage of the tool's capabilities matching the top anchor. It is not 4 because there are no meaningful gaps in the action coverage.

5 / 5

Completeness

The 'what' is clearly and concretely answered, but there is no 'Use when...' clause or any equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. It is not 2 because the 'what' half is strong, not vague.

3 / 5

Trigger Term Quality

Natural keywords are present ("regex", "grammars", "JSON/XML/code", "structured formats", "multi-step workflows", "Guidance", "constrained generation") but common variations users would say are missing (e.g., "structured output", "valid JSON", "schema"). Good coverage with a few natural terms missing fits anchor 4 rather than the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

Naming the distinct Guidance framework and its constrained-generation niche makes it mostly distinct, but generic phrasing like "enforce structured formats" and "valid JSON generation" overlaps with other structured-output skills (e.g., JSON-schema or validation skills). This is minor overlap risk with closely related skills (anchor 4) rather than a fully clear niche (anchor 5).

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (582 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.