CtrlK
BlogDocsLog inGet started
Tessl Logo

outlines

Outlines: structured JSON/regex/Pydantic LLM generation.

47

Quality

52%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/inference/outlines/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is technically strong — accurate v1 API coverage, executable examples across all backends, and good version-migration handling — but it is significantly over-long and duplicative. Moving backend configuration and the example patterns into the existing reference files, and cutting basic-Pydantic best-practices explanations, would fix both conciseness and structure without losing actionability.

Suggestions

Move the "Backend Configuration" section and Patterns 2-6 into the existing references (backends.md, examples.md), keeping only Quick Start, the v1 API note, and one canonical example inline; reference them via "See [references/backends.md](references/backends.md)" style links at point of need.

Cut the "Best Practices" items that restate basic Pydantic knowledge (typed fields vs strings, enums for fixed sets, Optional fields) and the unverifiable "Performance Characteristics" claims; keep only the pattern of validating output with `model_validate_json`.

Factor the repeated `outlines.from_transformers(...)` setup into one snippet defined once, and make later pattern snippets self-contained (import outlines or state 'using the model defined above').

DimensionReasoningScore

Conciseness

At ~660 lines the body repeats the `outlines.from_transformers(...)` setup block six times, spends a "Best Practices" section on basic Pydantic knowledge Claude already has (typed fields vs strings, enums, Optional), and includes marketing-style padding ("Zero overhead", "1.2-2x faster", star counts). Several unnecessary/padded sections matches anchor 2; above anchor 1 because the code itself is current and substantive, not beginner filler.

2 / 5

Actionability

Quick Start and Core Concepts give fully executable, copy-paste-ready code for the current v1 API (real model names, imports, `model_validate_json` parsing). Minor gap: Patterns 2-5 and several backend snippets invoke `model(...)` without defining or importing it, so those blocks are not standalone. Anchor 4 (mostly executable with minor gaps) rather than 5.

4 / 5

Workflow Clarity

For this single-purpose skill the action sequence (install → wrap model via `from_*` factory → call with output_type → `model_validate_json`) is unambiguous, validation of results is shown consistently — including per-item validation in the batch Pattern 6 — and the version-sensitive pre-1.0 API is properly quarantined in an "API note (Outlines 1.x)" callout. Anchor 4 rather than 5 because there is no error-recovery guidance (e.g. what to do when validation fails or the backend lacks a feature).

4 / 5

Progressive Disclosure

The three reference files (json_generation.md, backends.md, examples.md) all exist, are one level deep, and are clearly signaled with descriptions in "See Also" — but the 17KB body inlines full backend-configuration sections and example patterns that duplicate references/backends.md and references/examples.md, making SKILL.md itself a near-monolith. Content that should be separate is inline matches anchor 3; signaling and structure keep it above anchor 2.

3 / 5

Total

13

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is terse and names the right domain and technologies, but reads like a keyword fragment rather than a trigger description. It lacks any 'use when' guidance and common natural phrasings users would actually say, capping both completeness and trigger-term quality.

Suggestions

Add an explicit 'Use when...' clause, e.g. "Use when generating guaranteed-valid JSON, structured outputs, or regex-constrained text from local or API LLMs, or when the user mentions Outlines, Pydantic schemas, or structured generation."

Include natural trigger synonyms users would say: "structured output", "JSON schema", "guaranteed valid JSON", "constrained decoding", ".json outputs".

Replace the fragment with a full third-person sentence listing 2-3 concrete actions (e.g. "Generates guaranteed-valid JSON, regex-matched text, and Pydantic-typed outputs by constraining LLM token sampling...").

DimensionReasoningScore

Specificity

"Outlines: structured JSON/regex/Pydantic LLM generation" names the domain and technologies, but the only action is the generic "generation" — no concrete actions like extract, validate, or constrain. Matches the anchor 'Names the domain but actions are minimal or generic'; below anchor 3 because no specific concrete action is listed.

2 / 5

Completeness

The description conveys a terse 'what' (structured JSON/regex/Pydantic LLM generation) but contains no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3. Clear 'what' with 'when' entirely missing matches anchor 3 rather than 4.

3 / 5

Trigger Term Quality

Includes relevant keywords ("JSON", "regex", "Pydantic", "LLM", "structured", the library name "Outlines") but misses natural variations users would say, such as "structured output", "JSON schema", "constrained generation", or "guaranteed valid JSON". Some relevant keywords with missing synonyms matches anchor 3.

3 / 5

Distinctiveness Conflict Risk

Naming the specific library "Outlines" plus distinctive technologies (Pydantic, regex-based generation) gives a mostly distinct profile with only minor overlap risk against sibling structured-output libraries (Instructor, Guidance). Mostly distinct matches anchor 4; not 5 because the niche is not sealed with explicit trigger phrases.

4 / 5

Total

12

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (663 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.