CtrlK
BlogDocsLog inGet started
Tessl Logo

instructor

Structured LLM outputs validated with Pydantic.

48

Quality

55%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/instructor/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, highly executable reference for Instructor, with excellent code coverage of the common patterns. Its weaknesses are verbosity and structure: marketing/benefits padding and near-verbatim duplication of the bundle's own reference files bloat SKILL.md instead of delegating to progressive disclosure.

Suggestions

Delete the Provider Configuration and Common Patterns sections from SKILL.md and replace them with inline pointers, e.g. '**Provider setup**: See references/providers.md' — they are already duplicated in the bundle.

Trim Claude-knowledge padding: the 'Benefits:' lists after code samples, the GitHub-stars/battle-tested marketing line, and the Discord entry in Resources add no actionable information.

Move the 'See Also' reference pointers next to their related sections (validation patterns → references/validation.md, providers → references/providers.md) so navigation is signaled where a reader needs it.

DimensionReasoningScore

Conciseness

At ~740 lines the body pads noticeably: 'Benefits:' bullet lists of things Claude already knows (type hints, IDE autocomplete), marketing stats ('GitHub Stars: 15,000+ | Battle-tested: 100,000+ developers'), a resources section, and whole sections (Provider Configuration, Common Patterns) duplicated nearly verbatim in references/providers.md and references/examples.md. This matches 'Noticeably verbose; several unnecessary explanations or padded sections' — below the score-3 anchor, where padding would be occasional rather than structural.

2 / 5

Actionability

The body is dominated by executable, copy-paste-ready Python covering the common cases (extraction, classification, streaming, validation, error handling), with the Quick Start fully runnable as written. It falls short of score 5 only because many later snippets use placeholders (messages=[...], YourModel, undefined User/Sentiment/HttpUrl in some blocks), which is exactly 'Mostly executable guidance; concrete code or commands with minor gaps'.

4 / 5

Workflow Clarity

The usage flow is unambiguous and linearly presented (install → define model → build client → call with response_model), and the retry feedback loop is explicitly documented ('If invalid: Error message sent back to LLM... Repeats up to max_retries') with a try/except error-handling section. It sits at 'Clear sequence with most checkpoints present' rather than 5 because the sequence is implied by section order rather than stated as an explicit checklist, and validation guidance is illustrative rather than prescribed steps.

4 / 5

Progressive Disclosure

The three references (validation.md, providers.md, examples.md — all verified real and one level deep) are listed under 'See Also' with labels, but that listing sits at the very end instead of being signaled at the relevant sections, and the body inlines large blocks (provider setup, the five patterns) that duplicate the reference files almost verbatim. That matches 'references present but not clearly signaled; content that should be separate is inline' — better than score 2 (no headers, references buried), short of score 4 where placement would be mostly appropriate.

3 / 5

Total

13

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear, concrete capability but omits any usage triggers and under-represents the library's breadth (retries, streaming, multi-provider extraction). It is distinct and non-generic, just incomplete as a trigger surface.

Suggestions

Append an explicit trigger clause, e.g. 'Use when extracting structured data or JSON from LLM responses, when the user mentions Instructor, Pydantic response models, or schema-validated outputs.'

Broaden natural trigger terms to include synonyms users actually say: 'JSON extraction', 'data extraction', 'schema validation', 'response_model', 'structured output'.

Mention one or two more concrete capabilities (automatic retries on validation failure, multi-provider support) to lift specificity beyond the 1-2-action level.

DimensionReasoningScore

Specificity

The description 'Structured LLM outputs validated with Pydantic' names the domain (structured LLM outputs) and one concrete capability (validation via Pydantic), but stops there — no mention of extraction, retries, streaming, or multi-provider support. This matches the anchor 'Names domain and 1-2 concrete actions, but not comprehensive', not score 4 which requires several specific actions, nor score 2 whose actions are purely generic.

3 / 5

Completeness

The 'what' is clear (structured outputs validated with Pydantic), but there is no 'Use when...' clause or any equivalent trigger guidance. Per the judging guidelines a missing 'Use when...' clause caps completeness at 3, matching the anchor 'Has a clear what but when is missing'.

3 / 5

Trigger Term Quality

'Structured LLM outputs', 'Pydantic', and 'validated' are phrases users would naturally use, but common variations are missing: 'JSON', 'data extraction', 'schema', 'retry', and the library name 'Instructor'. Anchor 3 ('Some relevant keywords but missing common variations or synonyms') is the best fit — score 4 requires good coverage with only a few natural terms missing.

3 / 5

Distinctiveness Conflict Risk

'Structured LLM outputs validated with Pydantic' carves a fairly clear niche — few skills would compete for it, with only minor overlap risk against general JSON/validation skills. It fits 'Mostly distinct; minor overlap risk' better than score 3 ('could still overlap with similar skills') because Pydantic + structured outputs is a specific pairing, and better than score 5 which would require explicit distinct trigger phrases.

4 / 5

Total

13

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (745 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.