CtrlK
BlogDocsLog inGet started
Tessl Logo

instructor

Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results with Instructor - battle-tested structured output library

57

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/instructor/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, largely executable reference for Instructor with real model IDs and good coverage of core patterns, sequencing, and error feedback. Its weaknesses are noticeable padding and boilerplate repetition, and weak progressive disclosure: detailed material that already lives in the references/ files is duplicated inline with only a buried See Also list for navigation.

Suggestions

Replace the repeated client.messages.create(...) boilerplate in every pattern/advanced section with a single stated convention (e.g., 'All examples assume client and User from Quick Start'), and cut the marketing line, Benefits lists, and 'Custom Error Messages' section to tighten conciseness.

Move the Provider Configuration, Common Patterns, and detailed Validation sections into references/providers.md, references/examples.md, and references/validation.md respectively, keeping only one illustrative snippet inline with a clearly signaled 'See [references/validation.md](references/validation.md)' pointer at each relevant section.

Add per-item try/except error handling to the batch processing pattern to complete the feedback loop expected of batch operations.

DimensionReasoningScore

Conciseness

Most content is actionable instructor-specific code Claude would not know from memory, but it is padded with marketing claims ('Battle-tested: 100,000+ developers'), 'Benefits:' lists, a restated 'How it works' sequence, a misleading 'Custom Error Messages' section that actually shows json_schema_extra examples, and a ~20x-repeated client.messages.create(...) boilerplate block — 'mostly efficient but could be tightened' rather than the 'noticeably verbose' 2 anchor since little of it explains concepts Claude already knows.

3 / 5

Actionability

Examples use real imports, real model IDs (claude-sonnet-4-5-20250929, gpt-4o-mini), and runnable Pydantic definitions, but many later snippets depend on an undefined 'client', 'YourModel', or 'messages=[...]' placeholders, so guidance is 'mostly executable with minor gaps' rather than fully copy-paste ready.

4 / 5

Workflow Clarity

The document follows a coherent sequence (Installation → Quick Start → response models → validation/retry → error handling) and includes a real feedback loop (automatic retry with error feedback plus the ValidationError try/except pattern), but patterns like batch processing lack per-item error checkpoints, keeping it below the explicit-checkpoint 5 anchor.

4 / 5

Progressive Disclosure

All three references (validation.md, providers.md, examples.md) exist and are one level deep, but they are signaled only in a bottom 'See Also' list rather than at the relevant sections, and the body inlines substantial duplicate content (Provider Configuration vs providers.md, Common Patterns vs examples.md, the validation sections vs the 606-line validation.md), matching 'references present but not clearly signaled; content that should be separate is inline'.

3 / 5

Total

14

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific, third-person description that clearly states what the skill does across five concrete capabilities. Its main weakness is the complete absence of any 'Use when…' trigger guidance, which both caps completeness and slightly weakens trigger-term quality and distinctiveness.

Suggestions

Append an explicit trigger clause such as 'Use when extracting structured data or JSON from LLM responses, validating LLM outputs against a schema, or the user mentions Instructor, Pydantic response models, or structured outputs' to raise completeness from 3.

Add natural synonym phrases users would say (e.g., 'structured output', 'schema validation', 'type-safe LLM responses', 'function calling') to broaden trigger-term coverage.

Drop the marketing tail 'battle-tested structured output library', which adds no trigger value and consumes description budget.

DimensionReasoningScore

Specificity

The description lists five concrete, third-person actions — 'Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results' — which matches the comprehensive-coverage anchor rather than the 4 anchor that expects minor gaps.

5 / 5

Completeness

The 'what' is explicit and concrete, but there is no 'Use when…' clause or equivalent explicit trigger guidance anywhere in the description, which per the judging guidelines caps completeness at 3; the 'when' is absent rather than merely imprecise, ruling out 4.

3 / 5

Trigger Term Quality

Natural trigger terms are present ('structured data', 'LLM responses', 'Pydantic', 'JSON', 'validation', 'stream', 'structured output'), but common user variations such as 'schema', 'function calling', or 'parse LLM output' are missing, fitting the 'good coverage, a few natural terms missing' anchor rather than the synonym-complete 5 anchor.

4 / 5

Distinctiveness Conflict Risk

Naming 'Instructor', 'Pydantic', and 'structured output' carves a distinct niche, but generic phrases like 'parse complex JSON' and 'LLM responses' could overlap with sibling llm-tools skills, matching 'mostly distinct; minor overlap risk' rather than the minimal-conflict 5 anchor.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (750 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.