CtrlK
BlogDocsLog inGet started
Tessl Logo

dspy

DSPy: declarative LM programs, auto-optimize prompts, RAG.

51

Quality

58%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/research/dspy/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, code-heavy reference for DSPy with strong, mostly executable examples and a sensible learning progression. Its main weaknesses are token inefficiency from duplicating the reference bundle files inline and burying the pointers to those files in a trailing See Also section, plus a few non-executable snippets.

Suggestions

Move the detailed module and optimizer sections (### 2. Modules, ### 3. Optimizers) into references/modules.md and references/optimizers.md, keeping only one example per concept in SKILL.md and linking at the point of use (e.g., 'See references/modules.md for the full module guide') instead of a trailing See Also.

Fix the non-executable snippets: the broken quote in 'Optimize with Representative Data' (`answer="...).with_inputs(...)`), the undefined `validate_answer`/`trainset` in the RAG example, and `search_tool` returning the undefined `results`.

Trim duplication — the MultiHopQA and RAG System examples, repeated provider configuration blocks, and the LangChain comparison table overlap content already in references/examples.md and could be consolidated to cut token cost.

DimensionReasoningScore

Conciseness

The body is dense code rather than padded prose, but at ~580 lines it substantially duplicates content that already exists in the references/ files ("### 2. Modules" and "### 3. Optimizers" re-cover modules.md and optimizers.md), repeats `dspy.settings.configure(lm=lm)` in nearly every example, and presents the same RAG pipeline twice ("MultiHopQA" and "RAG System with Optimization"). Mostly useful material, but it could be tightened considerably by moving detail to the references and trimming redundant examples.

3 / 5

Actionability

The bulk of the guidance is concrete, copy-paste-ready Python covering the common cases (signatures, Predict/ChainOfThought, BootstrapFewShot, multi-stage modules, provider config). Minor gaps keep it below 5: "Best Practices #3" has a syntax error (`answer="...).with_inputs("question")`), the RAG example uses `validate_answer` and `trainset` that are never defined in that section, and `search_tool` returns an undefined `results` variable.

4 / 5

Workflow Clarity

The document follows a clear progression (Installation → Quick Start → Core Concepts → Common Patterns → Evaluation → Best Practices) and "Start Simple, Iterate" ("Start with Predict", "Add reasoning if need", "Add optimization when you have data") gives an explicit improvement sequence. It loses a point because the optimize→evaluate→compare loop ("Compare optimized vs unoptimized") is shown as fragments rather than one connected checkpoint workflow.

4 / 5

Progressive Disclosure

References exist (modules.md, optimizers.md, examples.md — all verified present) and there is a "See Also" section, but they are only surfaced at the very end rather than signaled at the point of use, and the inline "Modules"/"Optimizers"/RAG sections (~200+ lines) duplicate the reference files instead of deferring to them. This matches "references present but not clearly signaled; content that should be separate is inline".

3 / 5

Total

14

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is short, domain-specific, and anchored by the distinctive 'DSPy' name, but it is a telegraphic fragment: it names a couple of capabilities without explaining when to use the skill, and it omits the natural trigger phrases (prompt optimization, RAG, few-shot, signatures) that would help discovery. Expanding it into full third-person sentences with a 'Use when...' clause would move it up two levels.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when building RAG pipelines, agents, or classifiers in DSPy, or when the user wants to auto-optimize prompts instead of hand-tuning them.'

Expand the compressed fragments into third-person verb phrases and cover more concrete capabilities: 'Builds declarative LM pipelines with signatures and modules; optimizes prompts automatically with data-driven optimizers (BootstrapFewShot, MIPRO); supports RAG, agents, and classifiers.'

Include natural trigger terms and synonyms users would actually say — 'prompt optimization', 'prompt engineering', 'few-shot', 'teleprompter', 'Stanford DSPy' — to improve keyword coverage.

DimensionReasoningScore

Specificity

The description names the domain ("DSPy") and a couple of concrete actions ("declarative LM programs, auto-optimize prompts") plus "RAG", but the telegraphic abbreviation style keeps it from listing several well-formed actions. It sits between anchor 3 (domain plus 1-2 concrete actions, not comprehensive) and anchor 4 (several specific actions with minor gaps); the compressed, noun-heavy phrasing pulls it to 3.

3 / 5

Completeness

It conveys 'what' (declarative LM programming, automatic prompt optimization, RAG) but has no 'Use when...' clause or equivalent trigger guidance, which per the judging guidelines caps completeness at 3. The 'what' is clear, the 'when' is entirely absent — exactly anchor 3.

3 / 5

Trigger Term Quality

"DSPy" and "prompts" are natural terms a user would say, but the description misses common variations and synonyms users actually use: "prompt optimization", "signatures", "teleprompters", "few-shot", "prompt engineering". It has some relevant keywords but is missing common variations, matching anchor 3 rather than anchor 4's 'good keyword coverage'.

3 / 5

Distinctiveness Conflict Risk

Leading with the unique framework name "DSPy" gives it a clear niche and low conflict risk with other skills. It stops short of anchor 5 because the remaining text ("auto-optimize prompts", "RAG") consists of generic terms that could plausibly match a general prompt-engineering or RAG skill if the DSPy token were dropped or the name unrecognized.

4 / 5

Total

13

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (595 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.