CtrlK
BlogDocsLog inGet started
Tessl Logo

dspy

Build complex AI systems with declarative programming, optimize prompts automatically, create modular RAG systems and agents with DSPy - Stanford NLP's framework for systematic LM programming

53

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/dspy/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, code-rich reference for DSPy with strong actionability, held back by duplicated inline content that belongs in the existing reference files and the absence of validation checkpoints in its workflows. Tightening the main file to a quick start plus pointers would improve both conciseness and progressive disclosure.

Suggestions

Replace the inline Modules and Optimizers sections with 1-2 key examples each plus pointers to references/modules.md and references/optimizers.md, keeping the full catalogs in the reference files only.

Add explicit validation checkpoints, e.g. 'Evaluate a baseline with dspy.evaluate.Evaluate before compiling an optimizer, then re-evaluate the optimized module to confirm improvement.'

Fix the syntax error in the Best Practices trainset example (unbalanced quotes at line 527) and remove time-sensitive details like the GitHub star count.

Fill in or clearly label the placeholder tool implementation in the ReAct example so the code is runnable as written.

DimensionReasoningScore

Conciseness

Mostly efficient code-first content, but the Modules and Optimizers sections (~100 lines) duplicate material that lives in references/modules.md and references/optimizers.md, and padding like "# Now optimized_qa performs better!" and the time-sensitive "GitHub Stars: 22,000+" could be cut. Not 2 because there is little conceptual explanation of things Claude already knows.

3 / 5

Actionability

Abundant concrete, mostly copy-paste-ready Python covering signatures, modules, optimizers, providers, and patterns. Not 5 because of placeholder code ("# Your search implementation", "return results" in the ReAct tool) and a syntax error in the Best Practices trainset example (answer="...) at line 527).

4 / 5

Workflow Clarity

There is a sensible progression (install -> quick start -> concepts -> patterns -> evaluation) and the Evaluation section shows before/after comparison, but no explicit validation checkpoints or feedback loops (e.g., evaluate a baseline before optimizing, retry on failed compiles). Not 2 because the ordering and "Start Simple, Iterate" best practice give a usable rough sequence.

3 / 5

Progressive Disclosure

Good section structure with real one-level-deep references clearly listed in the "See Also" section, but the inline Modules and Optimizers sections duplicate content that belongs in those reference files, bloating the main file. Not 4 because content that should be separate is inline rather than only "minor organization gaps".

3 / 5

Total

13

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description with concrete named capabilities and a clearly identified framework, weakened by the complete absence of a "Use when..." trigger clause. Adding explicit trigger conditions would raise completeness and further sharpen distinctiveness.

Suggestions

Append a 'Use when...' clause, e.g. 'Use when the user mentions DSPy, wants to replace manual prompt engineering with data-driven optimization, or is building RAG, agents, or multi-stage LM pipelines in Python.'

Include natural trigger variations such as 'prompt engineering', 'automatic prompt optimization', and 'prompt optimizer' to improve keyword coverage.

Trim the abstract framing ('Build complex AI systems with declarative programming') in favor of the concrete actions it already lists.

DimensionReasoningScore

Specificity

Lists several concrete actions ("optimize prompts automatically", "create modular RAG systems and agents") alongside the more abstract "Build complex AI systems with declarative programming". Not 5 because the lead action and "systematic LM programming" are broad framing rather than concrete capabilities.

4 / 5

Completeness

The "what" is clear (build, optimize, create RAG/agents with DSPy), but there is no "Use when..." clause or equivalent trigger guidance anywhere in the description, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Natural user phrases like "DSPy", "optimize prompts", "RAG", "agents", and "Stanford NLP" are present. Not 5 because common variations such as "prompt engineering", "automatic prompt optimization", or framework-adjacent synonyms are missing.

4 / 5

Distinctiveness Conflict Risk

Naming DSPy and Stanford NLP gives it a clear niche, but trigger phrases like "RAG systems", "agents", and "optimize prompts" overlap with general LLM-framework skills (LangChain, LlamaIndex). Not 3 because the explicit framework name keeps it mostly distinct; not 5 because the generic capability terms create minor overlap risk.

4 / 5

Total

15

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (600 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.