CtrlK
BlogDocsLog inGet started
Tessl Logo

ai-engineering-toolkit

6 production-ready AI engineering workflows: prompt evaluation (8-dimension scoring), context budget planning, RAG pipeline design, agent security audit (65-point checklist), eval harness building, and product sense coaching.

28

Quality

21%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

Fix and improve this skill with Tessl

tessl review fix ./skills/ai-engineering-toolkit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

0%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This skill reads as a README/marketing document rather than an actionable skill file. It describes 6 ambitious workflows at a high level but provides none of the actual content Claude would need to execute any of them — no scoring rubrics, no checklists, no decision trees, no step-by-step procedures, no executable code. The content is simultaneously too verbose (explaining concepts Claude knows, repeating information across sections) and too shallow (never reaching actionable depth on any workflow).

Suggestions

Replace high-level descriptions with actual executable content: provide the 8-dimension scoring rubric with weights, the 65-point security checklist, the RAG decision tree, etc. — either inline or in referenced bundle files.

Remove marketing-style content (overview explaining what skills are, installation instructions, repository links, 'When to Use' bullets that restate descriptions) to free token budget for actual methodology.

Create separate bundle files for each of the 6 workflows containing the detailed procedures, and reference them from SKILL.md with clear one-level-deep navigation.

Add concrete step-by-step procedures with validation checkpoints — e.g., for the prompt evaluator: Step 1: Score each dimension (provide rubric), Step 2: Calculate weighted aggregate (provide formula), Step 3: Identify bottom 3, Step 4: Generate rewrite targeting those dimensions, Step 5: Re-score and compare.

DimensionReasoningScore

Conciseness

Extremely verbose for a skill file. The overview section explains what skills are and why consistency matters (Claude already knows this). Extensive descriptions of each sub-skill read like marketing copy rather than actionable instructions. The 'When to Use' section lists 6 bullet points that largely restate the skill descriptions. Installation instructions and repository links consume tokens without teaching Claude how to perform the tasks.

1 / 3

Actionability

Despite describing 6 workflows, the skill provides zero executable code, no concrete commands, no scoring rubric details, no actual checklists, and no decision trees. It describes what each skill does at a high level ('Scores prompts across 8 dimensions') but never provides the actual dimensions, weights, scoring criteria, or step-by-step procedures Claude would need to execute any of these workflows. The examples show expected outputs but not the process to produce them.

1 / 3

Workflow Clarity

No multi-step workflows are actually defined despite claiming to encode 'step-by-step decision frameworks.' The skill describes what workflows exist but never sequences the steps, provides validation checkpoints, or defines feedback loops. For example, the 65-point security audit mentions attack categories but provides zero actual test procedures or pass/fail criteria.

1 / 3

Progressive Disclosure

No bundle files are provided, yet the skill describes 6 complex workflows that clearly need detailed sub-documents (scoring rubrics, checklists, decision trees, templates). Everything is in a single monolithic file that paradoxically contains no actionable detail — it's all summary with no depth anywhere. There are no references to supporting files that would contain the actual methodology.

1 / 3

Total

4

/

12

Passed

Description

42%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description excels at listing specific, concrete capabilities with impressive detail (8-dimension scoring, 65-point checklist), but critically lacks any 'Use when...' guidance to help Claude know when to select this skill. The trigger terms are domain-appropriate but lean toward feature listing rather than matching natural user language patterns.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when the user asks about prompt engineering, RAG systems, LLM context windows, AI agent security, evaluation frameworks, or AI product development.'

Include natural user-facing trigger terms and variations such as 'LLM', 'large language model', 'prompt engineering', 'retrieval augmented generation', 'context window', 'AI agent', 'model evaluation'

Consider whether the six workflows should be separate skills to reduce overlap risk, or at minimum clarify the unifying theme (e.g., 'AI engineering best practices') to improve distinctiveness.

DimensionReasoningScore

Specificity

Lists six specific concrete workflows: prompt evaluation with 8-dimension scoring, context budget planning, RAG pipeline design, agent security audit with 65-point checklist, eval harness building, and product sense coaching. These are highly specific and actionable.

3 / 3

Completeness

The description answers 'what does this do' well but completely lacks any 'when should Claude use it' guidance. There is no 'Use when...' clause or equivalent explicit trigger guidance, which per the rubric should cap completeness at 2, and since the 'when' is entirely absent, a score of 1 is appropriate.

1 / 3

Trigger Term Quality

Contains some good domain-specific terms like 'RAG pipeline', 'prompt evaluation', 'agent security audit', 'eval harness', but lacks natural user-facing trigger terms. Users might say things like 'review my prompt', 'optimize context window', 'build RAG system', or 'AI safety' which aren't explicitly covered. The terms lean more toward listing features than matching user language.

2 / 3

Distinctiveness Conflict Risk

The six workflows are fairly specific to AI engineering, which helps distinguish it, but the breadth of coverage (prompt evaluation, RAG, security, coaching) means it could overlap with more focused skills in any of those individual areas. The combination is somewhat distinctive but the wide scope increases conflict risk.

2 / 3

Total

8

/

12

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
popey/claude-code-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.