CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-ai-assistant-claims

Evaluate a vendor's productivity claim about an AI coding assistant against the evidence checklist from the talk. Use when someone quotes a percentage gain and asks whether the team should adopt the tool.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is evaluate-ai-assistant-claims in jbaruch/shownotes-template

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary short, instruction-only skill: lean prose, concrete questions and recommendations, a defined output format with a worked example, and clean section organization with no unnecessary bundle files. The only minor gap is the absence of explicit validation checkpoints in the workflow, though the task carries no destructive or batch risk that would demand them.

DimensionReasoningScore

Conciseness

The ~27-line body is lean with every token earning its place: "Reported gains cluster in the first two", "Time saved typing is often time spent reading". It assumes Claude's competence, explains nothing the model already knows, and includes no padding, matching anchor 5 exactly.

5 / 5

Actionability

As an instruction-only skill, the guidance is fully concrete: exact questions to ask ("Faster than what?"), specific diagnostic criteria (the "vibe coding" tell), a concrete recommendation ("a two-week pilot on one real backlog slice"), a defined output table, and a worked example. Per the scoring notes, absence of code is not penalized when guidance is this actionable — anchor 5.

5 / 5

Workflow Clarity

The five steps are clearly numbered and sequenced, each feeding the defined output table with an example. It sits at anchor 4 ('clear sequence... minor validation gaps') rather than 5 because no explicit validation/checkpoint steps exist; it stays above 3 since the sequence is fully coherent and the task is non-destructive, so the destructive-operation cap does not apply.

4 / 5

Progressive Disclosure

The skill is under 50 lines with no need for external references — verified: no references/, scripts/, or assets/ directories exist and the body references no files. Per the rubric's simple-skill guideline, well-organized sections (Steps, Output format, Example) alone warrant anchor 5.

5 / 5

Total

19

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly answers both what the skill does and when to use it, with natural trigger phrasing and a distinct niche. Its only weaknesses are limited enumeration of concrete sub-actions and the slightly dangling reference to "the evidence checklist from the talk", which assumes external context.

DimensionReasoningScore

Specificity

The description names its domain ("a vendor's productivity claim about an AI coding assistant") and one concrete action ("Evaluate... against the evidence checklist"), but does not enumerate multiple specific actions. Anchor 3 fits: domain plus 1-2 concrete actions, not comprehensive — it stays below 4 because no list of several specific capabilities is given, and above 2 because the action named is concrete rather than generic.

3 / 5

Completeness

Both parts are explicit: the what is "Evaluate a vendor's productivity claim about an AI coding assistant against the evidence checklist" and the when is "Use when someone quotes a percentage gain and asks whether the team should adopt the tool" — a concrete trigger scenario. The 'when' clause is as explicit as the anchor-5 example, so it does not fall to anchor 4 ('when could be more specific').

5 / 5

Trigger Term Quality

Natural trigger phrases users would actually say are present: "quotes a percentage gain", "asks whether the team should adopt the tool", "productivity claim". Coverage is good but common variations (e.g. "X% faster", "should we buy/license Copilot-style tools", "ROI") are missing, matching anchor 4 rather than the comprehensive synonym coverage of 5.

4 / 5

Distinctiveness Conflict Risk

The niche (evaluating vendor productivity claims about AI coding assistants) is clear with distinct triggers (percentage-gain quotes, adoption questions), so conflict with other skills is minimal. It clearly matches anchor 5 rather than anchor 4's 'minor overlap risk with closely related skills'.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
jbaruch/shownotes
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.