CtrlK
BlogDocsLog inGet started
Tessl Logo

quality

Evaluates whether a GitHub issue is spam, empty, needs more information, or is OK to proceed.

58

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./tools/caretaker-agent/cloudrun/triage-worker/.gemini/skills/quality/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-structured, concise, and actionable for an instruction-only classification skill, with a clear output schema and a useful validation checkpoint. Its main gap is the absence of worked examples that would anchor the classification criteria.

Suggestions

Add 2-3 worked examples showing a sample issue and the expected JSON classification output to lift actionability to fully concrete.

Make the workflow sequence explicit (e.g. a short numbered list: read title/body, apply definitions, verify intent if OK, emit JSON) even though the task is simple.

Tighten the 'comment' field description by separating the required opening phrase from the variable missing-details guidance to reduce verbosity.

DimensionReasoningScore

Conciseness

The body is lean and task-specific, defining classification criteria and edge cases rather than padding with concepts Claude already knows; the only minor trim is the verbose comment-template instruction, fitting the 'efficient; minor instances of over-explanation' anchor.

4 / 5

Actionability

It gives a concrete JSON output schema with an exact enum and a specific comment template plus explicit edge-case rules (prompt injection -> SPAM), but it lacks worked examples mapping a real issue to a classification, leaving it just short of fully concrete.

4 / 5

Workflow Clarity

As a simple single-task classification skill it has an unambiguous action (analyze, classify, emit JSON) plus one validation checkpoint (verify user intent before classifying OK), but the sequence is implicit and there is no feedback loop, matching the 'clear sequence with most checkpoints present' anchor.

4 / 5

Progressive Disclosure

The skill is under 50 lines with no need for external references and is organized into clear sections (Instructions, Verification of User Intent, JSON Output Format, Quality Definitions), which per the guideline lets progressive disclosure score 5.

5 / 5

Total

17

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and third-person with a clear 'what', but it lacks any 'when to use' trigger guidance, capping completeness and limiting trigger-term coverage. It is distinguishable but would benefit from explicit use-when phrasing.

Suggestions

Add a 'Use when...' clause with natural trigger phrases, e.g. 'Use when triaging or moderating GitHub issues to detect spam, empty reports, or issues needing more information.'

Broaden trigger keywords to include synonyms users actually say such as 'triage', 'moderate', and 'review issue quality'.

Consider splitting the single 'Evaluates' verb into a couple of concrete actions (e.g. 'Classifies and triages GitHub issues...') to lift specificity.

DimensionReasoningScore

Specificity

Names the domain ('GitHub issue') and a concrete action ('Evaluates whether...') with an enumerated decision space (spam, empty, needs more information, OK), but it is a single action verb rather than several distinct actions, matching the 'names domain and 1-2 concrete actions' anchor.

3 / 5

Completeness

It clearly states what the skill does (evaluate issue quality into five outcomes) but provides no 'Use when...' clause or equivalent trigger guidance, which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Contains natural terms a user might say ('GitHub issue', 'spam') but misses common variations and synonyms such as 'triage', 'moderate', or 'review this issue', fitting the 'some relevant keywords but missing common variations' anchor.

3 / 5

Distinctiveness Conflict Risk

It carves a clear niche (GitHub issue quality classification) with low conflict risk, but the absence of an explicit trigger clause leaves minor overlap with general issue-management skills, matching the 'mostly distinct; minor overlap risk' anchor.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
google-gemini/gemini-cli
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.