CtrlK
BlogDocsLog inGet started
Tessl Logo

continual-learning

Nightly refinement of an existing per-repo review-style prompt using this reviewer's own finding outcomes. Read confirmed (resolved-by-commit / thumbs-up) and dismissed (thumbs-down) findings, promote the bug patterns the team actually fixes, demote the false-positive patterns, reconcile against the current prompt, and save the refined version. Use this once outcomes exist; use bootstrap-repo-analysis for a cold-start repo.

70

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-structured instruction-only skill: an unambiguous three-step workflow with named tool calls, concrete decision rules, and an explicit no-op guard for empty outcomes. Actionability is the only dimension short of top marks, due to the unspecified mechanism for retrieving the current custom_prompt and the absence of a worked save-call example.

DimensionReasoningScore

Conciseness

The body is lean: every line carries task-specific information Claude would not infer on its own ("A single dismissed finding is noise; the same class dismissed several times is a rule", "do not re-run a full PR crawl", the 400–1200 word constraint). There is no explanation of concepts Claude already knows and no padding, matching the score-5 anchor (every token earns its place); score 4 would require identifiable over-explanation, which is absent.

5 / 5

Actionability

Concrete guidance is present: named tool calls (`read_finding_outcomes`, `save_review_style_prompt`), the fields they return, the required payload fields (custom_prompt, analysis_summary, top_reviewers) with a word-count range and an example summary string. This matches the score-4 anchor (concrete commands with minor gaps) rather than 5, because no example of the actual save call with arguments is shown, and how the current custom_prompt is obtained ("it is summarized for you / available via the dashboard record") is left ambiguous.

4 / 5

Workflow Clarity

A clear three-step sequence (read outcomes → reconcile → save) with a decision heuristic ("Look for repetition, not one-offs") and an explicit guard for the destructive case — "If outcomes were empty and nothing changed... re-save the existing prompt unchanged rather than degrading it". This fits the score-4 anchor (clear sequence, most checkpoints present, minor validation gaps); it is not 5 because there is no post-save verification step or feedback loop, and not 3 because the main degradation risk is explicitly validated against.

4 / 5

Progressive Disclosure

The skill is under 50 lines with no bundle files (no references/, scripts/, or assets/ exist) and no external content is needed; the body is organized into three well-labeled sections that each fit on one screen. Per the rubric's simple-skill guideline (under 50 lines, no need for external references), this matches the score-5 anchor of well-organized structure with appropriate content placement.

5 / 5

Total

18

/

20

Passed

Description

85%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with fully concrete actions, an explicit use-when clause, and good disambiguation from the sibling cold-start skill. Its only weakness is that trigger terms lean on internal vocabulary (finding outcomes, confirmed/dismissed, review-style prompt) rather than the phrases a user would naturally type when reaching for this skill.

Suggestions

Add natural-language synonyms for the core concepts, e.g. "Use when the user wants to refine, update, or retrain the repo's review prompt from past feedback or findings".

Include a couple of plain-language trigger phrases users would actually say ("learn from dismissed findings", "stop repeating false positives") alongside the jargon terms.

Clarify who/what "this reviewer" refers to in the first sentence or rephrase to "the review agent's", since a user matching this skill may not share that referent.

DimensionReasoningScore

Specificity

Quotes: "Read confirmed (resolved-by-commit / thumbs-up) and dismissed (thumbs-down) findings, promote the bug patterns the team actually fixes, demote the false-positive patterns, reconcile against the current prompt, and save the refined version" — five concrete, domain-specific actions covering the full workflow. This matches the score-5 anchor (multiple specific concrete actions, comprehensive coverage) and exceeds the score-4 anchor only by being complete rather than having minor gaps.

5 / 5

Completeness

The what is explicit ("Read confirmed... and dismissed... findings, promote..., demote..., reconcile..., and save the refined version") and the when is explicit with a concrete trigger condition ("Use this once outcomes exist; use bootstrap-repo-analysis for a cold-start repo"). Both what and when are clearly answered, matching the score-5 anchor; score 4 would require the when to be less specific than it is here.

5 / 5

Trigger Term Quality

Relevant keywords exist — "refinement", "review-style prompt", "finding outcomes", "confirmed", "dismissed", "false-positive patterns" — but these are internal jargon rather than the natural phrases a user would say (e.g. "update the review prompt", "learn from past review feedback", "recurring bug patterns"). This fits the score-3 anchor (some relevant keywords, missing common variations/synonyms) and is below score 4, which requires broader natural-term coverage.

3 / 5

Distinctiveness Conflict Risk

The niche is sharply delimited — "Nightly refinement of an existing per-repo review-style prompt using this reviewer's own finding outcomes" — and it actively disambiguates from the nearest neighbor skill ("use bootstrap-repo-analysis for a cold-start repo"), matching the score-5 anchor of a clear niche with distinct triggers and minimal conflict risk.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
langchain-ai/open-swe
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.