CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-refinement

Agent skill for refinement - invoke with $agent-refinement

64

1.23x
Quality

46%

Does it follow best practices?

Impact

96%

1.23x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/agent-refinement/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, code-rich guide to the SPARC Refinement phase with a clearly sequenced TDD workflow and mostly executable examples, but it spends most of its tokens teaching patterns Claude already knows and inlines everything that belongs in reference files. A duplicate/stray YAML frontmatter block (name, hooks, capabilities) left in the body after line 4 adds confusion and should be removed or merged into the single real frontmatter.

Suggestions

Trim or drop the generic pattern tutorials (retry decorator, circuit breaker, complexity example) and keep only skill-specific guidance, or move them to references/error-handling.md and references/performance.md linked one level deep from SKILL.md.

Fix the code gaps that block copy-paste use: define or import sanitizeUser, generateToken, hash, SESSION_DURATION, and align the Green-phase constructor (logger) with the test setup.

Remove the stray second YAML frontmatter block (lines 6-30) from the body — or fold the hooks and capabilities into the real frontmatter — and add explicit validation commands (e.g., 'Run npm test; only proceed when green') between workflow steps.

DimensionReasoningScore

Conciseness

The ~530-line body extensively demonstrates concepts Claude already knows — TDD red/green/refactor, retry with exponential backoff, circuit breakers, cyclomatic complexity, and a generic authentication-service tutorial — matching 'Noticeably verbose; several unnecessary explanations or padded sections'. Not a 1 because the examples are relevant to the refinement phase rather than pure conceptual padding, and not a 3 because the bulk of the code teaches familiar patterns instead of adding skill-specific knowledge.

2 / 5

Actionability

Concrete, near-executable TypeScript/Jest examples dominate the body (mocked repositories, specific assertions like toHaveProperty and cache.set argument checks, a real SQL JOIN optimization, a Jest coverage config), matching 'Mostly executable guidance; concrete code or commands with minor gaps'. Not a 5 because snippets are illustrative rather than copy-paste ready — sanitizeUser, generateToken, hash, and this.SESSION_DURATION are undefined, and the Green-phase constructor requires a logger the test setup does not provide.

4 / 5

Workflow Clarity

The TDD cycle (Red - write failing tests, Green - implement to pass, Refactor - keep tests green) is clearly sequenced with the test itself acting as a validation checkpoint, and performance refinement follows identify-then-optimize steps, matching 'Clear sequence with most checkpoints present'. Not a 5 because no explicit run-the-suite commands or error-recovery steps appear in the body (they exist only in the misplaced hooks block), and not a 3 because the failing-test/pass-test loop is an inherent feedback loop.

4 / 5

Progressive Disclosure

Section headers (TDD Refinement Process, Performance Refinement, Error Handling Refinement, Quality Metrics, Best Practices) give the body structure, but all ~500 lines live inline in SKILL.md with no references to separate files and no bundle files exist, matching 'Some structure but could be better organized; content that should be separate is inline'. Not a 2 because the sections are clear and navigable rather than a wall of text; not a 4 because nothing is split out into reference files.

3 / 5

Total

13

/

20

Passed

Description

36%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a bare invocation stub: it names the skill's domain but says nothing about what it does, omits any 'Use when' trigger guidance, and offers no distinguishing keywords beyond the single word 'refinement'. It would rarely be surfaced for the right task because users asking to 'refactor', 'optimize', or 'improve' this code would not match it.

Suggestions

State concrete capabilities in third person, e.g., 'Performs the SPARC Refinement phase: writes failing tests (TDD), refactors code while keeping tests green, optimizes hot paths, and improves error handling.'

Add an explicit trigger clause, e.g., 'Use when the user asks to refine, refactor, optimize, or improve existing code, or when a SPARC workflow enters the Refinement phase.'

Include natural synonyms and concrete artifacts (code, tests, performance, coverage) so the description is distinguishable from generic code-review or testing skills.

DimensionReasoningScore

Specificity

The description 'Agent skill for refinement - invoke with $agent-refinement' names the domain (refinement) but contains no concrete action verbs whatsoever, matching the anchor 'Names the domain but actions are minimal or generic'. It is not a 1 because a domain is explicitly named rather than pure abstraction, and not a 3 because no 1-2 concrete actions (e.g., 'improves code quality through testing and refactoring') are listed.

2 / 5

Completeness

'Agent skill for refinement' is a vague what (refinement of what is unstated), and 'invoke with $agent-refinement' is an invocation instruction rather than a 'Use when...' trigger, matching 'Has a vague what and no when'. The missing-trigger guideline caps this dimension at 3, and the description sits below that cap because the what is also weak.

2 / 5

Trigger Term Quality

'refinement' is a term a user would naturally say, but there are no synonyms or variations such as refactor, optimize, clean up, improve code, or TDD, matching 'Some relevant keywords but missing common variations or synonyms'. Not a 4 because keyword coverage is essentially one word; not a 2 because 'refinement' is domain-specific rather than generic like 'works with files'.

3 / 5

Distinctiveness Conflict Risk

'refinement' alone is somewhat specific but overlaps with refactoring, code-review, testing, and even prompt-refinement skills, matching 'Somewhat specific but could still overlap with similar skills'. Not a 4 because no concrete triggers or scope (e.g., SPARC phase, code, testing) distinguish it; not a 2 because a domain is named rather than being entirely broad.

3 / 5

Total

10

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (530 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
ruvnet/ruflo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.