CtrlK
BlogDocsLog inGet started
Tessl Logo

reflection-coach

Reflect on another agent's recent execution record and propose the smallest durable instruction, skill, or tool-description change. Use for evidence-backed coaching proposals, never hot-swaps.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./packages/skills-catalog/catalog/bundled/paperclip-operations/reflection-coach/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality procedure skill: sequenced, gated, and deeply actionable, with concrete API commands, payload fields, and a verification checklist. Its one real weakness is repetition — the core safety rule is restated so often that the document is noticeably less token-efficient than it could be.

Suggestions

State the no-same-run-apply rule once in Hard guardrails and reference it from Pitfalls/step 10 instead of restating it five times — this alone would tighten conciseness without losing safety.

Move the proposal-document template and the server-enforced mutation target-key list into a references/ file (e.g. references/proposal-template.md and references/mutation-targets.md) linked from the relevant steps, keeping SKILL.md as a lean overview.

Trim editorial asides ('the discipline is the point', 'This is a lightweight stand-in for a real replay harness') — the gate itself is already clearly specified.

DimensionReasoningScore

Conciseness

Mostly dense and earning its tokens, but the no-same-run-apply rule is stated five separate times (intro, 'Two load-bearing rules', Hard guardrails, Pitfalls, step 10) and 'trajectories, not scores' is restated in Pitfalls; step 8 also editorializes ('the discipline is the point'). More than 'minor instances that could be trimmed', so it falls below the 4 anchor.

3 / 5

Actionability

Fully executable: copy-paste-ready curl commands with real endpoints and auth headers, exact request_confirmation payload fields (continuationPolicy, detailsMarkdown), the verbatim server-enforced mutation target keys, a concrete skill-sync JSON body, and a full markdown template for the proposal document. Specific examples cover the common cases.

5 / 5

Workflow Clarity

Ten explicitly numbered steps with hard guardrails, a benchmark-gate validation step (step 8: drop rules that would break past successes), an acceptance gate before any mutation (step 10), and a closing self-check checklist — exactly the explicit validation and feedback-loop discipline the rubric's top anchor requires for mutating operations.

5 / 5

Progressive Disclosure

No bundle files exist and nothing is nested or buried; sections are well-organized and one level deep. However, at ~190 lines, material such as the mutation target-key table and the full proposal-document template is inline where a reference file could keep SKILL.md leaner — 'good structure with minor organization gaps', clearly above the 3 anchor.

4 / 5

Total

17

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, third-person, with both a what and an explicit 'Use for' trigger plus a negative boundary. Its main gaps are on the trigger side — user-situation phrasing is thin and 'execution record' is platform jargon a requester may not echo verbatim.

Suggestions

Add user-side trigger situations to the 'Use for' clause, e.g. 'Use when an issue asks you to reflect on, coach, or review another agent's recent work' — concrete situations a requester would actually write.

Include one or two natural synonyms (e.g. 'retro', 'postmortem', 'agent performance') so the description matches how users phrase the request rather than platform jargon like 'execution record'.

Briefly name the output artifact (a reviewable change proposal with evidence quotes) so the what-side covers what the skill produces, not just what it examines.

DimensionReasoningScore

Specificity

Names the domain ('another agent's recent execution record') and several concrete actions — reflect on the record and propose 'the smallest durable instruction, skill, or tool-description change' across three named surfaces. Not a 5 because coverage is not comprehensive: it never says what the execution record contains or what the output proposal looks like.

4 / 5

Completeness

Both parts are present: a clear what ('Reflect on another agent's recent execution record and propose the smallest durable... change') and an explicit 'Use for...' clause. Not a 5 because the when-side describes the output ('evidence-backed coaching proposals') rather than concrete user-side trigger situations like 'Use when an issue asks you to reflect on or coach another agent'.

4 / 5

Trigger Term Quality

Good natural terms — 'Reflect on', 'coaching', 'evidence-backed', plus the negative trigger 'never hot-swaps'. A few natural phrases users would actually say are missing ('review an agent's work', 'postmortem', 'agent performance'), and 'execution record' leans on platform jargon, so it sits between the 3 and 5 anchors — noticeably above the midpoint.

4 / 5

Distinctiveness Conflict Risk

Clear niche — reflecting on and coaching another agent from its execution record — with the 'never hot-swaps' boundary adding distinctiveness. Minor overlap risk with generic review/retrospective skills keeps it below 5.

4 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 4 suspicious

Warning

Total

14

/

16

Passed

Repository
paperclipai/paperclip
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.