CtrlK
BlogDocsLog inGet started
Tessl Logo

trace-review

Use when asked to review skill telemetry, find week-over-week skill regressions, run a weekly Langfuse trace report, or check whether a skill got slower / more expensive / more error-prone - produces a per-skill regression report from the local Langfuse trace store and can schedule itself.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is trace-review in AndreJorgeLopes/devflow

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered skill body: executable commands with exact flags and thresholds, an explicit flow with validation checkpoints and degraded-mode rules, and dense non-generic domain knowledge. Weak points are mild redundancy across the mistake table and earlier sections, and a TEMPLATES.md reference that is not resolvable within the provided bundle.

Suggestions

Deduplicate the Common Mistakes table against 'Core pattern'/'Running it' — rows like the skill.name and degraded-data rules restate already-explained content and could be shortened to the delta only.

Make the TEMPLATES.md reference a proper markdown link and ensure the file actually ships in the skill bundle (e.g. references/TEMPLATES.md), since the report format and scheduler-extension guide are currently unreachable from the skill directory.

DimensionReasoningScore

Conciseness

The body is dense and almost entirely non-obvious domain knowledge (attribution ladder, per-trace cost fields, thresholds, scheduler trade-offs) with no explanations of concepts Claude already knows, but the Common Mistakes table restates rules already given in 'Core pattern', 'Running it', and 'Scheduling', and the mermaid + numbered attribution explanation carry some redundancy. Not 5: those duplicated statements could be trimmed without losing clarity.

4 / 5

Actionability

Copy-paste-ready commands cover the common cases: `devflow trace-review run`, `--json`, `--output reports/trace-review-$(date -u +%F).md`, `--window 14`, `schedule --backend cron --cron "0 9 * * 1"`, plus concrete seeding scripts (`eval/lib/tessl-push.sh <skill> <score-0-100>`) and exact thresholds and env overrides. Not 4: the commands are complete and specific down to arguments and defaults.

5 / 5

Workflow Clarity

The mermaid flow gives a clear sequenced flow with explicit validation checkpoints and error-recovery loops: preflight dependency check ("Required missing → STOP"), Langfuse-reachable branch with "STOP: devflow up, then retry", degraded-data rules ("Render the report anyway", NEW rows, `-` scores), and an AskUserQuestion decision for save/schedule. Not 4: feedback loops and decision points are explicit rather than implicit; the operation is read-only so no destructive-validation cap applies.

5 / 5

Progressive Disclosure

Good structure with clear section headers and a one-level-deep, well-signaled reference ("See `TEMPLATES.md` for the exact report format and the scheduler-extension guide"), keeping the report format out of the overview. Not 5: `TEMPLATES.md` is cited as plain text rather than a link and is not present in the skill's bundle (no references/scripts/assets directories exist here), and the other cited paths (`requirements.json`, `lib/trace-review.py`, `eval/lib/*.sh`) are likewise not verifiable in the bundle.

4 / 5

Total

18

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when' trigger list, concrete actions, and a clear statement of what is produced. Third-person voice is maintained and there is no fluff. The only gaps are a few missing natural trigger synonyms and minor overlap risk with sibling skill-review skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "find week-over-week skill regressions", "run a weekly Langfuse trace report", "check whether a skill got slower / more expensive / more error-prone", "produces a per-skill regression report from the local Langfuse trace store", "can schedule itself" — with the data source named. Not 4: the action list is comprehensive for this domain rather than having coverage gaps.

5 / 5

Completeness

Explicitly answers both: "Use when asked to review skill telemetry, ... or check whether a skill got slower / more expensive / more error-prone" (when) and "produces a per-skill regression report from the local Langfuse trace store and can schedule itself" (what). Not 4: the 'when' is given as a concrete trigger-phrase list, not merely implied or generic.

5 / 5

Trigger Term Quality

Good natural-phrase coverage: "review skill telemetry", "skill regressions", "weekly Langfuse trace report", "slower / more expensive / more error-prone" are phrases users would plausibly say. Not 5: common variations like "skill health", "did any skill regress", or "trace report" alone are absent.

4 / 5

Distinctiveness Conflict Risk

The Langfuse/telemetry/regression niche is mostly distinct with specific triggers, but "review skill telemetry" and regression-reporting could overlap with closely related skill-quality or review skills in the same suite. Not 5: minor overlap risk remains with generic skill-review/quality skills.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
AndreJorgeLopes/devflow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.