CtrlK
BlogDocsLog inGet started
Tessl Logo

build-review-interface

Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting structured feedback. Use when you need to build an annotation tool, review traces, or collect human labels.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/build-review-interface/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-organized, lean, and highly actionable, with concrete UI specs and a rigorous testing workflow including validation checkpoints. The main improvements are reducing redundancy between the Design Checklist and earlier sections and considering splitting the Testing/Additional Features detail into a reference file.

Suggestions

Trim the Design Checklist to items not already stated in Data Display / Feedback Collection / Navigation to reduce redundancy and lift conciseness.

Tighten the build sequence in the Overview into an explicit numbered list with validation checkpoints (e.g. "verify traces load before adding controls") to push workflow_clarity toward 5.

Consider moving the detailed 9-step Playwright test workflow into a references/testing.md file and linking to it, adding one-level-deep progressive disclosure.

DimensionReasoningScore

Conciseness

The body is lean and directive ("Emails should look like emails. Code should have syntax highlighting."), assuming Claude's competence without padding; it stays at 4 rather than 5 because the Design Checklist repeats several points already made in earlier sections (native-format rendering, full-trace access, auto-save, trace-level annotation).

4 / 5

Actionability

Concrete, specific guidance throughout — exact controls (Pass/Fail/Defer, free-text notes), explicit keyboard mappings, and a 9-step Playwright test workflow — provides mostly executable direction; it is an instruction-only skill so the absence of a code scaffold is acceptable, keeping it at 4.

4 / 5

Workflow Clarity

The Overview sequences the build (load traces → display → save labels → customize) and the Testing section gives a clearly numbered 9-step sequence with verification checkpoints ("verify labels persist", "verify content is accessible"), matching the 4 anchor with only minor validation gaps in the build steps themselves.

4 / 5

Progressive Disclosure

No bundle files exist and none are needed for this compact (~88 line) single-purpose skill, which is well-organized with clear section headers; it earns 4 for good self-contained structure rather than 5 because there is no progressive disclosure into deeper reference materials.

4 / 5

Total

16

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly states both capability and trigger conditions with natural, domain-specific keywords. The main weakness is the second-person voice ("you need to"), which the rubric penalizes, and slightly broad framing that risks minor overlap with generic web-building skills.

Suggestions

Rewrite in third person to avoid the second-person penalty: e.g. "Builds a custom browser-based annotation interface... Use when building an annotation tool, reviewing LLM traces, or collecting human labels."

Add a few more natural synonyms ("labeling tool", "evaluating LLM outputs", "golden dataset") to broaden trigger coverage.

Tighten the niche framing (e.g. mention "for LLM/agent traces") to further reduce overlap with generic web-app skills.

DimensionReasoningScore

Specificity

Lists several concrete actions ("Build a custom browser-based annotation interface", "reviewing LLM traces", "collecting structured feedback") which would rate a 4, but the second-person voice ("Use when you need to build") triggers the rubric's -1 specificity penalty, bringing it to 3.

3 / 5

Completeness

Explicitly answers both what ("Build a custom browser-based annotation interface... for reviewing LLM traces and collecting structured feedback") and when ("Use when you need to build an annotation tool, review traces, or collect human labels") with concrete trigger phrases, matching the 5 anchor.

5 / 5

Trigger Term Quality

Natural phrases a user would say ("annotation tool", "review traces", "collect human labels", "LLM traces") are present with good coverage, though synonyms like "labeling" or "evaluating LLM outputs" are missing, matching the 4 anchor.

4 / 5

Distinctiveness Conflict Risk

The LLM-trace annotation niche has distinct triggers and low conflict risk, but "build a custom browser-based annotation interface" is broad enough to overlap slightly with general web-app-building skills, so it sits at 4 rather than 5.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
hamelsmu/evals-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.