CtrlK
BlogDocsLog inGet started
Tessl Logo

test-observability

Integrate Playwright tests with OpenTelemetry, Grafana, Prometheus, Loki, and Tempo. Use when debugging test failures across distributed systems, measuring test performance, creating test dashboards, or correlating tests with backend traces.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is test-observability in fernandezbaptiste/Skrillz

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable, code-dense content with clear step sequences for setup and debugging. The main weaknesses are dangling references to bundle files that do not exist, a couple of non-executable code examples with undefined symbols, and redundant config blocks across sections.

Suggestions

Create the referenced bundle files (references/tracetest-integration.md, references/otel-reporter-setup.md, references/grafana-dashboards.md, dashboards/test-results-dashboard.json) or remove the References section, and move the advanced Tracetest/log-correlation/alerting sections into them to slim SKILL.md into a true overview.

Fix non-executable examples: define or import OTLPLogTransport in the Winston example and define spanId (or derive it from the active span context) in the trace-propagation example.

Consolidate the three near-identical playwright.config.ts reporter blocks into one and trim the marketing-style Benefits list and Overview bullets that duplicate the description.

DimensionReasoningScore

Conciseness

Largely code-driven with almost no explanation of concepts Claude already knows, but includes trimmable redundancy: three near-identical playwright.config.ts reporter blocks (Quick Start, Tracetest, Railway), a marketing-style Benefits list ('80% faster debugging'), and an Overview that repeats the frontmatter description. Not 5 due to this duplication; not 3 since there is no concept padding.

4 / 5

Actionability

Concrete npm installs, full configs, PromQL queries, YAML alert rules, and a curl verification command cover the common cases. Not 5 because the log-correlation example instantiates an undefined 'OTLPLogTransport', the trace-propagation example uses an undefined 'spanId', and the referenced dashboards/test-results-dashboard.json is not copy-paste usable because it does not exist in the bundle.

4 / 5

Workflow Clarity

Quick Start is a clean numbered 1-4 sequence and the debugging section is a concrete 4-step root-cause walkthrough; Railway includes a connection-verification curl checkpoint. Not 5 because the basic reporter setup and Grafana import steps lack explicit verification of success.

4 / 5

Progressive Disclosure

A dedicated References section clearly signals follow-up files, but none of the four referenced files (references/tracetest-integration.md, references/otel-reporter-setup.md, references/grafana-dashboards.md, dashboards/test-results-dashboard.json) exist in the bundle, and ~430 lines of advanced content that those files should hold (Tracetest, log correlation, alerting) are inlined in SKILL.md. Structure and signaling are present, keeping it above 2, but the dangling references and inline bulk keep it below 4.

3 / 5

Total

15

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description in third-person voice with an explicit 'Use when' trigger clause covering multiple concrete scenarios. Specific tool naming makes it highly distinct; only minor keyword variations (flakiness, CI) are missing.

DimensionReasoningScore

Specificity

Names the domain (Playwright + OpenTelemetry/Grafana/Prometheus/Loki/Tempo) and several concrete actions via the when-clause ('debugging test failures', 'measuring test performance', 'creating test dashboards', 'correlating tests with backend traces'). Falls short of 5 because the lead verb 'Integrate' is generic and capabilities like log aggregation, alerting, and trace-based assertions are absent.

4 / 5

Completeness

Explicitly answers both: 'Integrate Playwright tests with OpenTelemetry, Grafana, Prometheus, Loki, and Tempo' (what) and 'Use when debugging test failures..., measuring test performance, creating test dashboards, or correlating tests with backend traces' (when, with concrete trigger phrases). Matches the top anchor.

5 / 5

Trigger Term Quality

Good coverage of natural phrases users would say: 'Playwright tests', 'test failures', 'test performance', 'test dashboards', 'backend traces', plus tool names. Missing common variations such as 'flaky tests', 'CI failures', or 'OTel', so not a 5.

4 / 5

Distinctiveness Conflict Risk

Clear niche combining Playwright with a specific observability stack; the trigger scenarios (test debugging across distributed systems, test dashboards, trace correlation) are unlikely to fire for unrelated skills.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.