CtrlK
BlogDocsLog inGet started
Tessl Logo

test-observability

Integrate Playwright tests with OpenTelemetry, Grafana, Prometheus, Loki, and Tempo. Use when debugging test failures across distributed systems, measuring test performance, creating test dashboards, or correlating tests with backend traces.

62

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/test-observability/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill delivers strong, mostly executable guidance for wiring Playwright into an OTEL/Grafana stack, with clear sequenced workflows and a genuinely useful debugging example. Its main weaknesses are a bloated inline body that duplicates content promised (but not actually provided) in reference files, and a few code examples with undefined symbols that would not run as written.

Suggestions

Actually provide references/tracetest-integration.md, references/otel-reporter-setup.md, references/grafana-dashboards.md, and dashboards/test-results-dashboard.json — or remove the References section — since the body currently points users to files that do not exist.

Move the Tracestest integration, alerting, log correlation, and best-practices sections into the reference files, keeping SKILL.md to the Quick Start, architecture overview, and pointers.

Fix the broken examples: define/import OTLPLogTransport in the winston example and capture spanId in the trace-propagation example so both run as written.

DimensionReasoningScore

Conciseness

The body is mostly concrete code rather than concept explanation, but at ~430 lines it inlines advanced material (Tracetest setup, alerting rules, log correlation, best practices) that duplicates the content promised in the References section, plus padded sections like 'Benefits' ('80% faster debugging') and the closing tagline. Not the 'minor instances that could be trimmed' of anchor 4; several full sections could move to reference files.

3 / 5

Actionability

Most code blocks are concrete and copy-paste ready (install commands, playwright.config.ts snippets, PromQL queries, YAML alert rules), but a few have gaps: the winston example uses an 'OTLPLogTransport' that is never imported or defined, the trace-propagation example references an undefined 'spanId', and the dashboard import points to 'dashboards/test-results-dashboard.json' which does not exist in the bundle. These are real but limited gaps within otherwise executable guidance, matching anchor 4 rather than anchor 3 (the core paths run as written).

4 / 5

Workflow Clarity

Multi-step processes are clearly sequenced: Quick Start (install → configure → run → view in Grafana) and the debugging workflow (test fails → find trace → analyze → root cause) are both numbered and concrete, and the Railway section includes an explicit 'Verify Connection' curl check. It falls short of anchor 5 because the main Quick Start path has no checkpoint that the collector is actually receiving spans before sending the user to Grafana.

4 / 5

Progressive Disclosure

Sections are well organized and a dedicated References section clearly signals one-level-deep pointers ('references/tracetest-integration.md', 'references/otel-reporter-setup.md', 'references/grafana-dashboards.md', 'dashboards/test-results-dashboard.json'), but none of these files exist in the bundle — the navigation is dangling. Additionally, advanced content that the references are meant to hold is fully inlined, matching anchor 3 ('content that should be separate is inline') rather than anchor 4.

3 / 5

Total

14

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person-imperative, concise, with an explicit 'Use when' clause enumerating four concrete triggers and a well-defined niche. The only soft spots are the generic leading verb 'Integrate' and a few missing natural synonyms like 'flaky tests' or 'CI'.

DimensionReasoningScore

Specificity

It names the domain (Playwright tests) and stack (OpenTelemetry, Grafana, Prometheus, Loki, Tempo) and lists several concrete actions — 'debugging test failures', 'measuring test performance', 'creating test dashboards', 'correlating tests with backend traces' — but the headline action 'Integrate' is generic, leaving minor gaps versus the comprehensive multi-action anchor 5.

4 / 5

Completeness

It explicitly answers both: what ('Integrate Playwright tests with OpenTelemetry, Grafana, Prometheus, Loki, and Tempo') and when ('Use when debugging test failures across distributed systems, measuring test performance, creating test dashboards, or correlating tests with backend traces') with four concrete trigger phrases — a direct match for the anchor-5 example.

5 / 5

Trigger Term Quality

Good coverage of natural terms users would say — 'debugging test failures', 'test performance', 'test dashboards', 'backend traces', plus all five tool names — but common variations like 'flaky tests'/'flakiness', 'CI', 'monitoring', or 'OTEL' are missing, matching 'good keyword coverage; a few natural terms missing'.

4 / 5

Distinctiveness Conflict Risk

The Playwright + OpenTelemetry/Grafana-stack combination is a clear niche with mostly distinct triggers, but 'debugging test failures across distributed systems' overlaps with general debugging/tracing skills, so it is 'mostly distinct; minor overlap risk' rather than minimal.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.