CtrlK
BlogDocsLog inGet started
Tessl Logo

dt-obs-genai

Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.

67

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-architected overview skill: executable DQL per capability, clean one-level-deep reference structure with all referenced files present, and clearly sequenced investigation workflows. Minor conciseness gains are available by trimming prose intros and the inlined empty-state block.

Suggestions

Move the two empty-state presence-check queries and their explanatory paragraph into a reference file (e.g., references/empty-state.md), keeping only the rule and a pointer inline.

Tighten the prose intro under each Core Capability to one sentence so the DQL example leads sooner.

Add an explicit 'if no rows, run the empty-state check before responding' validation step at the start of each investigation workflow to close the feedback-loop gap.

DimensionReasoningScore

Conciseness

Mostly efficient with lean, copy-paste DQL examples and no basic-concept padding, though several capability prose intros and the fully inlined Empty-State Check block (two queries plus a long explanation) could be trimmed or moved to a reference.

4 / 5

Actionability

Every capability ships a complete, executable DQL query and the workflows reference named queries in the reference files; the guidance is copy-paste ready and covers the common cases.

5 / 5

Workflow Clarity

Five well-sequenced numbered investigation workflows with the Empty-State Check serving as a validation checkpoint; read-only analysis so the destructive-cap does not apply, but the workflows lack explicit error-recovery feedback loops for a 5.

4 / 5

Progressive Disclosure

SKILL.md is a clear overview with one starter query per capability and well-signaled one-level-deep '→ See [references/X.md]' links; all seven referenced files exist and are catalogued in a final References section, giving easy navigation.

5 / 5

Total

18

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, comprehensive description of GenAI/LLM debugging capabilities with strong distinctiveness, but it lacks an explicit 'Use when...' trigger clause, which caps its completeness.

Suggestions

Add an explicit 'Use when...' clause naming natural trigger phrases (e.g., 'Use when investigating LLM latency, token cost, agent failures, or evaluation quality in GenAI applications').

Soften platform jargon ('dt-evals') with a user-natural synonym so the trigger reads as something a user would actually say.

Include a couple of common synonyms (e.g., 'rate limits / throttling', 'safety filtering') to broaden natural trigger coverage.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions — 'token cost & caching by prompt, model & provider', 'latency/errors', 'agent & tool loops/failures', 'conversations', 'guardrails', 'evaluations', 'OpenTelemetry/dt-evals setup' — with comprehensive coverage of the GenAI debugging domain.

5 / 5

Completeness

The 'what' is explicit and detailed, but there is no 'Use when...' clause or equivalent trigger guidance; per the judging guideline a missing explicit 'when' caps completeness at 3.

3 / 5

Trigger Term Quality

Good natural keywords ('token cost & caching', 'latency/errors', 'agent & tool loops/failures') with the GenAI/LLM synonym pair, but jargon-heavy ('OpenTelemetry/dt-evals') and missing common synonyms/extensions, so below the comprehensive anchor.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (GenAI/LLM app observability) tied to a specific platform via 'OpenTelemetry/dt-evals setup', with distinct triggers and minimal overlap risk with generic observability skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Dynatrace/dynatrace-for-ai
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.