CtrlK
BlogDocsLog inGet started
Tessl Logo

azure-diagnostics

Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, analyze logs, KQL, insights, image pull failures, cold start issues, health probe failures, resource health, root cause of errors, troubleshoot event hubs, troubleshoot service bus, messaging SDK error, AMQP connection failure, message lock lost, service bus dead letter.

59

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.github/plugins/azure-skills/skills/azure-diagnostics/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

53%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-organized with a useful service-routing table and concrete CLI/KQL examples, but it suffers from duplicated command and trigger sections, a diagnosis flow lacking validation checkpoints, and serious progressive-disclosure failures: half the routing links point to nonexistent files and none of the 12 bundled scripts are surfaced. Navigation structure exists on paper but does not match the actual bundle.

Suggestions

Fix or create the missing troubleshooting/ files for AKS, VM connectivity, and Messaging (or repoint those routes to existing references), since 3 of 6 service routes currently dead-end.

Reference the scripts/ bundle (e.g. a table mapping scripts like aks-baseline.sh, appservice-diagnostics.sh, test-messaging-connectivity.sh to their use cases) so the 12 bundled scripts are discoverable.

De-duplicate: remove the repeated 'az resource show' / 'az monitor activity-log list' block and collapse the Triggers, Rules, and Routing sections, which restate each other and the frontmatter description.

DimensionReasoningScore

Conciseness

Mostly efficient tables and command blocks, but with real duplication: 'az resource show' and 'az monitor activity-log list' appear verbatim in both 'Common Diagnostic Commands' and 'Check Azure Resource Health', the Triggers list restates the frontmatter description, the Rules section largely restates the Routing section, and the 'AUTHORITATIVE GUIDANCE - MANDATORY COMPLIANCE' blockquote adds no information. It fits 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than the 2-anchor, since the majority of the body is genuinely useful.

3 / 5

Actionability

Mostly executable guidance: concrete 'az' CLI commands and a working KQL snippet ('traces | where timestamp > ago(1h) | order by timestamp desc | take 50') with placeholders. The MCP blocks ('mcp_azure_mcp_applens' with intent/command/parameters key-value pseudo-notation) are descriptive rather than executable calls, which is a minor gap keeping it below the fully copy-paste-ready 5-anchor.

4 / 5

Workflow Clarity

The 'Quick Diagnosis Flow' presents a clear 5-step sequence (identify symptoms, check resource health, review logs, analyze metrics, investigate changes) but each step is a one-line question with no validation checkpoints or error-recovery guidance. This matches 'steps listed but validation gaps; checkpoints missing or implicit'. The operations are read-only diagnostics, so the destructive-operation cap does not apply, but the 4-anchor's 'most checkpoints present' is not met either.

3 / 5

Progressive Disclosure

Scored against the actual bundle: 3 of the 6 service routes (troubleshooting/aks/aks-troubleshooting.md, troubleshooting/compute/vm-troubleshooting.md, troubleshooting/messaging/README.md) point to a directory that does not exist, so AKS, VM, and Messaging navigation is broken. Additionally, all 12 scripts in scripts/ (aks-baseline, appservice-diagnostics, pod-evidence, test-messaging-connectivity, run-ig, etc.) are never mentioned in SKILL.md, leaving substantial bundle content undiscoverable. Despite a well-formed routing table, references are effectively buried/broken for half the services, fitting the 2-anchor; the 3-anchor's functional reference structure is not achieved.

2 / 5

Total

12

/

20

Passed

Description

86%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an exceptionally comprehensive, natural-language WHEN trigger list and an explicit what/when structure. Its main weaknesses are a somewhat terse 'what' (a single debug action plus tool names, with the redundant phrasing 'on Azure ... on Azure') and generic Kubernetes terms that slightly raise conflict risk with non-Azure Kubernetes skills.

Suggestions

List the covered services (App Service, Container Apps, Functions, AKS, VMs, Event Hubs/Service Bus) explicitly in the 'what' clause so capabilities are stated, not only implied by triggers.

Remove the duplicated 'on Azure' and scope Kubernetes trigger terms to AKS (e.g. 'AKS pod pending', 'AKS node not ready') to reduce overlap with generic Kubernetes skills.

DimensionReasoningScore

Specificity

The description names the domain and a concrete action ("Debug Azure production issues") plus tooling ("using AppLens, Azure Monitor, resource health, and safe triage"), but stops at 1-2 actions rather than listing several distinct capabilities. It matches the 'names domain and 1-2 concrete actions' anchor, not the 4-anchor which requires several specific actions.

3 / 5

Completeness

Both parts are explicit: the 'what' ("Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage") and an explicit 'WHEN:' clause with concrete trigger phrases. This clearly matches the anchor requiring both what and when with concrete triggers; the 4-anchor's caveat ('when' could be more explicit) does not apply since the WHEN clause is highly explicit.

5 / 5

Trigger Term Quality

The WHEN clause is a comprehensive list of natural user phrases with synonyms and variations: "troubleshoot app service", "app service high CPU", "VM black screen", "can't connect to VM", "reset VM password", "kubectl cannot connect", "crashloop", "message lock lost", "service bus dead letter". Coverage mirrors how users actually phrase these problems, matching the comprehensive-coverage anchor.

5 / 5

Distinctiveness Conflict Risk

The Azure production-diagnostics niche is clear and mostly distinct, but generic Kubernetes trigger terms ("kubectl cannot connect", "kube-system/CoreDNS failures", "crashloop", "node not ready", "pod pending") create minor overlap risk with any general Kubernetes-troubleshooting skill. This fits 'mostly distinct; minor overlap risk' rather than the minimal-conflict 5-anchor.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 8 missing, 13 deeper-than-1-level

Warning

referenced_paths_exist

Referenced path issues: 5 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
microsoft/azure-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.