CtrlK
BlogDocsLog inGet started
Tessl Logo

perf-analysis

Interpret pinned managed benchmark evidence for the /review performance GitHub Agentic Workflow. Produce a narrative for independently validated reporting; never execute measurements or publish directly.

69

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

90%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary operational skill body: token-efficient, densely actionable with exact field names, thresholds, enums, and a copy-paste output template, and a clearly phased workflow with explicit completeness rules. The only weaknesses are a missing fix-and-retry loop for narrative validation and two under-signaled bundle references.

Suggestions

Add an explicit error-recovery step to Phase 6: if Validate-PerformanceReport.ps1 rejects the narrative, name which fields to revise and re-submit rather than leaving the fix loop implicit.

Signal `references/platform-scenarios.json` by path where Phase 4 cites 'the catalog', and mention `references/benchmark-families.json` where sampled families are classified, so all reference files are discoverable from SKILL.md.

DimensionReasoningScore

Conciseness

Lean and telegraphic throughout: no concept explanations, no filler, and it assumes Claude's competence (uses "ABBA", "library-TFM", "captured closures, boxing" without glossing them). Every sentence carries an operational constraint or decision rule — e.g., "Timing flags require non-overlapping run-level ranges and at least a 15% median delta" — so every token earns its place.

5 / 5

Actionability

Fully executable instruction-only guidance: an exact evidence file table for Phase 0, concrete field paths (`.suites[]`, `.sampledProductFiles[]`, `.coverage`), precise thresholds (non-overlapping byte gap, 15% median delta), closed enums (`none`/`warning`/`error`, `accidental`/`deliberate`/`unknown`), named validating scripts, and a copy-paste-ready JSON output template with all required fields. Specific guidance covers the common cases completely.

5 / 5

Workflow Clarity

Phases 0–6 are clearly sequenced with per-phase inputs and outputs, and completeness checkpoints are explicit ("Missing evidence, failed execution, or mismatched identities mean incomplete, never clean") plus a named external validator ("Validate-PerformanceReport.ps1 checks them independently"). Falls short of 5 because there is no fix-and-revalidate feedback loop for the skill's own output — nothing instructs what to do if the narrative fails validation.

4 / 5

Progressive Disclosure

Good structure: well-headed phases, a one-level-deep explicit reference ("Read `references/recommendation-policy.json` from the trusted skill directory"), and data appropriately split into the three real reference JSON files and six scripts. Minor gaps keep it below 5: Phase 4 refers to "the platform scenario catalog" without its path (`references/platform-scenarios.json`), and `references/benchmark-families.json` is never mentioned in the body.

4 / 5

Total

18

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, disciplined description: concrete third-person actions, an explicit workflow trigger context, and a sharp boundary statement that eliminates overlap with measurement or publishing skills. Its main gaps are that the 'when' is embedded in a 'for...' phrase rather than an explicit trigger clause, and a few natural synonyms (regression, profiling, PR review) are absent.

Suggestions

Add an explicit trigger clause (e.g., "Use when the /review performance workflow supplies a pinned benchmark evidence bundle") so the 'when' is stated as trigger guidance rather than carried by the 'for the ... Agentic Workflow' phrase.

Include natural trigger synonyms such as "benchmark regression", "perf evidence", or "pull request performance review" to round out keyword coverage beyond 'benchmark' and 'performance'.

DimensionReasoningScore

Specificity

Lists several concrete actions — "Interpret pinned managed benchmark evidence", "Produce a narrative for independently validated reporting", "never execute measurements or publish directly" — in third person with no vague filler. Falls short of 5 because major activities from the body (coverage classification, static hot-path review, native-gap reporting) are not surfaced, and the actions are umbrella-level rather than fully enumerated.

4 / 5

Completeness

The 'what' is clear and explicit (interpret pinned evidence, produce a validated narrative, never measure or publish directly). The 'when' is present via "for the /review performance GitHub Agentic Workflow" — explicit trigger context naming the invoking workflow — but it is not phrased as a "Use when..." clause and could be more specific, matching anchor 4 rather than 5; it clearly exceeds anchor 3 because the trigger context is stated, not merely implied.

4 / 5

Trigger Term Quality

Good keyword coverage: "benchmark", "performance", "evidence", "narrative", "reporting", "measurements", plus the concrete workflow identifier "/review performance GitHub Agentic Workflow". A few natural terms are missing — users would also say "regression", "profiling", or "pull request/PR performance review" — so it sits between anchor 4 and the synonym-complete anchor 5.

4 / 5

Distinctiveness Conflict Risk

Clear niche with distinct triggers: "pinned managed benchmark evidence for the /review performance GitHub Agentic Workflow" plus the explicit boundary "never execute measurements or publish directly" separates it sharply from generic profiling, benchmarking, and code-review skills. Minimal conflict risk.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
dotnet/maui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.