CtrlK
BlogDocsLog inGet started
Tessl Logo

aip-tracker

Track Airflow Improvement Proposal (AIP) implementation progress by comparing Confluence specs against codebase evidence. Use when asked to assess, report on, or compare AIP status.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-crafted instruction skill: concrete tool usage, a tightly specified output contract, and a mandatory verification checklist that enforces its own rules. Its only costs are mild redundancy between the rules and the self-verification section, and a monolithic single-file layout that is slightly long for inline-only content.

DimensionReasoningScore

Conciseness

The body is dense and almost entirely non-obvious instruction (tool-call strategy, deliverable granularity rules, evidence-citation mandates, a fixed output template). The only notable padding is that the Self-Verification section restates several Assessment Rules verbatim (no percentages, no fabricated PRs, no editorializing), which is deliberate redundancy a tighter phrasing could reduce — 'efficient; minor instances that could be trimmed' (anchor 4), not the fully lean anchor 5.

4 / 5

Actionability

Guidance is fully concrete: named tools with call signatures and argument strategy ('search_github_prs with both the AIP number (e.g. "AIP-76") AND topic keywords (e.g. "asset partition")'), a priority-ordered extraction procedure, and a complete copy-paste output template with placeholders. For an instruction-only skill this matches the anchor-5 fully-executable bar.

5 / 5

Workflow Clarity

The sequence is explicit and ordered — three-source evidence gathering, deliverable extraction with priority rules, assessment rules, fixed output format — and it closes with a mandatory self-verification checklist including an arithmetic consistency check ('shipped + in_progress + not_started + beyond_spec = total'). This is the validate-and-fix feedback loop the anchor-5 example describes, and no destructive/batch cap applies.

5 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/), so all guidance lives in a single well-sectioned 143-line file. Sections are clear and in workflow order, but the file exceeds the 'under 50 lines, no external references needed' simple-skill case, and parts (e.g., the full report template or the self-verification checklist) could arguably live in a reference file. That minor organization gap fits anchor 4 better than 5.

4 / 5

Total

18

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit third-person what-statement plus a concrete 'Use when' trigger clause, with a clearly bounded niche. The only soft spot is that the what-statement covers a single composite action rather than enumerating the skill's several concrete capabilities.

Suggestions

Enumerate two or three concrete capabilities in the what-statement, e.g. 'Track AIP implementation progress: compare Confluence specs against code and PR evidence, produce per-AIP deliverable counts and cross-AIP progress reports.'

Add one or two natural trigger variations users might phrase differently, such as 'AIP progress' or 'AIP implementation status'.

DimensionReasoningScore

Specificity

The description names the domain (AIP implementation progress) and one concrete mechanism ('comparing Confluence specs against codebase evidence'), which matches the 'names domain and 1-2 concrete actions, but not comprehensive' anchor. It does not list multiple distinct capabilities (e.g., producing per-AIP breakdowns or cross-AIP reports), which keeps it below anchor 4.

3 / 5

Completeness

Both parts are explicit: 'Track Airflow Improvement Proposal (AIP) implementation progress by comparing Confluence specs against codebase evidence' answers what, and 'Use when asked to assess, report on, or compare AIP status' answers when with concrete trigger phrases — a direct match for the anchor-5 good example. It is not anchor 4 because the when-clause is already specific, not merely present-but-weak.

5 / 5

Trigger Term Quality

'Use when asked to assess, report on, or compare AIP status' provides good natural-verb coverage ('assess', 'report on', 'compare') plus the domain noun 'AIP status'. A few natural phrasing variations are missing (e.g., 'track AIP progress', 'implementation status'), so it falls just short of the comprehensive-synonyms anchor 5.

4 / 5

Distinctiveness Conflict Risk

The niche is narrow and unmistakable (Airflow Improvement Proposals, Confluence-spec-vs-code comparison), with trigger terms unlikely to fire for unrelated skills. This matches the 'clear niche with distinct triggers; minimal conflict risk' anchor.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
apache/airflow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.