CtrlK
BlogDocsLog inGet started
Tessl Logo

fla-correctness-coverage

Guidelines for kernel correctness testing and coverage in fla/ops/** and related modules, including common Triton grid/addressing pitfalls. Helps decide what tests to add or run before an MR.

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/fla-correctness-coverage/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable body with excellent progressive disclosure — the reference files are clearly signaled, one level deep, and guarded by an explicit load-only-when-needed rule. The remaining gaps are minor: no mechanism or example for the 'list the coverage matrix' and 'add tests' workflow steps, and no failure-recovery guidance when tests fail.

Suggestions

Add a concrete command or snippet for workflow step 1 (e.g., how to enumerate existing tests/coverage for the op being touched) and a minimal new-test skeleton showing fixture/marker conventions.

Add a short failure-recovery loop to the Workflow section: what to do when a test fails or a safety check flags narrow program-ID arithmetic (fix, re-run the affected axes, re-run cross-backend shape).

Trim or relocate the 'What NOT to put in this skill' section — it is maintenance guidance about the skill itself, not task-execution content, and competes with the context window.

DimensionReasoningScore

Conciseness

The body is efficient — bullet lists, a coverage table, and copy-paste pytest commands with no explanation of concepts Claude already knows. It misses the 5 anchor because the 'What NOT to put in this skill' section is skill-maintenance meta-content that does not earn tokens during task execution, and a few checklist items could be tightened.

4 / 5

Actionability

Mostly executable: concrete pytest commands, exact helper names ('fla.utils.device', 'IS_NVIDIA', 'tl.make_tensor_descriptor'), and a concrete coverage-axes table. It falls short of fully executable because step 1 of the workflow ('List the current coverage matrix') and the 'add tests' step give no command or minimal test skeleton showing how to enumerate existing coverage or structure a new test.

4 / 5

Workflow Clarity

The Workflow section gives a clear 4-step sequence ending in a validation checkpoint ('Run the relevant tests and make sure they pass'), and the safety-checks section adds a concrete cross-backend verification requirement. It is not a 5 because there is no error-recovery loop (what to do when a test fails or a check flags int32 arithmetic) and no checklist tying the safety checks into the step sequence.

4 / 5

Progressive Disclosure

The body is a lean overview that lists each of the four existing references/ files with a one-line description, explicitly instructs 'read only the relevant reference file' and 'Do not load every reference by default' — well-signaled, one-level-deep navigation that matches the 5 anchor. All four referenced paths exist in the bundle (as one-line pointers to the operator READMEs).

5 / 5

Total

17

/

20

Passed

Description

73%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, clearly-scoped description with an explicit niche and solid trigger vocabulary. Its main weakness is that the 'what' is stated as abstract 'guidelines' rather than the concrete actions the skill actually performs, and the 'when' relies on an implied pre-MR situation rather than an explicit 'Use when...' clause.

Suggestions

Replace the abstract framing 'Guidelines for kernel correctness testing and coverage' with verb-first concrete actions, e.g. 'Checks kernel test coverage across varlen/backward/state axes, flags Triton int64 grid/addressing pitfalls, and runs the relevant pytest suites in fla/ops/**'.

Add an explicit trigger clause with natural phrases, e.g. 'Use when adding or modifying a Triton kernel in fla/ops/ (KDA, GDN, GLA, DeltaNet, NSA) or when deciding which correctness tests to run before an MR'.

Include one or two more user-natural terms such as 'linear attention', 'backward pass', or 'varlen' to broaden trigger coverage.

DimensionReasoningScore

Specificity

The description names the domain precisely ("kernel correctness testing and coverage in fla/ops/**", "common Triton grid/addressing pitfalls") and gives 1-2 actions ("Helps decide what tests to add or run before an MR"), but the actions themselves are generic — it never says it checks coverage axes, int64 address arithmetic, or device helper usage, so it sits at the 'domain plus 1-2 concrete actions, not comprehensive' anchor rather than 4.

3 / 5

Completeness

Both halves are present: the 'what' is explicit ("Guidelines for kernel correctness testing and coverage... including common Triton grid/addressing pitfalls") and the 'when' is given equivalent trigger guidance ("Helps decide what tests to add or run before an MR"). It is not the 5 anchor because there is no explicit 'Use when...' clause with concrete trigger phrases, and the 'when' could be more specific (e.g., 'when adding or modifying a kernel in fla/ops/').

4 / 5

Trigger Term Quality

Good natural keyword coverage — "kernel", "correctness", "testing", "coverage", "Triton", "MR", "tests" are all phrases a user touching fla kernels would say. A few natural variations are missing (e.g., "fla", "linear attention", "backward pass", "varlen"), which keeps it below the 5 anchor's synonym/extension completeness.

4 / 5

Distinctiveness Conflict Risk

The niche is unmistakable — "fla/ops/**", "Triton grid/addressing pitfalls", "before an MR" are unique to this domain, so it would not fire for unrelated testing or general Triton skills. Clear niche with distinct triggers, minimal conflict risk.

5 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
fla-org/flash-linear-attention
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.