CtrlK
BlogDocsLog inGet started
Tessl Logo

fla-optimization-loop

Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness. Synthesizes the task-contract / three-phase / iteration-protocol / silent-bug-catalog discipline of agent kernel-optimization frameworks (KDA, the MLSys FlashInfer contest workflow, AKO4ALL/AKO4X), and anchors all of it on FLA's frozen pytest (forward AND backward, under NaN poisoning) as the immutable correctness gate. Use when iterating on `fla/ops/**` performance over multiple rounds.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary process skill: a frozen-contract rule up front, a phased workflow with per-iteration validation and feedback loops, and concrete commands, paths, and templates throughout. The only cost is a handful of provenance asides that could be trimmed for token efficiency.

DimensionReasoningScore

Conciseness

The body is dense and assumes expertise throughout — no explanations of what Triton/NCU/pytest are, every section prescribes action. Minor trims are possible: the three '(Borrowed from KDA/AKO4ALL/AKO4X)' provenance asides and the inlined banned/allowed table edges add tokens without adding executable guidance, placing it at anchor 4 (efficient, minor instances that could be trimmed) rather than the fully lean anchor 5.

4 / 5

Actionability

Guidance is fully executable: exact commands ('python -m benchmarks.ops.verify --op <op> [--gate-k <subset>]'), concrete file paths (tests/ops/test_<op>.py, naive.py, docs/draft.md, OPT_LOG.md, dispatch.md), a named scratch-workspace layout, and a specific banned-vs-allowed table. Per the code_vs_instruction note, this instruction-only process skill needs concrete guidance rather than code, and it provides it copy-paste ready.

5 / 5

Workflow Clarity

The multi-step process is explicitly sequenced (task contract → three phases → iteration protocol → promotion) with a validation checkpoint embedded in every iteration ('Run verify.py — gate must stay green'), explicit feedback loops (stall handling after 3 no-improvement iterations, re-profile and re-assess), defined stopping criteria, a no-go bar with five named requirements, and a promotion checklist. This matches the anchor-5 pattern of clear sequence plus explicit validation and error-recovery loops.

5 / 5

Progressive Disclosure

The body is an overview of the loop that points to exactly two real, one-level-deep bundle files, both clearly signaled with their purpose ('Read references/TRAPS.md before trusting any number', templates in 'references/opt-log-template.md'), and both verified to exist with matching content. Detail is appropriately split (trap catalog and log templates live in references; the loop discipline lives here), matching the anchor-5 structure of well-signaled one-level-deep references.

5 / 5

Total

19

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-targeted description with an explicit and concrete 'Use when' trigger and a clearly stated purpose. Its main weakness is the middle provenance sentence naming external frameworks, which spends tokens on jargon instead of enumerating the skill's concrete capabilities.

Suggestions

Replace or compress the framework-provenance sentence ('Synthesizes the task-contract / three-phase / iteration-protocol ... AKO4ALL/AKO4X') with concrete capability verbs, e.g. 'Locks the op's pytest as a frozen correctness gate, runs profile-guided optimization in three phases, and logs every iteration' — this would raise specificity.

Add natural trigger synonyms such as 'optimize', 'speed up', or 'benchmark' an FLA kernel so users phrasing the need differently still match.

Keep the explicit 'Use when iterating on `fla/ops/**` performance over multiple rounds' clause — it is the strongest part of the description and should be preserved in any rewrite.

DimensionReasoningScore

Specificity

The description names the domain concretely ("making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness") but offers only one core capability; the middle sentence ("Synthesizes the task-contract / three-phase / iteration-protocol / silent-bug-catalog discipline of ... (KDA, the MLSys FlashInfer contest workflow, AKO4ALL/AKO4X)") is provenance framing rather than additional concrete actions. This matches anchor 3 (domain plus 1-2 concrete actions, not comprehensive) better than anchor 4, which expects several listed specific actions.

3 / 5

Completeness

It explicitly answers both: what ("Disciplined, reproducible loop for making an FLA kernel faster ... without ever breaking or gaming correctness") and when ("Use when iterating on `fla/ops/**` performance over multiple rounds") with a concrete trigger phrase. The 'when' clause is explicit and specific, matching the anchor-5 example structure rather than anchor 4's weaker 'when'.

5 / 5

Trigger Term Quality

Natural phrases a user would say are present — "making an FLA kernel faster", "iterating on `fla/ops/**` performance over multiple rounds" — plus backend names. Common synonyms like "optimize", "speed up", or "benchmark" are missing, and the framework-name jargon (KDA, AKO4X, NaN poisoning) adds no trigger value, so it sits at anchor 4 (good coverage, a few natural terms missing) rather than 5.

4 / 5

Distinctiveness Conflict Risk

A clear niche (FLA kernel optimization loops with a frozen correctness gate) with distinct triggers tied to `fla/ops/**` and multi-round iteration. Minor overlap risk remains with the closely related sibling performance skills (fla-nvidia-performance, fla-ascend-performance) whose trigger space also covers fla/ops performance work, so anchor 4 fits better than 5.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fla-org/flash-linear-attention
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.