CtrlK
BlogDocsLog inGet started
Tessl Logo

fla-nvidia-performance

Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo. Covers profiling workflow, hardware baselines, and MR-ready performance evidence requirements. Uses an installed ncu-report-skill when a task needs detailed Nsight Compute collection and diagnosis.

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary lean, action-oriented skill body: concrete NCU and benchmark commands, a clear pre-MR evidence checklist, and clean self-contained organization. The only meaningful gap is the absence of an explicit recovery loop when a regression or failed profile run is detected.

DimensionReasoningScore

Conciseness

Lean bullet-based sections with no explanation of concepts Claude already knows; every line carries repo-specific value (hardware baseline policy, artifact layout, exact commands), so every token earns its place.

5 / 5

Actionability

Fully executable, copy-paste-ready commands: complete ncu invocations with sections and kernel-regex flags, concrete benchmark commands with real flags, and a filled-in artifact example (profile/kda_chunk_bwd_20250603/); placeholders like <kernel_regex> are necessary parameterization, not pseudocode.

5 / 5

Workflow Clarity

Clear sequencing from day-to-day sanity checks to the numbered four-part pre-MR evidence checklist, with checkpoints (same-script before/after comparison, NCU-unavailable fallback), but no explicit fix-and-re-validate feedback loop when a regression is found — only 'Flag any backend or shape that regressed and explain why'.

4 / 5

Progressive Disclosure

A self-contained ~110-line skill with well-organized sections; nothing inlined that belongs in a separate file, no bundle files needed, and the delegation to the external user-level ncu-report-skill is clearly signaled.

5 / 5

Total

19

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-scoped description with strong domain keywords and minimal conflict risk, but it describes topic coverage rather than actions and entirely lacks an explicit 'when to use this skill' trigger clause. Adding a 'Use when...' sentence would raise both completeness and action framing.

Suggestions

Add an explicit trigger clause, e.g. 'Use when optimizing or benchmarking Triton/Gluon/TileLang/CUDA kernels in this repo, or when an MR touches fla/ops/ performance.'

Rephrase topic nouns as concrete actions, e.g. 'Profile kernels with NCU, run throughput benchmarks, and assemble MR-ready performance evidence.'

Include natural user vocabulary such as 'benchmark', 'throughput', 'latency', and 'GPU optimization' alongside the framework names to improve trigger-term coverage.

DimensionReasoningScore

Specificity

Lists several concrete coverage items ('profiling workflow, hardware baselines, and MR-ready performance evidence requirements', 'NCU collection and diagnosis'), but they are framed as topics ('Guidelines for...', 'Covers...') rather than executable actions, leaving minor gaps versus the comprehensive anchor.

4 / 5

Completeness

The 'what' is clear (profiling workflow, hardware baselines, MR evidence requirements), but there is no 'Use when...' clause or equivalent explicit trigger guidance; the only 'when' phrase ('when a task needs detailed Nsight Compute collection') scopes the external sub-skill, not this skill, so completeness is capped at 3.

3 / 5

Trigger Term Quality

Good natural-term coverage ('NVIDIA GPU kernel', 'Triton', 'CUDA', 'performance', 'profiling', 'Nsight Compute'), but common variations users would actually say such as 'benchmark', 'throughput', and 'optimize/speed up' are missing.

4 / 5

Distinctiveness Conflict Risk

Clear niche with distinct triggers — 'FLA repo', Triton/Gluon/TileLang/CUDA kernel work, NCU/Nsight Compute — making overlap with other skills minimal.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fla-org/flash-linear-attention
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.