CtrlK
BlogDocsLog inGet started
Tessl Logo

flags

Use when you need to check feature flag states, compare channels, or debug why a feature behaves differently across release channels.

88

1.03x
Quality

83%

Does it follow best practices?

Impact

98%

1.03x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-structured instruction-only skill: a concrete delegated command with fully enumerated options, channels, and output legend, plus a useful Common Mistakes section. The only nit is the absence of worked example invocations and the slightly implicit definition of 'meaningful' differences in --diff output.

DimensionReasoningScore

Conciseness

Every section earns its tokens: a five-row options table, a compact channel list, a one-line legend, three numbered steps, and two targeted mistakes. Nothing explains concepts Claude already knows and there is no padding, matching anchor 5 (lean and efficient; every token earns its place).

5 / 5

Actionability

The core command `yarn flags $ARGUMENTS` is copy-paste ready and the option syntax (`--diff <ch1> <ch2>`, `--cleanup`, `--csv`), channel names, and legend are concrete. It falls short of anchor 5 only because there are no worked example invocations (e.g., `yarn flags --diff www next`) and 'highlight meaningful differences' leaves what counts as 'meaningful' implicit — minor gaps typical of anchor 4.

4 / 5

Workflow Clarity

This is a simple, single-purpose skill under 50 lines whose single action is unambiguous: 'Run `yarn flags $ARGUMENTS`', then explain, with a --diff-specific follow-up — the rubric's simple-skill exception allows a 5. The read-only reporting command involves no destructive or batch operations, so no validation cap applies, and 'Common Mistakes' supplies the guidance a checklist would.

5 / 5

Progressive Disclosure

The skill is under 50 lines with no need for external references (no references/, scripts/, or assets/ exist, and the body references none), and its sections — Options, Channels, Legend, Instructions, Common Mistakes — are well-organized, which per the rubric's simple-skill note earns anchor 5. Nothing that belongs in a separate file is inlined.

5 / 5

Total

19

/

20

Passed

Description

73%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-targeted, reasonably concise description with a clear explicit trigger clause and a distinct niche. Its main weaknesses are second-person voice and action coverage that omits the skill's full option surface (cleanup grouping, CSV output), plus a few missing natural synonyms like 'toggle'.

Suggestions

Rewrite in third person to avoid the specificity penalty and separate the two concerns, e.g.: 'Checks feature flag states, compares flags across release channels, and groups flags by cleanup status or exports CSV. Use when the user mentions feature flags, flag states, channel differences (canary/next/experimental), or a feature behaving differently across channels.'

Add natural synonyms users might say — 'toggle', 'gate', 'experiment' — and concrete channel names (canary, next, experimental, rn) to broaden trigger coverage.

State the 'what' declaratively (list the actual capabilities: show all flags, --diff channels, cleanup status, --csv) instead of embedding it entirely inside the 'when' clause.

DimensionReasoningScore

Specificity

The description lists three concrete actions — "check feature flag states", "compare channels", "debug why a feature behaves differently across release channels" — which would otherwise land at 4, but it omits capabilities the body documents (listing all flags, cleanup status, CSV output) and uses second-person voice ("Use when you need to"), which the judging guidelines penalize by 1. The result is anchor 3: names the domain and 1-2 effective concrete actions without comprehensive coverage.

3 / 5

Completeness

Both what ("check feature flag states, compare channels, or debug...") and when ("Use when you need to...") are present in a single clause, so both questions are answered but not as two explicit, maximally concrete statements — matching anchor 4 rather than anchor 5. It is above anchor 3 because the when-clause is explicit, not merely implied.

4 / 5

Trigger Term Quality

Natural phrases like "feature flag states", "compare channels", "release channels", and "debug why a feature behaves differently" map well onto what a user would say. It misses common synonyms and variations (e.g., "flag", "toggle", "gate", specific channel names like canary/next), so it fits anchor 4 (good keyword coverage, a few natural terms missing) rather than anchor 5.

4 / 5

Distinctiveness Conflict Risk

The feature-flag/release-channel niche is distinct with trigger phrasing ("feature flag states", "release channels") unlikely to fire for unrelated skills — matching anchor 5's clear niche with distinct triggers and minimal conflict risk. Anchor 4's 'minor overlap risk with closely related skills' would apply only if a generic debugging or release-management skill existed with overlapping triggers, which the wording here avoids.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
facebook/react
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.