CtrlK
BlogDocsLog inGet started
Tessl Logo

aif-grounded

Reliability gate for answers. Forces evidence-based reasoning, explicit uncertainty, and “insufficient information” instead of guesses. Use when user says “be 100% sure”, “no hallucinations”, “only if verified”, “grounded answer”, or when stakes are high.

68

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a crisp, gated workflow with concrete rules and copy-paste-ready output templates, and its sequencing with an explicit confidence checkpoint is exemplary. The main weakness is verbosity in the skill-context and policy sections, which repeat emphasis and read as boilerplate.

Suggestions

Condense Step 0's skill-context instructions to 3-4 bullet points: the current version spends ~15 lines repeatedly emphasizing compliance ('CRITICAL', 'Enforcement') where one 'skill-context rules override these defaults and apply to all outputs' statement plus the enforcement note would suffice.

Trim or merge the 'Artifact Ownership and Config Policy' section into a single sentence under a general notes heading, since 'no artifact ownership' and 'config-agnostic' carry little operational weight for a gate skill.

State the retry loop explicitly in the workflow (e.g., 'After the user supplies the missing evidence, restart at Step 2') so the feedback path is a sequenced step rather than only implied by the output template.

DimensionReasoningScore

Conciseness

The workflow core (Steps 1-5, output templates, guardrail) is efficient, but Step 0 belabors skill-context rule application across multiple repetitive paragraphs ("CRITICAL: ... you MUST comply", the separate "Enforcement" paragraph restating it), and the Artifact/Config policy section is boilerplate. Mostly efficient with unnecessary explanation that could be tightened fits anchor 3; not 4 because the redundancy exceeds minor instances, not 2 because there is no padding that explains concepts Claude already knows.

3 / 5

Actionability

For an instruction-only skill the guidance is fully executable: a concrete three-way classification taxonomy, hard rules for evidence vs. unknowns, a numeric confidence-gate rule, and two copy-paste-ready output format templates covering both gate outcomes. Matches the top anchor; not 4 because the templates leave no gaps for the common cases.

5 / 5

Workflow Clarity

Steps 0-5 are clearly sequenced with an explicit validation checkpoint (the Step 4 confidence gate requiring zero open unknowns) and a feedback loop ("To reach 100: 1-3 concrete asks or commands for the user to run and paste output"). The implementation guardrail adds validate-before-act protection; this is not a destructive/batch skill so no cap applies. Fits anchor 5 rather than 4 because both the gate and the retry path are explicit.

5 / 5

Progressive Disclosure

The single-file skill (~110 lines) is well-sectioned with no bundle files, and the only reference (.ai-factory/skill-context/aif-grounded/SKILL.md) is one level deep and clearly signaled. Good structure with minor organization gaps — the lengthy Step 0 block could be condensed — fits anchor 4; not 5 because the file exceeds the simple-skill threshold and contains a padded section, not 3 because nothing that belongs in a separate file is inlined.

4 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities, an explicit 'Use when' clause with quoted natural trigger phrases, and a clearly distinct reliability-gating niche. The only gaps are a few missing natural trigger synonyms and slight breadth in "when stakes are high".

DimensionReasoningScore

Specificity

"Forces evidence-based reasoning, explicit uncertainty, and 'insufficient information' instead of guesses" lists three concrete behaviors, matching the several-specific-actions anchor with minor gaps (confidence-gate and output-format mechanics are not mentioned). Not 5 because coverage is not comprehensive; not 3 because it goes beyond 1-2 actions.

4 / 5

Completeness

The description explicitly answers both what ("Reliability gate for answers. Forces evidence-based reasoning, explicit uncertainty, and 'insufficient information' instead of guesses") and when ("Use when user says 'be 100% sure', 'no hallucinations'...") with concrete trigger phrases, matching the top anchor. Not 4 because the 'when' clause is already fully explicit rather than needing more specificity.

5 / 5

Trigger Term Quality

Quoted phrases like "be 100% sure", "no hallucinations", "only if verified", and "grounded answer" are natural things a user would say, but common variants such as "don't guess", "cite your sources", or "no assumptions" are missing. Good coverage with a few natural terms absent fits anchor 4 rather than the comprehensive synonym coverage of 5.

4 / 5

Distinctiveness Conflict Risk

"Reliability gate for answers" with triggers like "no hallucinations" and "grounded answer" carves a mostly distinct epistemic-reliability niche with minor overlap risk against research/verification skills, and "when stakes are high" is somewhat broad. Fits anchor 4; not 5 due to that residual overlap, not 3 since the triggers are distinctive rather than merely somewhat specific.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
lee-to/ai-factory
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.