CtrlK
BlogDocsLog inGet started
Tessl Logo

confidence

Rates confidence that the current work fully solves the stated requirement. Supports plan validation, code review, and analysis (root-cause, refactor, diagnose) modes. Plan mode combines LLM judgment with deterministic rule checks (multi-signal gate); a failed rule caps the gate at 89% regardless of LLM score. Use before committing to autonomous execution, after implementation, or during investigation. Triggers on "confidence check", "validate plan", "rate confidence", "quality gate", "/confidence".

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced evaluation skill: every rule check is a verbatim command, the gate logic is unambiguous, and validation/feedback loops are explicit. Its weakness is token efficiency — the same justifications (verbatim-command notes, alias deprecation, 89% cap) are repeated multiple times, and the rules table could be split into a reference file.

Suggestions

State the no-shell-pipes/verbatim justification for the awk commands once (e.g., in a note above the rule table) instead of repeating it in rules #5, #6, #7, #9, and #10.

Consolidate the deprecated `bug-analysis` alias handling into a single paragraph — it is currently explained in the metadata tag comment, the mode table, a dedicated alias-handling paragraph, and again in the Fix Mode section.

Move the 11 deterministic rule-check commands to a one-level-deep reference file (e.g. references/plan-rules.md), keeping a rule-number/pass-fail summary table in SKILL.md to cut context cost for code and analysis invocations that never run them.

DimensionReasoningScore

Conciseness

Mostly efficient, load-bearing tables, but with real repetition: the justification "Single awk command with no shell pipes, so the table row executes verbatim" is repeated across five rules, the deprecated `bug-analysis` alias is explained in four separate places (metadata tag comment, mode table, alias-handling paragraph, Fix Mode section), and the 89% cap is restated in the callout, rule intro, Step 3, and Output Format. This fits anchor 3 (could be tightened) better than anchor 4's "minor instances".

3 / 5

Actionability

Fully executable guidance: copy-paste-ready `test -f`, `grep`, and single-command `awk` one-liners for all 11 deterministic rule checks, an exact output-format template with a worked table skeleton, explicit dimension weights, and numeric score thresholds. Specific examples cover the plan/code/analysis common cases — anchor 5.

5 / 5

Workflow Clarity

Plan mode is explicitly sequenced (Step 1 LLM dimensional scoring → Step 2 deterministic rule checks → Step 3 combined gate) with a hard validation checkpoint ("A failing rule is a blocker the gate must surface even if the LLM dimensional score is high"), an iteration protocol that re-runs the assessment after each of up to 2 iterations, and a Post-Fix Re-Assessment feedback loop. This matches anchor 5's explicit validation, feedback loops, and checklists.

5 / 5

Progressive Disclosure

No bundle files exist, and the single-file body is well navigated via a Contents list and clean section headers; the Output Format and Fix Mode sections are cleanly gated ("omit for code/analysis", "Skip this section entirely if not in Fix Mode"). The ~60-line rule-check table with long inline awk commands is a plausible candidate for a one-level-deep reference file, which keeps this at anchor 4 rather than 5.

4 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly covers what the skill does, when to use it, and how it works internally (multi-signal gate), with a dedicated trigger-phrase list. The main gaps are the unmentioned Fix Mode capability and a few missing natural trigger variations.

DimensionReasoningScore

Specificity

"Rates confidence that the current work fully solves the stated requirement. Supports plan validation, code review, and analysis (root-cause, refactor, diagnose) modes" names several concrete capabilities with mechanism detail ("a failed rule caps the gate at 89% regardless of LLM score"). Minor gaps — Fix Mode and the % report output are never mentioned — keep it at anchor 4 rather than 5, and it is well above the 1–2-action level of anchor 3.

4 / 5

Completeness

It clearly answers both questions: what ("Rates confidence that the current work fully solves the stated requirement") and when ("Use before committing to autonomous execution, after implementation, or during investigation") with concrete trigger phrases. This matches the anchor-5 example structure exactly; it is not anchor 4 because the 'when' is already explicit and specific rather than improvable.

5 / 5

Trigger Term Quality

The explicit list `Triggers on "confidence check", "validate plan", "rate confidence", "quality gate", "/confidence"` plus timing phrases ("before committing to autonomous execution, after implementation, or during investigation") gives good natural-term coverage. Common user variations (e.g. "is this right", "double-check this plan") are missing, so it does not reach the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

The confidence-rating niche is distinct with dedicated triggers ("rate confidence", "/confidence"), but "quality gate" and "code review" phrasing creates minor overlap with general review/quality skills — anchor 4 (mostly distinct, minor overlap) rather than 5 (minimal conflict risk).

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.