CtrlK
BlogDocsLog inGet started
Tessl Logo

calibrate-claim-confidence

When the agent's epistemic state (GCCRF) indicates low empowerment and falling certainty, hedges out confident absolutes ("definitely", "always", "100%") in outgoing messages.

62

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/calibrate-claim-confidence/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-organized body for a simple automated interceptor: it explains the framework-specific state model, gives concrete before/after hedging examples, and points to the implementation with operational limits. The only real slack is one self-promotional sentence, and the hedging guidance could enumerate the mapping a little more completely.

DimensionReasoningScore

Conciseness

The body is lean — roughly 120 words with no padding or explanation of concepts Claude already knows — and the GCCRF framing is framework-specific knowledge Claude would not have. It misses 5 because the sentence 'This is the canonical example of state-binding: no other agent framework reads gccrf.empowerment to decide whether to hedge an outgoing statement' is a comparative marketing claim that could be trimmed without losing operational value.

4 / 5

Actionability

Concrete, specific guidance is present: exact word substitutions ('likely' instead of 'definitely', 'typically' instead of 'always', 'it appears' instead of 'obviously'), the conditions under which they apply, the implementation path, and the 6-fires-per-session limit. It falls short of 5 because the mapping is illustrative rather than exhaustive — no guidance on absolutes beyond the three examples, or on how heavily to hedge.

4 / 5

Workflow Clarity

This is a simple single-purpose skill (under 50 lines, one action), and per the rubric's simple-skill provision it can score 5 when the single action is unambiguous — which it is: when empowerment is low and certainty is falling, rewrite confident absolutes into hedged language, otherwise leave them intact. The modify-only-when-uncertain vs. leave-intact-when-confident branching is stated explicitly ('When it is confident (high empowerment, rising certainty), absolutes are left intact').

5 / 5

Progressive Disclosure

The skill is under 50 lines with no need for external references, and no bundle files (references/, scripts/, assets/) exist; per the rubric's provision, well-organized sections alone earn 5. The two sections ('What you'll see', 'Implementation') are appropriately scoped, and the single referenced path is an informational source location rather than a navigation hop.

5 / 5

Total

18

/

20

Passed

Description

63%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A distinctive, concrete, single-purpose description that clearly states what it does and the internal-state conditions under which it fires. Its main weakness is trigger-term quality: the 'when' is expressed purely in framework jargon (GCCRF, empowerment, certainty) rather than natural language a user or agent would recognize.

Suggestions

Rewrite the trigger in natural vocabulary alongside the state conditions, e.g. 'Use when the agent notices it is making confident absolute claims ("definitely", "always", "100%") while its knowledge is unverified or uncertain' — this preserves the state binding while adding recognizable trigger phrases.

Add common synonyms for the hedging domain (overclaiming, hedging, confidence calibration, uncertain ground) so the skill is discoverable beyond readers who already know the GCCRF terminology.

State the scope boundary explicitly (e.g. applies only to outgoing messages, not internal reasoning or tool inputs) to make the what/when split even sharper.

DimensionReasoningScore

Specificity

The description names a concrete action — "hedges out confident absolutes ('definitely', 'always', '100%') in outgoing messages" — with enumerated example terms. It stays at 4 rather than 5 because it covers a single action rather than the multiple comprehensive actions of the top anchor, and at 3 because the enumerated absolutes go beyond the '1-2 concrete actions' of the mid anchor.

4 / 5

Completeness

It explicitly answers both parts: the 'what' is hedging confident absolutes in outgoing messages, and the 'when' is present as an explicit clause — "When the agent's epistemic state (GCCRF) indicates low empowerment and falling certainty". It falls short of 5 because the 'when' is expressed in internal state jargon rather than concrete user-facing trigger phrases, matching the anchor where 'when' could be more explicit.

4 / 5

Trigger Term Quality

The trigger conditions are framed almost entirely in framework jargon — "epistemic state (GCCRF)", "low empowerment", "falling certainty" — that no user would naturally say; only the quoted absolutes ('definitely', 'always', '100%') are natural terms. It scores above 1 because those quoted words and "outgoing messages" are genuinely natural vocabulary, but well below 3-4 because the common user phrasings around hedging, confidence, or overclaiming are absent.

2 / 5

Distinctiveness Conflict Risk

This occupies a clear niche — state-conditioned hedging of outgoing messages keyed to GCCRF empowerment and certainty — with no plausible overlap with document, code, or data skills. The trigger conditions are so specific that wrong-skill activation risk is minimal.

5 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
Bitterbot-AI/bitterbot-desktop
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.