CtrlK
BlogDocsLog inGet started
Tessl Logo

outcome

This skill should be used when the user asks to "run the outcome eval", "paired bare vs protocol", "which decisions did the protocol surface", "count what the AI asked or presented", "does /inquire surface what the request left out", or wants to see, from transcripts, which decisions reached the user as a question or something to recognize instead of having to be written into the opening prompt. Project-local contributor tooling.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Excellent body: lean, entirely project-specific, with a fully executable runbook, explicit scope boundaries (what the eval does not measure, boundary with /realize), and clean progressive disclosure to a single real reference. The only soft spot is workflow clarity, where per-turn loop and validation details live in the referenced runbook rather than inline.

DimensionReasoningScore

Conciseness

The body is dense and project-specific throughout — the measure definition ('asked'/'presented'), the boundary with /realize, design constraints (frozen case sha256), and a compact runbook — with no explanation of concepts Claude already knows and no padding; every section earns its place, matching the 5 anchor.

5 / 5

Actionability

The Runbook block gives copy-paste-ready commands with real flags and variants ('node $S plan --runner claude --model claude-sonnet-5 --reps 2 --dry-run', codex auth variants, turn/note/report/teardown), plus a concrete test command and prerequisites — fully executable guidance for the common cases.

5 / 5

Workflow Clarity

The sequence (plan → setup → turn → note → report → teardown) is clear and in order, with checkpoints present (a --dry-run argument check, sha256 frozen-case refusal, a model-free test), but the per-turn loop mechanics and validation detail ('as the oracle says', how to write the notes' items) are delegated to the reference rather than stated inline — minor validation gaps put this at 4 rather than 5.

4 / 5

Progressive Disclosure

The body is a well-organized overview with one clearly signaled, one-level-deep reference ('Read references/runbook.md before the first run: the turn procedure, how to write the notes' items, authentication and isolation per runner…'), and the referenced file exists and itself contains no nested references; scripts are referenced by real paths in ./scripts/.

5 / 5

Total

19

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit use-when triggers in the user's own likely words, a concrete statement of what the skill shows and from where (transcripts), third-person voice, and a distinct niche. The only minor gap is that it describes a single workflow rather than enumerating several discrete actions.

DimensionReasoningScore

Specificity

The description states the concrete capability — 'wants to see, from transcripts, which decisions reached the user as a question or something to recognize instead of having to be written into the opening prompt' — including the paired bare-vs-protocol mechanism, but describes one workflow rather than several discrete actions, so it sits between the 3 and 5 anchors.

4 / 5

Completeness

Both 'what' (shows, from transcripts, which decisions reached the user as a question or something to recognize) and 'when' (an explicit 'This skill should be used when the user asks to…' clause with concrete trigger phrases) are clearly and explicitly answered, matching the 5 anchor.

5 / 5

Trigger Term Quality

It quotes five natural trigger phrases users would actually say ('run the outcome eval', 'paired bare vs protocol', 'which decisions did the protocol surface', 'count what the AI asked or presented', 'does /inquire surface what the request left out'), giving comprehensive coverage of phrasings.

5 / 5

Distinctiveness Conflict Risk

It occupies a clear niche (transcript-based outcome eval comparing bare vs protocol arms) with distinctive triggers like 'outcome eval', 'paired bare vs protocol', and '/inquire', and scopes itself as 'Project-local contributor tooling', so conflict risk with other skills is minimal.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
jongwony/epistemic-protocols
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.