CtrlK
BlogDocsLog inGet started
Tessl Logo

realize

This skill should be used when the user asks to "run the eval", "test whether the protocol actually works at runtime", "check type realization", "measure protocol fulfillment", "run the type-realization suite", "does the gate actually stop", or wants runtime evidence that a protocol's formal transition contract is realized rather than an assessment of downstream artifact quality. Invoke explicitly with /realize and a skill argument.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/realize/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable body for a complex eval harness, with concrete commands and clear sequencing. It is held back mainly by rationale prose that could be trimmed and by references to bundle paths that do not exist in the provided package.

Suggestions

Trim the philosophical/rationale prose — e.g. the multi-paragraph sham-arm justification citing published work and the Judgment boundary exposition — to only what is needed to execute the harness, improving token efficiency.

Make the integrity check an explicit validate→inspect→retry checkpoint inside the Workflow section (e.g. inspect the integrity column after run.sh and before teardown) rather than describing it only under 'Reading results'.

Fix or remove references to bundle paths that are not present (evals/scaffold.sh, evals/inquire-underspecified/, evals/inquire-fully-specified/, .github/workflows/type-realization.yml, harness.config.json) so navigation is not broken.

DimensionReasoningScore

Conciseness

Substantive and dense rather than padded with basics, but it carries multi-paragraph rationale (e.g. the sham-arm justification citing published work, the Judgment boundary prose) that assumes less than it could and could be tightened.

3 / 5

Actionability

The Workflow provides concrete, mostly executable bash commands (setup/run/teardown with Claude and Codex variants) and names real scripts, with only minor gaps such as the token placeholder and deferral to runbook.md.

4 / 5

Workflow Clarity

A clear setup→run→teardown→report sequence with integrity checks and non-zero exit on failure is present; validation is described in 'Reading results' rather than as an explicit inline checkpoint in the Workflow itself.

4 / 5

Progressive Disclosure

Well-organized sections with a clearly signaled one-level-deep 'Additional Resources' list, but several referenced paths (evals/scaffold.sh, evals/inquire-*, .github/workflows/type-realization.yml, harness.config.json) are absent from the bundle, leaving minor navigation gaps.

4 / 5

Total

15

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, highly specific description that explicitly covers both what the skill does and when to invoke it, with rich natural trigger phrasing and negligible conflict risk. Its only soft spot is specificity of enumerated actions, since it conveys a single core capability rather than a list of concrete operations.

DimensionReasoningScore

Specificity

States a clear core action ("wants runtime evidence that a protocol's formal transition contract is realized") and many trigger phrases, but does not enumerate multiple distinct concrete actions, matching the 'names domain and 1-2 concrete actions, not comprehensive' anchor.

3 / 5

Completeness

Explicitly answers both: a clear 'what' (runtime evidence of contract realization vs. downstream artifact quality) and an explicit 'when' ("This skill should be used when the user asks to…") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Lists numerous natural trigger phrases with synonyms ("run the eval" / "run the type-realization suite"; "check type realization" / "measure protocol fulfillment"), though several skew technical rather than conversational, sitting just below comprehensive.

4 / 5

Distinctiveness Conflict Risk

Occupies a sharply specialized niche ("/realize", "type-realization suite", "does the gate actually stop") with distinct triggers and minimal overlap with other skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
jongwony/epistemic-protocols
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.