CtrlK
BlogDocsLog inGet started
Tessl Logo

elicit

The user knows roughly what they want but not which decisions it turns on: trace them from their own material (code, rules, past sessions) and the domain's usual decisions; ask until it settles.

49

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./euporia/skills/elicit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

42%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill encodes an unusually rigorous, self-consistent dialogue contract, and its one reference file is well signaled, but the delivery works against it: a ~340-line Lean formal block dominates the file at massive token cost, and the executable guidance (what a surface actually looks like) is never demonstrated. The prose sections that remain are terse and specific, so the deficit is structural placement and verbosity rather than vagueness.

Suggestions

Move the Lean formal block to a reference file (e.g., references/contract.md), keeping in SKILL.md only the FLOW summary, the gate rules, and the surface composition requirements — the formal encoding is occasion-bound verification material, not overview content.

Add one worked example of a surface round (seed utterance → surfaced coordinates with sources → answer slots) so the abstract record/focus rules become visibly actionable; this is the single highest-leverage fix for both actionability and workflow clarity.

Compress the axiom doc comments: rules like StandingSupported are restated across the Definition, the axiom comment, and TOOL GROUNDING; state each judgment rule once in plain prose beside the axiom it governs.

DimensionReasoningScore

Conciseness

The body spends ~340 of its ~390 lines on a Lean 4 formal block whose content — origin grounding rules, a record structure, a recursive dialogue loop — could be stated in a small fraction of the tokens; each axiom's doc comment re-explains judgment rules at length (e.g., the multi-paragraph StandingSupported comment). This matches 'noticeably verbose; several unnecessary explanations or padded sections' rather than a 3, because the padding is structural: the entire formal encoding is an expensive restatement of prose rules, not a few local over-explanations.

2 / 5

Actionability

There is real concrete guidance — TOOL GROUNDING's `surface` entry specifies what to present (one-sentence read-back, per-coordinate sources, contrary grounds before answer slots), the Intensity table maps situations to formats, and round-composition.md carries concrete placement rules. But it is not a 4: no worked example of an actual surface or dialogue round exists anywhere, and the Protocol section delegates execution by pointer ("Present what TOOL GROUNDING's `surface` entry names"), leaving the model to reconstruct behavior from a dense formal block. It is above a 2 because the surface entry's composition rules are specific enough to act on.

3 / 5

Workflow Clarity

The FLOW comment lays out a real sequence — surface, fuse each utterance, check withdrawal/resolution gates, then converge or continue — and the Morphism section orders the phases. It is not a 4 because the sequence lives inside a Lean code block that the reader must decode, the Phase Transitions section defines steps by cross-reference, and there are no explicit checkpoints (e.g., what to do when a surface is misunderstood or a run stalls); it is above a 2 because the gates and exit conditions are enumerated unambiguously.

3 / 5

Progressive Disclosure

The single reference, references/round-composition.md, is genuinely well signaled — the body names it with explicit load conditions ("Read ... before composing when a term must remain stable...") and it is one level deep. But the structure as a whole matches 'content that should be separate is inline': the ~340-line formal contract is a monolithic inlined block that would more naturally live in a reference file, and only one reference exists to absorb the occasion-bound material, so organization is only partly realized.

3 / 5

Total

11

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description answers both what and when with a distinctive niche and mostly concrete actions, in appropriately terse third-person voice. Its weaknesses are the thin set of natural trigger phrases and the high-level wording of its actions, which limit a user's ability to recognize the skill by their own words alone.

Suggestions

Broaden natural trigger coverage with phrases users actually say, e.g. "Use when the user says they 'know roughly what they want' but can't pin down the requirements, decisions, or open questions behind it."

Make the actions more operational — e.g., name what 'trace' involves (reading code, rules, past sessions to surface the open decisions) and what a round looks like (surface each decision with its source, ask, repeat until settled).

Consider adding a couple of synonyms (requirements, open questions, decision points) so the skill triggers on varied user phrasings.

DimensionReasoningScore

Specificity

The description names two concrete actions — "trace them from their own material (code, rules, past sessions) and the domain's usual decisions" and "ask until it settles" — with the material sources enumerated, matching the anchor for naming the domain plus 1-2 concrete actions without comprehensive coverage. It is not a 4 because the actions are given at a high level ("trace", "ask") without the variety of specific operations that would fill out minor gaps.

3 / 5

Completeness

Both what and when are present: the what is "trace them from their own material... and the domain's usual decisions; ask until it settles", and the when is the opening condition "The user knows roughly what they want but not which decisions it turns on". This matches the anchor where both are present but the 'when' could be more explicit — it is an embedded condition rather than a "Use when..." clause, which keeps it below 5.

4 / 5

Trigger Term Quality

Phrases like "knows roughly what they want", "code, rules, past sessions", and "ask until it settles" are words a user might plausibly use, giving some relevant keywords. It is not a 4 because common natural variations and synonyms (e.g., "requirements", "clarify", "decide", "figure out what I want") are absent, and it is not a 2 because the keywords present are domain-relevant rather than purely generic.

3 / 5

Distinctiveness Conflict Risk

The niche — eliciting the decision coordinates behind an underspecified intent from the user's own material — is fairly distinct and unlikely to fire for unrelated skills. It is not a 5 because the description could overlap with general clarification or requirements-gathering skills whose triggers ("knows roughly what they want") resemble each other, and not a 3 because the sourcing from the user's own material and the domain's decisions carves out a recognizable niche.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
jongwony/epistemic-protocols
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.