CtrlK
BlogDocsLog inGet started
Tessl Logo

legreffier-eval

Evaluate rendered packs against scenarios and author new gap-test scenarios with adversarial baseline gating. Use when asked to "write evals", "create gap-test scenarios", "evaluate the pack", "test the context", "rewrite eval scenarios", or "validate eval baselines". Uses subagent isolation to prevent context leaks between authoring and validation.

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

91%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an actionable, well-sequenced workflow with strong validation gates and clean one-level-deep reference structure. Its only weakness is mild redundancy where the isolation rationale is conveyed in both prose and diagram form.

Suggestions

Collapse the ASCII orchestrator box diagram and the prose 'Why subagent isolation matters' section into a single concise treatment to reduce redundancy (conciseness).

The 'Information flow diagram' checklist largely restates the contracts already specified in the loop steps — consider trimming it to only the non-obvious ✓/✗ contrasts (conciseness).

The 'Phases (from #566)' forward-looking section references ticket numbers and future work that may drift; relocate to a reference file or trim to keep the body evergreen (conciseness / time-sensitivity).

DimensionReasoningScore

Conciseness

Mostly lean and assumes Claude's competence (commands, gate tables, contracts), but the ASCII orchestrator diagram and information-flow grid restate the isolation rationale already explained in prose, which could be trimmed.

4 / 5

Actionability

Fully executable guidance throughout — concrete eval commands with flags, a delta-report table, a gate-check table with thresholds, and pointed references to real bundle files (scenario-format.md, author-subagent-prompt.md).

5 / 5

Workflow Clarity

Both modes are explicitly sequenced with numbered steps, a measurement gate (run ≥2 times, baseline thresholds), and an explicit iterate-or-accept feedback loop keyed on the gate result.

5 / 5

Progressive Disclosure

Well-signaled one-level-deep references to real bundle files (scenario-format.md, gap-test-principles.md, fixture-ref-discovery.md, author-subagent-prompt.md) keep the body an overview while detail lives in references/; no nested chains.

5 / 5

Total

19

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is third-person, concise, and clearly answers both what the skill does and when to use it, with comprehensive natural trigger phrases. It is specific and distinct with minimal conflict risk.

DimensionReasoningScore

Specificity

Names the domain and several concrete actions ('Evaluate rendered packs against scenarios', 'author new gap-test scenarios', 'adversarial baseline gating', subagent isolation), with only minor gaps in coverage of the run-mode reporting.

4 / 5

Completeness

Explicitly answers both what (evaluate packs / author gap-test scenarios with adversarial baseline gating) and when (clear 'Use when asked to ...' clause with concrete trigger phrases).

5 / 5

Trigger Term Quality

Comprehensive coverage of natural trigger phrases users would say — 'write evals', 'create gap-test scenarios', 'evaluate the pack', 'test the context', 'rewrite eval scenarios', 'validate eval baselines' — including synonyms.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (eval scenario authoring with subagent-isolated baseline gating) with distinct triggers unlikely to fire for unrelated skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
getlarge/themoltnet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.