CtrlK
BlogDocsLog inGet started
Tessl Logo

agentic-evaluator

Evaluates any repository's agentic development maturity. Use when auditing a codebase for best practices in agents, skills, instructions, MCP config, and prompts. Produces a scored report with specific remediation steps.

81

1.50x
Quality

73%

Does it follow best practices?

Impact

93%

1.50x

Average score across 3 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.github/skills/agentic-evaluator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable, with explicit scan paths, point tables, schemas, and a report template that make the evaluation workflow easy to execute end-to-end. Its main weakness is that it violates its own right-sizing and progressive-disclosure guidance: it is roughly double its recommended token budget, inlines large reference material, and cites supporting files that are absent from the bundle.

Suggestions

Actually create the referenced 'checklist.md' and 'report-template.md' bundle files (or remove the 'Supporting Files' section) — the current references dangle and the Phase 7 report template should live in report-template.md rather than inline.

Move the two full example reports, 'Remediation Patterns', 'Size Guidelines Reference', and 'Skill Quality Dimensions' into bundled reference files, keeping SKILL.md to the workflow and point tables within its own 80–300 line / 1–2K token budget.

Deduplicate the SkillsBench statistics and the 'Skill Development Best Practices' section (they repeat the Lean Context Principle material), and specify how to measure the token budgets used in the right-sized checks (e.g., a token-count command or lines-as-proxy rule).

DimensionReasoningScore

Conciseness

The scoring tables are dense and signal-heavy, but sections like 'Skill Development Best Practices', 'Lean Context Principle', and two full example reports restate guidance and repeat SkillsBench statistics across three separate places, pushing the body to ~460 lines / ~4K tokens — roughly double the skill's own 80–300 line, 1–2K token budget. Mostly efficient but could be significantly tightened; not 2 since nearly all content is task-relevant rather than concept-explanation Claude already knows.

3 / 5

Actionability

Gives concrete scan paths, per-check point tables, frontmatter schemas, a copy-paste report template, and invocation examples — mostly executable guidance. Not 5 because the right-sized checks depend on token budgets ('1–2K tokens') with no method given for measuring them, leaving a minor verifiability gap.

4 / 5

Workflow Clarity

Phases 1–7 (Discovery → Foundation → Skills → Agents → Instructions → Consistency → Report) are clearly sequenced with explicit per-check criteria and a structured output. Not 5 because validation checkpoints are implicit — no step verifies the measured counts/line totals before scoring, and the referenced validation checklist does not exist. Not 3 since the sequence and criteria are fully explicit and the workflow is read-only (no destructive-operation cap applies).

4 / 5

Progressive Disclosure

Section structure is good and references are signaled, but the entire skill is a single ~460-line file: the example reports, 'Remediation Patterns', 'Size Guidelines Reference', and 'Skill Quality Dimensions' are inlined content that — per the skill's own guidance — belongs in bundled files, and the 'Supporting Files' section points to 'checklist.md' and 'report-template.md' that do not exist in the bundle. Fits anchor 3: some structure, references present but unresolvable, separable content inline.

3 / 5

Total

14

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description in third-person voice with an explicit 'Use when' trigger clause, named target domains, and a concrete output promise (scored report with remediation steps). Both what and when are clearly answered; only minor refinements — more specific capability wording and additional trigger synonyms — would raise it further.

DimensionReasoningScore

Specificity

Lists several specific actions ('Evaluates any repository's agentic development maturity', 'Produces a scored report with specific remediation steps') and names five concrete target domains ('agents, skills, instructions, MCP config, and prompts'). Not 5 because 'maturity' is somewhat abstract and the description never states what the evaluation concretely covers (e.g., the scored categories).

4 / 5

Completeness

Explicitly answers both what ('Evaluates any repository's agentic development maturity... Produces a scored report with specific remediation steps') and when ('Use when auditing a codebase for best practices in...'), with concrete trigger phrases matching the anchor-5 example pattern.

5 / 5

Trigger Term Quality

'Use when auditing a codebase for best practices in agents, skills, instructions, MCP config, and prompts' provides good natural keyword coverage (auditing, best practices, agentic, codebase). Not 5 because common synonyms like 'review', 'assess', or 'audit our repo's AI setup' are missing.

4 / 5

Distinctiveness Conflict Risk

'Agentic development maturity' plus named artifact types (skills, MCP config, instructions) forms a clear niche distinct from generic code-review skills. Not 5 because 'auditing a codebase for best practices' could minorly co-trigger with general audit/lint/review skills.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
0xRabbidfly/Eric-Cartman
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.