CtrlK
BlogDocsLog inGet started
Tessl Logo

alethia

Use when the user asks to run E2E tests, verify a web page, generate tests for an app, prove destructive actions are blocked, check if a UI element is visible, fill out a form, or drive a browser with natural language. Returns per-step results with safety classifications, policy decisions, DOM diffs, structured page context, and a signed audit trail.

95

2.80x
Quality

92%

Does it follow best practices?

Impact

98%

2.80x

Average score across 5 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is vitron-ai/alethia

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body with concrete NLP syntax, named tool calls, and validation-backed workflows including error-recovery loops. The main gaps are minor redundancy across response-field sections and external (non-bundle) references with some reference material inlined.

Suggestions

Consolidate the nearMatches/suggestedFix/pageContext description into one section instead of restating it across 'Workflow A', 'Structured Response Fields', and 'Understanding the PlanRun Response'.

Move the PlanRun field reference and reason-code table into a bundled reference file (e.g. references/planrun.md) and link to it, so SKILL.md stays an overview.

Either bundle the referenced ../../docs/index.md and ../../rules/alethia.md files or drop the external links to avoid dangling references.

DimensionReasoningScore

Conciseness

The body is mostly lean and tool-specific (NLP syntax, MCP tool names, reason codes), but nearMatches/suggestedFix/pageContext are restated across Workflow A, 'Structured Response Fields', and 'Understanding the PlanRun Response', and lines like 'the primitive that makes Alethia a verifiable-safety framework, not just a test runner' are mild padding — efficient with minor trimming possible.

4 / 5

Actionability

Copy-paste-ready NLP examples (navigate/assert/type/click/wait/expect block), explicit MCP tool names with their return shapes, and concrete argument examples (target URL, optional profile) fully cover the common cases.

5 / 5

Workflow Clarity

Workflows A/B/C are clearly sequenced with explicit validation checkpoints (alethia_status liveness, run.ok check) and a feedback loop (on failure, apply suggestedFix and retry), satisfying the error-recovery anchor for destructive/batch operations.

5 / 5

Progressive Disclosure

Well-organized into clear sections with one-level-deep, clearly signaled references ('For deeper references:'), but the referenced paths point outside the bundle (../../docs/index.md, ../../rules/alethia.md) and no bundle files exist, and some reference detail (PlanRun fields, reason codes) is inlined rather than split out.

4 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that concretely states both what the skill returns and when to invoke it, with natural trigger phrasing. The only soft spot is mild overlap risk with broader verification or browser-automation skills.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('run E2E tests, verify a web page, generate tests for an app, prove destructive actions are blocked, check if a UI element is visible, fill out a form, or drive a browser') plus a concrete output contract ('safety classifications, policy decisions, DOM diffs, structured page context, and a signed audit trail'), matching the comprehensive-coverage anchor.

5 / 5

Completeness

An explicit 'Use when the user asks to...' clause answers when, and 'Returns per-step results with safety classifications, policy decisions, DOM diffs...' answers what, both with concrete trigger phrases — matching the top anchor exactly.

5 / 5

Trigger Term Quality

Phrases like 'run E2E tests', 'verify a web page', 'check if a UI element is visible', 'fill out a form', and 'drive a browser with natural language' are natural things a user would actually say, giving comprehensive trigger coverage.

5 / 5

Distinctiveness Conflict Risk

The E2E/safety/audit-trail framing is a clear niche, but generic triggers like 'verify a web page', 'fill out a form', and 'drive a browser' carry minor overlap risk with general verify/run-browser skills, so it sits just below the minimal-conflict anchor.

4 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 suspicious

Warning

Total

15

/

16

Passed

Repository
vitron-ai/alethia-mcp
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.