CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-production-validator

Agent skill for production-validator - invoke with $agent-production-validator

49

1.22x
Quality

21%

Does it follow best practices?

Impact

98%

1.22x

Average score across 3 eval scenarios

SecuritybySnyk

Medium

Suggest reviewing before use

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/agent-production-validator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

25%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body has a coherent theme — verifying no mock/stub implementations remain and testing against real infrastructure — and a useful shell-grep checklist, but it is delivered as a monolithic dump of corrupted, non-executable test-suite examples with no ordered workflow, no feedback loops, and no external reference structure. The broken syntax ('$' for '/' throughout the code) makes most of it unrunnable as written, and it needs both a repair pass and a split into reference files.

Suggestions

Fix the corrupted code syntax — every regex literal and URL uses '$' where '/' belongs ($mock[A-Z]\w+$g, https:/$api.stripe.com$v1) — and either define or parameterize the placeholder classes (UserRepository, PaymentService, APIClient) so the examples are actually executable.

Restructure around an ordered validation workflow (scan for mocks -> verify environment -> run integration tests -> load test -> report) with explicit checkpoints and fix-and-retry loops, instead of presenting five unrelated example test suites.

Split the test-suite listings into separate files under references/ (e.g. references/integration-tests.md, references/performance-tests.md) and keep SKILL.md as a concise overview with the runnable grep checklist, cutting the redundant 'Best Practices' prose.

DimensionReasoningScore

Conciseness

The ~370-line body pads heavily: five full test-suite listings, a redundant 'Best Practices' section whose bullets restate what the code already demonstrates ('Test against actual databases, not in-memory alternatives'), and a duplicated agent-config YAML block that is dead weight. It is above 1 because the material is domain-specific rather than explaining concepts Claude already knows, but several sections could be cut or condensed to a fraction of their length.

2 / 5

Actionability

Concrete-looking code dominates the file, but it is not executable: regex literals use '$' delimiters instead of slashes ($mock[A-Z]\w+$g), URLs are corrupted ('https:/$api.stripe.com$v1'), and every class (UserRepository, PaymentService, RedisCache, APIClient, EmailService) is undefined. This lands at 'minimal concrete guidance / high-level hints' in practice — the shell checklist section is the only genuinely runnable part. It stays above 1 because the test-suite shapes do communicate a recognizable pattern.

2 / 5

Workflow Clarity

The document is a catalog of example test suites, not a sequenced process: there is no ordered set of steps telling Claude how to run a production validation engagement, and no validation checkpoints or fix-and-retry feedback loops despite covering batch grep scans and load testing (which the guidelines cap at 3 regardless). A rough implicit order (scan code, check environment, run integrations, load test) exists, placing it at 2 rather than 1.

2 / 5

Progressive Disclosure

The file is monolithic: no references/, scripts/, or assets/ directories exist, and ~370 lines of test-suite examples and checklists that clearly belong in separate reference files are all inlined. There is some heading structure ('Validation Strategies', 'Validation Checklist', 'Best Practices') which keeps it above the wall-of-text anchor 1, but it squarely matches 'content that clearly belongs in separate files is inlined'.

2 / 5

Total

8

/

20

Passed

Description

17%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

This is a mechanical wrapper description generated around a skill name rather than a description of the skill's capabilities. It answers neither what the skill concretely does nor when to invoke it, and contains no natural trigger terms. It should be rewritten to state the skill's actual function (e.g. 'Verifies applications are fully implemented with no mock/stub code, tests against real databases and APIs, and validates deployment readiness. Use when preparing to deploy or when you need to verify production readiness.')

Suggestions

Replace the boilerplate with a concrete capability statement drawn from the body, e.g. 'Scans codebases for mock/stub/TODO implementations, runs E2E tests against real databases, APIs, and infrastructure, and validates deployment readiness.'

Add an explicit 'Use when...' trigger clause naming natural user phrases such as 'production readiness', 'pre-deploy validation', 'e2e testing', or 'remove mock implementations'.

Drop the 'invoke with $agent-production-validator' invocation syntax — it is harness metadata, not a capability or trigger, and crowds out useful signal within the character budget.

DimensionReasoningScore

Specificity

The description only names the domain ('Agent skill for production-validator') with no concrete actions whatsoever — it never says what the skill actually does (e.g. verify implementations, run E2E tests, scan for mocks). It fits the anchor 'Names the domain but actions are minimal or generic' rather than 1 only because the domain term 'production-validator' is present; it cannot reach 3 since no specific capability is stated.

2 / 5

Completeness

It has a vague 'what' ('Agent skill for production-validator') and no 'when' clause at all — exactly matching the anchor 'Has a vague what and no when'. It does not reach 3 because even the 'what' is a label rather than a concrete capability statement, and the guideline capping completeness at 3 for a missing 'Use when...' clause applies a fortiori here.

2 / 5

Trigger Term Quality

The only phrase beyond the boilerplate is 'invoke with $agent-production-validator', which is internal invocation syntax, not a natural keyword. No user would say these words when they need this skill; there are no natural terms like 'production readiness', 'e2e testing', or 'mock implementations' that a user would actually utter.

1 / 5

Distinctiveness Conflict Risk

The text is generic agent-skill wrapper boilerplate ('Agent skill for X - invoke with $X') that would be nearly identical for any generated agent, creating high overlap risk with every similarly-formatted skill. It is above 1 only because the 'production-validator' token gives one distinguishing marker, but it lacks the concrete trigger phrases needed for a clear niche.

2 / 5

Total

7

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ruvnet/ruflo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.