CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/qa-vendor-evaluator

Build-an-X workflow that produces a side-by-side **commercial-vendor** evaluation matrix for QA tools - test-management platforms (TestRail / Qase / Xray / Zephyr / TestCollab), no-code platforms (mabl / Testim / Functionize / TestSigma / Reflect), visual regression services (Applitools / Percy / Chromatic), and commercial AI copilots - scoring each on capability fit, cost model, integration depth, vendor lock-in risk, exit cost, contractual posture, and customer-reference data. Scoped to commercial procurement - contract, lock-in, and exit-cost axes - not to choosing an open-source code-first framework on architectural fit. Use for commercial procurement decisions only - refuses to recommend a winner; the team owns the procurement choice.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable procurement procedure with explicit input gates, a concrete scoring scale, a specified output template, and clean one-level-deep reference delegation. Its only real weakness is conciseness: the Capgemini statistic and the 'not a recommendation' framing are each repeated several times.

Suggestions

Cite the Capgemini WQR 37% integration-friction statistic once (Step 2 A3, where it is load-bearing) and reference it elsewhere by name only instead of restating the figure.

Consolidate the 'decision-support, not recommendation / refuses to pick a winner' framing into one place (e.g., the Overview) and have Step 3, Step 4, and the Anti-patterns row point back to it rather than restating it.

Tighten the Overview to remove the overlap with the frontmatter description - the body's Overview can assume the description and jump straight to the WQR motivation and the decision-support stance.

DimensionReasoningScore

Conciseness

Mostly efficient and information-dense with no concept-explanation Claude already knows, but the Capgemini 'integration friction (37%)' stat is cited three times (Overview, Step 2 A3, Anti-patterns) and the 'refuses to pick a winner / not a recommendation' framing recurs across Overview, Step 3, Step 4, and an Anti-patterns row. Not score 3 because that repetition means not every token earns its place; not score 1 because it is not padded with basic concept explanation.

2 / 3

Actionability

Concrete, executable procedure for an instruction-only skill: the A3 integration scale 'native (1.0) / API-buildable (0.7) / community-plugin (0.5) / not available (0.0)', the 'Halts on 1' / recommends >=3 candidate gate, the fully-specified Step 3 output document structure, and the 5-step Step 4 hand-off. Not score 2 because the guidance is specific and complete rather than pseudocode or incomplete.

3 / 3

Workflow Clarity

Clear four-step sequence (capture inputs, score seven axes, emit matrix, hand off) with explicit checkpoints: the 'halts with INSUFFICIENT_INPUT' input gate, the 'every score citing its source' verification requirement, and the Anti-patterns table as a failure-mode checklist. Not score 2 because checkpoints are explicit rather than implicit; no fix/retry loop is required since this is a non-destructive analysis skill.

3 / 3

Progressive Disclosure

SKILL.md is an overview/procedure that delegates the detailed per-sub-axis rubric to references/scoring-rubric.md and the full worked matrix to references/example-matrix.md - both real, one-level-deep files clearly signaled via markdown links in Step 2 and Step 3. Not score 2 because references are clearly signaled and appropriately split rather than inlined or nested.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and distinctive: it names seven concrete scoring axes, an explicit trigger clause, and explicit de-confliction against the open-source framework-choice sibling skill. It is third-person throughout and free of vague fluff, though it is a long, dense single sentence.

DimensionReasoningScore

Specificity

Lists multiple concrete actions - 'produces a side-by-side evaluation matrix' and 'scoring each on capability fit, cost model, integration depth, vendor lock-in risk, exit cost, contractual posture, and customer-reference data' - naming seven specific scoring axes and the output artifact. Not score 2 because the action set is comprehensive rather than partial.

3 / 3

Completeness

Clearly answers both what (produces the seven-axis evaluation matrix) and when via the explicit 'Use for commercial procurement decisions only' trigger clause. Not score 2 because an explicit trigger clause is present, so the missing-trigger cap does not apply.

3 / 3

Trigger Term Quality

Good coverage of natural terms a QA manager would say - 'QA tools', 'test-management platforms', 'no-code platforms', 'visual regression', 'vendor', 'procurement' - alongside the brand enumeration. Not score 2 because common natural variations are present, not just jargon.

3 / 3

Distinctiveness Conflict Risk

Clear niche with explicit de-confliction - 'Scoped to commercial procurement' and 'not to choosing an open-source code-first framework on architectural fit' - plus 'refuses to recommend a winner'. Not score 2 because it actively scopes against the sibling skill, making wrong-skill triggering unlikely.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents