CtrlK
BlogDocsLog inGet started
Tessl Logo

tests-run

Execute Unity tests (`EditMode` or `PlayMode`) and return per-test results. Supports filtering by test assembly, namespace, class, and method. Refreshes the AssetDatabase first; defers execution across domain reloads if scripts changed. Precondition: every open scene must be saved — dirty scenes abort the run.

60

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./Unity-MCP-Plugin/.claude/skills/tests-run/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, mostly actionable tool reference with good error-recovery documentation, but it is held back by triple-documented parameters and two large raw JSON schemas inlined in the main file. Moving the schemas to reference files and consolidating the parameter documentation would materially improve both conciseness and progressive disclosure.

Suggestions

Move the Input and Output JSON schemas to a references/ file (e.g. references/schemas.md) and keep only a short summary of the response shape inline; the full .NET-named $defs dump is the single largest source of bloat.

Consolidate the Filters/Response toggles prose with the Input table — one canonical parameter listing instead of three overlapping enumerations.

Replace the "string_value" placeholder example with one fully-real invocation (e.g. a PlayMode run filtered to one test class) so the example is directly copy-paste executable.

DimensionReasoningScore

Conciseness

The body documents the same 11 parameters three times — prose in "Filters"/"Response toggles", the Input table, and the Input JSON Schema — and then dumps a ~140-line raw output JSON schema with .NET type names like "System.Collections.Generic.List(com.IvanMurzak.Unity.MCP.Editor.API.TestRunner.TestLogEntry)". Most content is information-bearing (not concept re-explanation), so it clears score 2, but the redundancy and inline schema bloat mean it could be tightened considerably.

3 / 5

Actionability

"unity-mcp-cli run-tool tests-run --input ..." plus the --input-file and stdin heredoc forms give a mostly executable, copy-paste-ready invocation, and the Input table supplies concrete example values. The gap keeping it from 5: the primary example uses placeholders ("string_value", bare booleans) rather than one fully-real call, e.g. a real EditMode run filtered to a test class.

4 / 5

Workflow Clarity

For a single-tool skill the action is unambiguous, and failure/recovery paths are documented: "If any open scene is dirty... save them and retry", "Pre-existing compilation errors short-circuit the run... so the caller can fix the project first", the domain-reload Processing/resume flow, and CLI-not-found fallbacks (npm install -g / npx). Minor gaps keep it at 4: these checkpoints are scattered across sections rather than sequenced, and there is no explicit verify step for confirming tests actually ran post-reload.

4 / 5

Progressive Disclosure

Section structure is clear (Filters, Response toggles, Domain reloads, How to Call, Input, Output) and the /unity-initial-setup reference is one level deep, but ~160 lines of raw input/output JSON schemas are inlined in SKILL.md — content that clearly belongs in reference files given no bundle directory exists. This matches the anchor "some structure but... content that should be separate is inline" rather than 2, since structure and navigation are otherwise good.

3 / 5

Total

14

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description with concrete, comprehensive capability statements and excellent niche distinctiveness. Its one structural weakness is the missing "Use when..." trigger clause, which caps completeness and leaves discovery to keyword overlap alone.

Suggestions

Append a trigger clause such as: "Use when the user asks to run Unity tests, a test suite, or specific test classes/methods, or wants per-test pass/fail results."

Include common user phrasings like "run tests", "test suite", or "unit tests" alongside the Unity-specific terms to broaden natural trigger coverage.

DimensionReasoningScore

Specificity

"Execute Unity tests (`EditMode` or `PlayMode`) and return per-test results. Supports filtering by test assembly, namespace, class, and method. Refreshes the AssetDatabase first; defers execution across domain reloads if scripts changed" — multiple specific concrete actions with comprehensive coverage of the tool's behavior, including boundary conditions (dirty scenes abort the run). No filler or vague language.

5 / 5

Completeness

The "what" is clearly and thoroughly answered, but there is no "Use when..." clause or equivalent explicit trigger guidance — the rubric explicitly caps completeness at 3 in that case. Not score 2 because the "what" is concrete, not vague; not score 4+ because the "when" is entirely absent rather than merely implicit.

3 / 5

Trigger Term Quality

Good natural keywords: "Unity tests", "EditMode", "PlayMode", "test assembly", "namespace", "class", "method" — phrases a user working in Unity would naturally say. A few common variations are missing (e.g., "run tests", "test suite", "unit tests"), so it falls just short of the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

"Execute Unity tests (`EditMode` or `PlayMode`)" with Unity-specific triggers (EditMode, PlayMode, AssetDatabase, domain reloads) carves out a clear niche with minimal overlap risk against generic test- or document-related skills. Re-checked against the 4 anchor: this is more distinct than "minor overlap risk" — the Unity-specific terminology makes mis-triggering unlikely.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
IvanMurzak/Unity-MCP
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.