CtrlK
BlogDocsLog inGet started
Tessl Logo

test-reasoning

Validate that reasoning parameters are correctly serialized and sent to provider APIs. Use when the user asks to test reasoning serialization, run reasoning tests, verify reasoning config fields, or check that ReasoningConfig maps correctly to provider-specific JSON (OpenRouter, Anthropic, GitHub Copilot, Codex).

76

Quality

95%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary SKILL.md body: lean, fully executable, and correctly structured as an overview over a bundled script. The dense provider/expectation table provides all specifics needed without any padding, and the manual path carries an explicit verification step.

DimensionReasoningScore

Conciseness

Every section earns its place: a one-line quick-start command, a single manual-test snippet, a dense expectation table, and four reference links. No concepts Claude already knows are explained, and there is no padding. Matches the lean top anchor.

5 / 5

Actionability

Commands are copy-paste ready ("./scripts/test-reasoning.sh" and a complete env-var invocation with FORGE_DEBUG_REQUESTS), and the coverage table gives exact config fields and expected JSON fields per provider/model. Fully executable with specific examples covering the common cases.

5 / 5

Workflow Clarity

The primary path is a single unambiguous action (run the bundled script), and the manual path includes an explicit validation checkpoint ("Then inspect .forge/forge.request.json for the expected fields") plus a documented failure mode ("non-zero exit, no request written"). Under the simple-skill exception the single action is unambiguous with no sequence gaps.

5 / 5

Progressive Disclosure

The ~500-line test logic is properly offloaded to the real bundle file scripts/test-reasoning.sh (verified present), which the body references by exact path. The body is a well-organized overview with clear sections, one-level-deep references, and easy navigation. Matches the top anchor for a bundle-backed skill.

5 / 5

Total

20

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description in third-person voice with an explicit Use-when clause listing concrete trigger phrases and a well-scoped provider set. The only weakness is that its trigger terms and actions are variations of a single phrase rather than a diverse set of synonyms.

Suggestions

Broaden trigger-term coverage with natural synonyms a user might say, e.g. "check that reasoning effort/thinking settings are applied" or "debug reasoning config not being sent".

Mention the expected outcome artifacts (e.g. verifying the captured request JSON) to make the capability set feel more comprehensive rather than one action restated.

DimensionReasoningScore

Specificity

"Validate that reasoning parameters are correctly serialized and sent to provider APIs" plus "check that ReasoningConfig maps correctly to provider-specific JSON" names several concrete actions with an enumerated provider list. Not 5 because the actions are variations of a single validation task rather than comprehensive, distinct capabilities.

4 / 5

Completeness

Clearly answers what ("Validate that reasoning parameters are correctly serialized and sent to provider APIs") and when via an explicit "Use when the user asks to..." clause with concrete trigger phrases. Both elements are explicit, matching the top anchor.

5 / 5

Trigger Term Quality

"test reasoning serialization", "run reasoning tests", "verify reasoning config fields" are natural user phrasings with good coverage. Not 5 because the terms are repetitive variations of one phrase and miss synonyms or alternate wordings a user might naturally say.

4 / 5

Distinctiveness Conflict Risk

The project-specific "ReasoningConfig" term and named providers (OpenRouter, Anthropic, GitHub Copilot, Codex) carve out a clear niche with distinct triggers. Conflict risk with other skills is minimal.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tailcallhq/forgecode
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.