CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-tool-builder

Tools are how AI agents interact with the world. A well-designed tool is the difference between an agent that works and one that hallucinates, fails silently, or costs 10x more tokens than necessary. This skill covers tool design from schema to error handling.

46

Quality

48%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/antigravity-agent-tool-builder/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

47%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean overview with a well-signaled, real one-level-deep reference — good progressive disclosure. But it offers no workflow of its own, and its single concrete artifact (the Python example) is non-executable fabricated-API code wrapped in quotes rather than a fenced block, leaving the body more descriptive than instructional.

Suggestions

Fix or remove the Python example: fence it as ```python, drop the non-existent 'beta_tool'/'tool_runner' API in favor of the real tool-use API, and add the missing 'import json' — or move it into the detailed guide.

Add a minimal design workflow to the body (e.g. 1. draft the schema, 2. write tool/parameter descriptions, 3. add enum constraints and error returns, 4. validate against the checklist in the guide) so the skill instructs even before the reference is loaded.

De-duplicate the intro (it repeats the description verbatim) and collapse the 12 'User mentions or implies:' bullets into a single trigger line.

DimensionReasoningScore

Conciseness

The intro duplicates the frontmatter description verbatim ('Tools are how AI agents interact with the world. A well-designed tool is the difference between...') and 'When to Use' pads 12 near-identical one-line bullets ('User mentions or implies: X') that could be a single line. The 'Key insight' paragraph does earn its place, so this is 'mostly efficient but includes some unnecessary explanation' rather than score 2's 'several padded sections'.

3 / 5

Actionability

The 'Python Example' looks concrete but is not executable: it is wrapped in triple quotes instead of a fenced code block, uses a non-existent API surface ('from anthropic import beta_tool', 'client.beta.messages.tool_runner'), and omits 'import json'. The body otherwise delegates all guidance to the reference file, matching 'some concrete guidance but incomplete; pseudocode instead of executable code; missing key details'.

3 / 5

Workflow Clarity

No steps are enumerated in the body — the only sequence is the implied 'read the detailed guide, then apply it', with no checkpoints or validation guidance for the design work itself. This fits 'rough sequence present but many gaps; steps poorly defined'; it cannot score 3 because no steps are actually listed, and the destructive/batch cap does not apply since this is not such a skill.

2 / 5

Progressive Disclosure

There is exactly one reference (references/detailed-guide.md, verified to exist, 650 lines, well-sectioned), clearly signaled ('Read the detailed guide before executing this skill... It retains the complete procedure') and only one level deep. Minor gaps keep it from 5: the 40-line inlined Python example arguably belongs in the guide, and 'load the relevant sections' does not enumerate which sections the guide contains.

4 / 5

Total

12

/

20

Passed

Description

50%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a genuine niche (agent tool design) and hints at coverage from schema to error handling, but two-thirds of it is motivational padding rather than capability statements. Its biggest structural flaw is the complete absence of any 'Use when...' trigger guidance, which caps completeness and leaves it indistinguishable from neighboring tooling skills.

Suggestions

Replace the two motivational sentences with concrete capability statements, e.g. 'Design agent tool schemas, write tool and parameter descriptions, and handle tool errors. Covers JSON Schema best practices and the MCP standard.'

Add an explicit trigger clause: 'Use when the user mentions agent tools, function calling, tool schemas, tool use, or building an MCP server/tool.'

Trim the '10x more tokens' over-claim; state what the skill does, not why tool design matters.

DimensionReasoningScore

Specificity

The only capability statement is 'This skill covers tool design from schema to error handling', which names the domain plus two concrete facets; the two preceding sentences ('Tools are how AI agents interact with the world. A well-designed tool is the difference between an agent that works and one that hallucinates...') are motivational fluff with no capabilities. It clears score 2 because schema and error handling are named concrete topics, but falls short of score 4's 'several specific actions'.

3 / 5

Completeness

The 'what' is stated ('covers tool design from schema to error handling') but the 'when' is entirely missing — there is no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. It cannot score 4 because the 'when' is not even weakly implied.

3 / 5

Trigger Term Quality

Relevant keywords include 'tools', 'AI agents', 'tool design', 'schema', and 'error handling', but common natural variations users would actually say — 'function calling', 'tool use', 'MCP', 'build a tool' — are absent. This matches 'some relevant keywords but missing common variations or synonyms' rather than score 4's 'good keyword coverage'.

3 / 5

Distinctiveness Conflict Risk

'Tool design' is a real niche, but the generic opening ('Tools are how AI agents interact with the world') and the absence of distinguishing trigger phrases mean it could overlap with MCP-server, API-integration, or general function-calling skills. This fits 'somewhat specific but could still overlap with similar skills' rather than score 4's 'mostly distinct'.

3 / 5

Total

12

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
boisenoise/skills-collections
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.