CtrlK
BlogDocsLog inGet started
Tessl Logo

memory-safety

Run AddressSanitizer and UndefinedBehaviorSanitizer on the Z3 test suite to detect memory errors, undefined behavior, and leaks. Logs each finding to z3agent.db.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong example of a script-backed skill: fully executable commands, a well-gated three-step workflow with explicit error-recovery branches, and clean separation between overview and implementation. Only trivial trimming of the opening paragraph is possible.

DimensionReasoningScore

Conciseness

The body is efficient: the Action/Expectation/Result scaffold, sanitizer-effect enumerations, and parameters table each earn their tokens, with only minor verbosity (the opening paragraph partly restates the frontmatter description).

4 / 5

Actionability

Every code block is executable and copy-paste ready ('--sanitizer asan', '--skip-build --build-dir build/sanitizer-asan', concrete z3db.py run/query commands), and the parameters table documents all flags with types and defaults.

5 / 5

Workflow Clarity

Three clearly sequenced steps with explicit checkpoints — build success gates Step 2, and each step enumerates outcome branches (clean/findings/timeout/error) with recovery actions like 'inspect build logs before re-running' and 'increase the deadline or investigate a possible infinite loop'.

5 / 5

Progressive Disclosure

Well-organized sectioned overview that delegates implementation to the real bundle script (scripts/memory_safety.py, verified present), with only one-level-deep references and a parameters table — nothing is inlined that belongs in a separate file.

5 / 5

Total

19

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and highly distinctive, clearly stating what the skill does and where results go. Its main weakness is the complete absence of 'when to use' trigger guidance, and it lacks the abbreviations (ASan/UBSan) users naturally say.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user mentions sanitizers, ASan/UBSan, memory leaks, undefined behavior, or wants memory-safety testing of Z3.'

Include natural synonyms and abbreviations ('ASan', 'UBSan', 'sanitizer run', 'memory-safety check') so the description matches how users actually phrase the request.

Optionally mention deduplication and longitudinal tracking in the description, since that differentiates it from a one-off sanitizer run.

DimensionReasoningScore

Specificity

Names several concrete actions — 'Run AddressSanitizer and UndefinedBehaviorSanitizer on the Z3 test suite', 'detect memory errors, undefined behavior, and leaks', 'Logs each finding to z3agent.db' — but omits parts of the workflow (build configuration, dedup/triage), so it falls just short of comprehensive coverage.

4 / 5

Completeness

The 'what' is clear and specific, but there is no 'Use when...' clause or equivalent explicit trigger guidance — the rubric explicitly caps completeness at 3 in that case.

3 / 5

Trigger Term Quality

Good natural-keyword coverage ('AddressSanitizer', 'UndefinedBehaviorSanitizer', 'memory errors', 'leaks'), but misses common abbreviations and synonyms users would say such as 'ASan', 'UBSan', or standalone 'sanitizer'.

4 / 5

Distinctiveness Conflict Risk

Tightly scoped to sanitizer runs on the Z3 test suite with logging to a specific database — a clear niche with distinct triggers and minimal conflict risk with other skills.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Z3Prover/z3
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.