CtrlK
BlogDocsLog inGet started
Tessl Logo

semantic-consistency-auditor

Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.

39

Quality

49%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Academic Writing/semantic-consistency-auditor/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body has genuinely executable CLI guidance that matches the packaged script, but it is buried in a padded, repetitive document with circular cross-references and duplicated sections. Consolidating to one coherent workflow and moving config/format details into the existing reference file would substantially improve it.

Suggestions

Remove the boilerplate and duplication: the circular 'See ## X above' pointers, the repeated py_compile commands (Quick Check vs Audit-Ready Commands vs Example Usage), the two References sections, and the generic Output Requirements/Response Template filler — this could cut the body roughly in half.

Consolidate Example run plan, Workflow, Implementation Details, and Error Handling into a single sequenced workflow with explicit validation checkpoints (e.g. verify results.json is written and parses, check pass counts against thresholds before reporting).

Fix inaccurate examples: the `from semantic_consistency_auditor import ...` Python API snippet (the class lives in scripts/main.py), the hardcoded `cd "20260318/..."` path, and the `~/.openclaw/skills/...` config path versus the script's actual `--config` flag — and clearly signal references/audit-reference.md once, near the top.

DimensionReasoningScore

Conciseness

The ~360-line body contains substantial padding: circular self-references ("See `## Prerequisites` above", "See `## Usage` above", "See `## Workflow` above" — pointing to sections that appear later), duplicated dependency listings, the same `py_compile` command repeated in three sections, two separate References sections, and generic filler (Output Requirements, Response Template, Input Validation) that adds no skill-specific value. This matches 'noticeably verbose; several unnecessary explanations or padded sections' rather than 1, since the algorithm/config/usage sections do carry real information.

2 / 5

Actionability

The CLI examples (`python scripts/main.py --ai-generated ... --gold-standard ... --output results.json`, batch `--input-file`, model overrides) match the actual argparse definitions in scripts/main.py and are executable, as are the config YAML and documented input/output JSON formats. Minor gaps keep it below 5: the Python API example imports `semantic_consistency_auditor` as a package though only `scripts/main.py` exists, and the Example Usage hardcodes a nonexistent path (`cd "20260318/scientific-skills/..."`).

4 / 5

Workflow Clarity

Steps exist (Example run plan, Workflow, Error Handling) but are scattered across three overlapping sections with circular cross-references, and the batch-evaluation workflow lacks result-level validation checkpoints. Per the rubric, a batch operation without validation is capped at 3, which is also where 'steps listed but validation gaps' fits; a 4 would require one coherent sequence with most checkpoints explicit.

3 / 5

Progressive Disclosure

The body has section structure and one one-level-deep reference (references/audit-reference.md, which exists), but that reference is only linked at the end in a second 'References' section and duplicates SKILL.md boilerplate, while content that belongs in separate files (config docs, input/output JSON schemas, Python API details) is inlined in a 360-line monolithic body. This fits 'references present but not clearly signaled; content that should be separate is inline'.

3 / 5

Total

12

/

20

Passed

Description

25%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is generic boilerplate that fails to convey the skill's real purpose (BERTScore/COMET-based semantic consistency evaluation of AI-generated clinical notes against expert gold standards). It would rarely trigger for the right task and frequently risks mis-triggering for unrelated structured-writing tasks.

Suggestions

State the concrete capability first, e.g. "Evaluates semantic consistency between AI-generated clinical notes and expert gold standards using BERTScore and COMET, reporting precision/recall/F1 plus a composite consistency score."

Add explicit trigger phrases users would actually say, e.g. "Use when comparing AI-generated medical or clinical text against reference/gold-standard text, running BERTScore or COMET evaluation, or auditing semantic entailment."

Drop the generic process language ("structured execution, explicit assumptions, clear output boundaries") — it does not distinguish this skill from any other structured workflow skill.

DimensionReasoningScore

Specificity

The description names a domain ("academic writing workflows") but lists no concrete actions — "structured execution, explicit assumptions, and clear output boundaries" are abstract process properties, not capabilities. It never mentions the skill's actual functions (BERTScore/COMET evaluation of clinical notes against gold standards), matching the 'names the domain but actions are minimal or generic' anchor; a score of 1 would require no domain naming at all, and 3 would require 1-2 concrete actions.

2 / 5

Completeness

The 'what' is vague (no mention of what the auditor actually does) and the 'when' is a weak domain qualifier ("for academic writing workflows that need structured execution") with no explicit 'Use when...' trigger clause, which caps completeness at 3 anyway. It sits at the 'vague what and no real when' anchor rather than 3, which requires a clear 'what'.

2 / 5

Trigger Term Quality

Only "semantic consistency auditor" and "academic writing" appear as keywords, and neither is a phrase a user would naturally say when needing this skill. Natural triggers like "BERTScore", "COMET", "clinical notes", "evaluate against gold standard" are absent, fitting the 'one or two generic keywords; missing the natural phrases users say' anchor.

2 / 5

Distinctiveness Conflict Risk

"Structured execution, explicit assumptions, and clear output boundaries" could describe almost any structured analysis or writing skill, creating high overlap risk with many similar skills. Only the embedded skill name provides any distinctiveness, so it fits 'very broad; high overlap risk' rather than 3's 'somewhat specific'.

2 / 5

Total

8

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.