CtrlK
BlogDocsLog inGet started
Tessl Logo

architecture-review

Traceability matrix mapping GDD requirements to ADRs. Finds gaps, cross-ADR conflicts, engine compatibility. PASS/CONCERNS/NOT ASSESSED/FAIL.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/architecture-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exceptionally actionable, well-sequenced review workflow: exact commands, templates, validation checkpoints, and fail-open handling of every degenerate input case. Its weaknesses are token efficiency (repeated rationale passages that could be stated once) and the total absence of progressive disclosure — everything, including large output-format templates, is inlined in a single very long file.

Suggestions

Move the stable output templates (RTM file format, reflexion-log entry, session-state block, traceability-index format) into a references/ file and link to them from the phases, cutting the main body substantially.

State the "skipped check is indistinguishable from a passed check" rule once and reference it from Phases 5-7 instead of re-explaining it three times; the same applies to the extended blockquote justifications in Phase 3 and the Error Recovery preamble.

Trim the meta-explanations of why rules exist (e.g. why 🟡 caps the verdict, why parentheticals break status matching) to a single sentence each — the rule itself is actionable without the narrative.

DimensionReasoningScore

Conciseness

The body is dense with project-specific, non-obvious semantics (denominator interpretation, fail-open rules, verdict gating), so it is not noticeably padded overall — but there is real trim material: the "a skipped check is indistinguishable from a check that passed" rationale is repeated in Phase 5, Phase 6, and Phase 7, and several long blockquotes (e.g. Phase 3's 🟡 justification, the Error Recovery preamble) explain reasoning the executing model could take on trust. Not 4 because the repetition and meta-justification are unnecessary tokens; not 2 because most content is load-bearing project knowledge rather than concepts Claude already knows.

3 / 5

Actionability

Guidance is copy-paste ready throughout: exact Grep patterns with glob and -A counts, exact Bash invocations (review-receipts.sh check/hash, adr-dep-graph.sh), complete output templates for the matrix, RTM file, conflict entries, and session-state block, plus scripted AskUserQuestion option lists. Specific examples cover the common cases of every mode.

5 / 5

Workflow Clarity

Nine clearly sequenced phases with explicit validation checkpoints everywhere: a freshness/receipt check before any scan, denominator counts with fail-open interpretation tables (including the 0-match malformed-ADR case), per-ADR escalation rules, artifact-on-disk verification before treating a phase as done, and an error-recovery protocol that requires a partial report. Feedback loops (re-scan, re-ask, re-validate) are explicit at every write.

5 / 5

Progressive Disclosure

No bundle files exist — the entire ~860-line skill lives in one inline SKILL.md. Internal structure is strong (phased headers, per-mode branching, tables), so it is not a 2, but content that clearly belongs in separate reference files (the RTM output format, reflexion-log entry format, session-state block, collaborative protocol) is inlined, and the only external pointers are to project docs rather than skill-managed references — so it does not reach 4.

3 / 5

Total

16

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is compact and specific about what the skill does, with a clear niche and a concrete verdict vocabulary. Its main weaknesses are the complete absence of a "Use when…" trigger clause and trigger terms that lean on project jargon rather than natural user phrasing, which together cap completeness and trigger-term quality at 3.

Suggestions

Append an explicit trigger clause, e.g. "Use when the user asks to review architecture coverage, check GDD-to-ADR traceability, find requirement gaps, or before the Pre-Production gate."

Add natural trigger phrases users would say — "architecture review", "requirements coverage", "traceability" — alongside the existing jargon (GDD, ADR).

Mention the dependency-cycle/implementation-order analysis and the rtm mode (story/test linkage) so the capability list is comprehensive for the skill's scope.

DimensionReasoningScore

Specificity

"Traceability matrix mapping GDD requirements to ADRs. Finds gaps, cross-ADR conflicts, engine compatibility" names several concrete, distinct actions (build the matrix, find gaps, detect cross-ADR conflicts, check engine compatibility). It is not 5 because coverage is not comprehensive — dependency-cycle detection and the RTM story/test linkage are absent — and "engine compatibility" is elliptical rather than a stated action; it is above 3 because multiple specific capabilities are listed rather than 1-2.

4 / 5

Completeness

The "what" is clear and concrete (maps GDD requirements to ADRs, finds gaps, cross-ADR conflicts, engine compatibility issues, emits a verdict), but there is no "Use when…" clause or equivalent explicit trigger guidance — the verdict list ("PASS/CONCERNS/NOT ASSESSED/FAIL") describes output, not usage timing. Per the rubric, a missing trigger clause caps completeness at 3.

3 / 5

Trigger Term Quality

Relevant domain keywords are present ("gaps", "cross-ADR conflicts", "engine compatibility", "traceability matrix", "GDD", "ADR"), but the natural phrases a user would actually say — "review the architecture", "architecture coverage", "requirements coverage" — are missing, as are common synonyms and the RTM framing. Not 4: keyword coverage is partial and leans on project jargon rather than natural user language.

3 / 5

Distinctiveness Conflict Risk

"Traceability matrix mapping GDD requirements to ADRs" carves a clear niche that would rarely trigger the wrong skill. Not 5: within this project's ecosystem of adjacent review skills (test-evidence-review, gate-check, propagate-design-change) there is minor overlap risk in the shared "review/coverage" trigger space; not 3: the GDD→ADR traceability scope is far more specific than a generic review skill.

4 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (870 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
Donchitos/Claude-Code-Game-Studios
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.