CtrlK
BlogDocsLog inGet started
Tessl Logo

grill-me

针对方案或设计的高强度追问式面试(adversarial design review / grill session),暴露假设漏洞与缺失约束,过程中同步维护领域模型(术语表和 ADR)。手动调用 /grill-me。

56

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/grill-me/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-written, imperative instruction skill with concrete behavioral examples, explicit checkpoints (consensus gate, ADR three-condition gate, code cross-check), and lean, project-specific guidance. Its main structural weakness is that everything lives in one ~200-line file: the CONTEXT.md and ADR format specifications read as reference material that should be split into references/ files.

Suggestions

Move the two '# 参考:…格式' sections into references/context-format.md and references/adr-format.md, leaving one-line pointers in the main body — this keeps SKILL.md a lean overview and fixes the progressive-disclosure gap.

Add one concrete example scenario under "讨论具体场景" (e.g., a partial-cancel vs full-cancel boundary case for Order) so the abstract instruction becomes actionable.

Deduplicate the lazy-creation rule (stated at lines 63, 157, and 167) into a single statement to tighten conciseness.

DimensionReasoningScore

Conciseness

The body is dense and project-specific — file layouts, lazy-creation rules, glossary/ADR templates, and exact behaviors with sample dialogue — with almost no explanation of concepts Claude already knows. Not 5 because there is minor redundancy: lazy-creation of files is stated three times (lines 63, 157, 167) and the single- vs multi-context structure is described both in the file-structure section and again in the CONTEXT.md format section. Not 3 because the padding is minor and the prose is consistently imperative rather than explanatory.

4 / 5

Actionability

Most guidance is directly executable: concrete sample challenges ("你的术语表把'取消'定义为 X,但你现在似乎是指 Y——到底是哪个?"), copy-ready templates for CONTEXT.md and ADR entries, exact numbering rules ("扫描 docs/adr/ 找到当前最大编号,加一"), and a precise three-condition ADR gate. Not 5 because a few behaviors stay abstract — "构造探索边界条件的场景" (construct boundary-condition scenarios) gives no example scenario, and the questioning flow itself has no sample exchange showing question → recommendation → consensus. Not 3 because the majority of instructions are concrete enough to act on immediately.

4 / 5

Workflow Clarity

The workflow is well sequenced: numbered startup checks (read CONTEXT.md/CONTEXT-MAP.md and docs/adr/ first), then the questioning loop with an explicit hard gate — "在我明确确认达成共识之前,不要开始执行方案" (do not begin executing until consensus is explicitly confirmed) — plus a three-condition checklist before any ADR is created and a cross-check-with-code feedback behavior. Not 5 because there is no recovery guidance for common failure modes (e.g., what to do when the user disputes a terminology conflict or the discussion drifts off-context in a multi-context repo beyond "问"). Not 3 because validation checkpoints (consensus gate, ADR gate, code cross-check) are explicit, not implicit.

4 / 5

Progressive Disclosure

The single SKILL.md is well sectioned, and the two reference blocks are clearly labeled ("# 参考:CONTEXT.md 格式", "# 参考:ADR 格式"), but there are no bundle files at all (references/, scripts/, assets/ are absent) and ~100 lines of format templates are inlined in the main skill file. Anchor 3 fits: structure exists but content that could be separate (the two format references, especially the multi-context map example) is inline. Not 4 because the format sections are substantial enough that they belong in separate reference files to keep the overview lean; not 2 because sections are clearly headed and navigation within the file is easy.

3 / 5

Total

15

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a clear, fairly specific purpose with an unusual differentiator (domain-model maintenance during the review), but its trigger guidance is weak: it explains how the skill is invoked rather than when a user should want it, and natural trigger phrases are sparse beyond the English parenthetical.

Suggestions

Add an explicit 'when to use' clause, e.g. "适用于:用户提出方案/设计并希望被严格挑战、在动工前暴露隐藏假设时" (use when the user proposes a plan/design and wants it challenged before implementation), to satisfy the completeness dimension.

Broaden natural trigger terms to include phrases users actually say, such as "挑战我的方案", "帮我把关设计", "design review", "stress-test my plan", alongside the existing "grill session".

Keep the concrete action list but consider stating the one-question-at-a-time interactive format, which distinguishes it from one-shot review skills and further reduces conflict risk.

DimensionReasoningScore

Specificity

The description names the domain (方案或设计 "plans or designs") and several concrete actions: "高强度追问式面试" (adversarial questioning), "暴露假设漏洞与缺失约束" (exposing assumption gaps and missing constraints), and "同步维护领域模型(术语表和 ADR)" (maintaining a glossary and ADRs). Not score 5 because it does not enumerate the full range of behaviors found in the body (e.g., cross-checking terminology against code, one-question-at-a-time pacing); not 3 because more than 1-2 concrete actions are explicitly stated.

4 / 5

Completeness

The "what" is clear (adversarial design review that exposes assumption holes and maintains a domain model), but the "when" guidance is limited to "手动调用 /grill-me" (manual invocation), which says how it is invoked, not the situations that warrant it — there is no 'Use when...' clause or equivalent. Per the judging guideline, a missing explicit 'Use when...' clause caps completeness at 3. Not 4 because the invocation note is not equivalent to explicit trigger situations; not 2 because the what is fully clear, not vague.

3 / 5

Trigger Term Quality

It includes some natural keywords users would say — "adversarial design review / grill session" and "追问" — but misses common variations a user might naturally utter such as "challenge my design", "poke holes in this plan", "pre-mortem", or "design critique". Not 4 because the natural-phrase coverage is thin beyond the English parenthetical; not 2 because the included terms are genuinely natural rather than pure jargon.

3 / 5

Distinctiveness Conflict Risk

The niche is distinct — an interactive adversarial review session with built-in domain modeling (glossary + ADR maintenance) differs from code-review or documentation skills, and the explicit manual-invocation note further reduces accidental triggering. Minor overlap risk remains with general design-review or architecture-discussion skills. Not 5 because "方案或设计" (plans or designs) review is a fairly broad category; not 3 because the glossary/ADR maintenance component is a clear differentiator.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 3 missing, 3 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
feiskyer/claude-code-settings
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.