CtrlK
BlogDocsLog inGet started
Tessl Logo

research-lab

使用独立 Research Core 把模型比较、软件评估或其他可证伪问题转化为可恢复、可复现、证据可追溯的研究。 用于用户要求研究、对照实验、基线比较、模型替代评估、消融、盲评、证据审计,或问题不能仅凭常识可靠下结论的场景。 不用于简单事实查询、无需实验的低风险判断,或用户只要求普通文案/解释的场景。

66

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/research-lab/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a tight, actionable protocol that names specific runtime tools, gives a concrete minimal Spec, and sequences the workflow with validation checkpoints. Its main weaknesses are an implicit retry loop and an external example reference that is not present in the bundle.

Suggestions

Make the fix-retry feedback loop explicit in 标准流程 (e.g. 'INVALID → revise Spec → re-run research_validate → only then research_execute') to push workflow_clarity to 5.

Either include `examples/llm-replacement.md` in the bundle or drop/soften the pointer so the single external reference is verifiable.

Trim the '目标' framing paragraph to one line since the 触发门禁 and 硬规则 sections already convey scope.

DimensionReasoningScore

Conciseness

Dense and mostly lean — terse hard rules and a focused workflow assume Claude already knows research methodology; the '目标' paragraph and the illustrative JSON spec carry minor over-explanation that could be trimmed, keeping it just below score 5.

4 / 5

Actionability

Concrete, specific guidance throughout — named `research_*` tool calls per step, a copy-paste-ready minimal Spec JSON, and enumerated Decision values; the caveat '字段以安装版本的 research validate 为准' makes the spec illustrative rather than fully executable as written, a minor gap.

4 / 5

Workflow Clarity

A clear 7-step sequence with explicit checkpoints (执行前检查 via `research_inspect`, structured gap return, `research_status`, `research_continue` recovery) and INVALID/INCONCLUSIVE/UNSUPPORTED handling; just short of score 5 because an explicit 'fix-and-re-validate' retry loop is implied rather than spelled out.

4 / 5

Progressive Disclosure

Well-organized sections (目标, 触发门禁, 硬规则, 标准流程, Spec, 输出契约) with a single clearly-signaled one-level reference to `examples/llm-replacement.md`; not score 5 because that referenced file is absent from the review bundle, so the pointer cannot be verified.

4 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is third-person, lean, and clearly states both capability and trigger conditions with a helpful exclusion clause. It is specific and well-differentiated, with only minor room to widen trigger synonyms and sharpen operational action verbs.

Suggestions

Add a few more colloquial trigger synonyms (e.g. '跑个对比', '哪个更好', '复现上次实验') to lift trigger-term coverage toward score 5.

Lead with one concrete operational verb (e.g. '编排可恢复实验') before the scenario list so the 'what' reads as an action rather than a domain enumeration.

DimensionReasoningScore

Specificity

Names the domain and several concrete scenarios — '模型比较、软件评估...可证伪问题转化为可恢复、可复现、证据可追溯的研究' — covering comparison, evaluation, ablation, blind review and evidence audit; just short of score 5 because it enumerates use-cases rather than concrete operational actions.

4 / 5

Completeness

Explicitly answers both 'what' (convert falsifiable questions into reproducible, evidence-traceable research) and 'when' via a concrete '用于...场景' trigger clause plus a '不用于...' exclusion clause, matching the score-5 anchor.

5 / 5

Trigger Term Quality

Good coverage of natural user phrasing including synonyms — '研究', '对照实验', '基线比较', '模型替代评估', '消融', '盲评', '证据审计' — though it stops short of the exhaustive synonym/extension coverage of score 5.

4 / 5

Distinctiveness Conflict Risk

Clear niche (falsifiable, experiment-backed research) with specific triggers and an explicit exclusion clause that reduces overlap; not score 5 because the broad term '研究' could still brush against general research-adjacent skills.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ooooooooooooooooooop/agent-tools
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.