CtrlK
BlogDocsLog inGet started
Tessl Logo

research-lab

使用独立 Research Core 把模型比较、软件评估或其他可证伪问题转化为可恢复、可复现、证据可追溯的研究。 仅用于需要真实执行对照实验的场景:模型/方案/系统的实证对比、基线比较、模型替代评估、消融、盲评、可复现实验。 不用于简单事实查询、只读审计/盘点/调查(即使涉及"证据")、无需实验的低风险判断,或用户只要求普通文案/解释的场景。

72

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, highly actionable protocol with a well-sequenced workflow and explicit validation feedback loops. The only gap is a referenced example file (examples/llm-replacement.md) that does not exist in the bundle.

DimensionReasoningScore

Conciseness

Dense, protocol-style prose that assumes Claude's competence — no padding explaining what baselines/ablations/experiments are — and every section (hard rules, workflow, spec, output contract) earns its place.

5 / 5

Actionability

Names concrete tools (research_validate, research_create, research_inspect, research_execute, research_continue, research_status, research_compare, research_evidence), provides a copy-paste-ready minimal Spec JSON, and enumerates exact Decision values, giving fully executable guidance.

5 / 5

Workflow Clarity

A clear 7-step sequence with explicit validation checkpoints (research_validate, 执行前检查 returning a structured gap on missing conditions, no hot-polling) and a feedback loop via the INVALID/INCONCLUSIVE/UNSUPPORTED/SUPPORTED Decision enum; the batch-operation cap does not apply because validation is present.

5 / 5

Progressive Disclosure

Well-organized into clear sections with a one-level-deep, clearly signaled reference ('相关示例见 examples/llm-replacement.md'), but that referenced example file is not present in the bundle, so navigation does not fully resolve.

4 / 5

Total

19

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly demarcates a niche and answers both what and when with explicit positive triggers and negative boundaries. Its main weakness is specificity: the stated action is a conceptual transformation rather than a concrete list of operations.

Suggestions

Add 1-2 concrete operational actions (e.g., '建立 ResearchSpec、调用 research_ 工具执行对照实验、依据 Evidence 返回 Decision') so the 'what' reads as operations, not just a category.

Consider including a couple more natural trigger phrasings or synonyms (e.g., 'A/B 对比', '跑 benchmark') to broaden trigger coverage toward the comprehensive anchor.

DimensionReasoningScore

Specificity

Names the domain ('模型比较、软件评估或其他可证伪问题') and a clear transformation ('转化为可恢复、可复现、证据可追溯的研究'), but the described action is conceptual rather than a list of concrete operations, and coverage is categorical not comprehensive.

3 / 5

Completeness

Explicitly answers both what ('转化为可恢复、可复现、证据可追溯的研究') and when ('仅用于需要真实执行对照实验的场景') with concrete trigger phrases and explicit negative boundary guidance.

5 / 5

Trigger Term Quality

Includes natural phrases a user would say ('能否替代', '是否显著更好', '哪个更可靠', '恢复', '复现') with some synonym coverage, though not the full comprehensive-with-all-variants level.

4 / 5

Distinctiveness Conflict Risk

A clear niche (empirical comparison experiments requiring reproducibility) with explicit exclusions ('不用于简单事实查询、只读审计...') that sharply reduce conflict with adjacent skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ooooooooooooooooooop/agent-tools
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.