CtrlK
BlogDocsLog inGet started
Tessl Logo

research-lab

使用独立 Research Core 把模型比较、软件评估或其他可证伪问题转化为可恢复、可复现、证据可追溯的研究。 仅用于需要真实执行对照实验的场景:模型/方案/系统的实证对比、基线比较、模型替代评估、消融、盲评、可复现实验。 不用于简单事实查询、只读审计/盘点/调查(即使涉及"证据")、无需实验的低风险判断,或用户只要求普通文案/解释的场景。

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is research-lab in ooooooooooooooooooop/agent-tools

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable protocol with concrete tool calls, a runnable Spec example, and a clear sequenced workflow with validation checkpoints. It is held just below top marks by minor placeholder gaps, the absence of an explicit retry loop, and a referenced example file that is not bundled.

Suggestions

Add an explicit 'if research_validate returns INVALID → fix the Spec → re-validate' retry loop in 标准流程 to close the workflow feedback-loop gap.

Fill the placeholder '...' values in the minimal Spec example (case input/expected) so the JSON is fully copy-paste-runnable.

Provide the actual invocation shape for the research_* tools (arguments/flags) rather than only naming them, or ship the referenced examples/llm-replacement.md so the signaled reference resolves.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence (no concept tutorials), with every section earning its place; minor redundancy between the 触发门禁 section and the frontmatter description keeps it just below anchor 5.

4 / 5

Actionability

Provides concrete tool calls (research_validate, research_create, research_inspect, research_execute, research_continue, research_compare, research_evidence), a copy-pasteable JSON Spec, and a Decision enum; minor gaps remain (placeholder '...' fields, no full command invocation syntax).

4 / 5

Workflow Clarity

标准流程 gives a clear 7-step sequence with explicit validation checkpoints (执行前检查 with research_inspect, returning structured gaps) and error-recovery enums (UNSUPPORTED/INVALID/INCONCLUSIVE), but lacks an explicit 'validate → fix → re-validate' retry loop, so it sits just under anchor 5.

4 / 5

Progressive Disclosure

Well-organized into clearly headed sections with content appropriately inline and a single clearly signaled one-level reference ('examples/llm-replacement.md'); the referenced file is not present in the bundle, a minor organization gap versus anchor 5.

4 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, clearly bounded, and explicitly covers both what the skill does and when to use it, with strong distinctiveness from explicit exclusions. Trigger-term coverage is good but could add a few more natural synonyms.

DimensionReasoningScore

Specificity

Names the research domain and several concrete experiment types ('模型比较、软件评估…实证对比、基线比较、模型替代评估、消融、盲评、可复现实验'), giving multiple specific capabilities, though framed as use-cases rather than discrete actions like the anchor-5 example.

4 / 5

Completeness

Explicitly answers both 'what' ('把…可证伪问题转化为可恢复、可复现、证据可追溯的研究') and 'when' with concrete triggers ('仅用于需要真实执行对照实验的场景…不用于简单事实查询…'), matching the anchor-5 example of clear what + when with trigger phrases.

5 / 5

Trigger Term Quality

Includes natural user-facing phrases ('模型比较', '软件评估', '模型替代评估', '可证伪问题') a user would say, with good coverage but missing a few common synonyms or explicit file/extension-style triggers.

4 / 5

Distinctiveness Conflict Risk

A clear niche (controlled empirical research via a dedicated Research Core) with explicit '不用于' exclusions (audits, simple queries, copywriting) that sharply reduce overlap with other skills, matching the anchor-5 distinct-trigger profile.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ooooooooooooooooooop/personal-ai
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.