CtrlK
BlogDocsLog inGet started
Tessl Logo

statistical-theory-analysis

Analyze theoretical properties of statistical methods under the formal formulation: identifiability, bias, variance, consistency, asymptotics, coverage, error bounds, robustness, and limitations.

59

Quality

69%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./external/agents/stat_research_agent/skills/statistical-theory-analysis/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A compact, well-structured theory skill that respects the token budget and gives a useful output template, but it lacks concrete derivation techniques and an explicit sequenced workflow with validation. Adding a short 'how to derive each output' guide and an ordered process would raise the weaker dimensions.

Suggestions

Add a brief ordered process (e.g. 1. state assumptions, 2. derive the target property, 3. write the proof sketch, 4. generate an empirical prediction) with a self-check that the prediction is testable.

For each Theory Output, give one concrete technique or starting point (e.g. 'consistency: show estimator converges in probability via LLN/Slutsky') instead of only naming the deliverable.

Include a validation checkpoint confirming assumptions are explicit and the experimental prediction names a concrete baseline and stress condition.

DimensionReasoningScore

Conciseness

The body is lean (~50 lines), assumes Claude already knows what bias/variance/consistency mean, and every section earns its place with no padding or over-explanation.

5 / 5

Actionability

It gives a concrete deliverable checklist and a fill-in theorem template, but the actual analytical methods for deriving each property are left unspecified, so guidance is actionable as a structure yet incomplete on execution.

3 / 5

Workflow Clarity

The Overview places the skill in a pipeline ('after method proposal and before final experimental comparison') and sections imply an order, but there is no explicit numbered sequence or validation checkpoint for the theory-producing process.

3 / 5

Progressive Disclosure

The skill is under 50 lines, single-purpose, needs no external references, and is organized into clear well-labeled sections (Overview, Theory Outputs, Theorem Template, Experimental Predictions), satisfying the simple-skill exception.

5 / 5

Total

16

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-targeted description that names many concrete theoretical deliverables, but it omits any explicit 'Use when...' trigger guidance, which caps completeness. Adding a usage clause would lift the weakest dimension.

Suggestions

Append an explicit trigger clause, e.g. 'Use when a user requests proofs, consistency/asymptotic arguments, bias-variance analysis, coverage, or robustness checks for a statistical method.'

Vary the action verbs beyond a single 'Analyze' (e.g. 'Derive bias/variance, prove consistency, establish asymptotic distributions') to strengthen specificity.

Surface a couple of natural synonyms users say ('proof', 'theory', 'guarantees') inside the description, not only in trigger-keywords metadata.

DimensionReasoningScore

Specificity

The description lists many specific analysis targets ('identifiability, bias, variance, consistency, asymptotics, coverage, error bounds, robustness, and limitations'), but they all hang off a single verb 'Analyze' rather than multiple distinct action verbs, so it sits just below the comprehensive multi-action anchor.

4 / 5

Completeness

It clearly answers 'what' (analyze theoretical properties of statistical methods) but contains no 'Use when...' clause or equivalent explicit trigger guidance, so per the rubric completeness is capped at 3.

3 / 5

Trigger Term Quality

It surfaces natural domain terms a researcher would say ('bias', 'variance', 'consistency', 'coverage', 'robustness'), giving good keyword coverage, though common synonyms and the bare term 'theory'/'proof' live only in metadata rather than the description itself.

4 / 5

Distinctiveness Conflict Risk

The formal-theory niche (identifiability, asymptotics, error bounds) is fairly distinct with specific triggers, with only minor overlap risk against general statistical-analysis skills.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
aiming-lab/AutoResearchClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.