CtrlK
BlogDocsLog inGet started
Tessl Logo

risk-analysis

Risk measurement and stress testing — VaR/CVaR/max drawdown calculation, Monte Carlo simulation, extreme-value tail-risk analysis, and historical scenario stress testing.

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./agent/src/skills/risk-analysis/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, highly actionable body whose module-specific conventions (sign discipline, seed requirements, stderr-based tail typing) are exactly what a skill should add, but it carries a layer of textbook finance theory Claude already knows and inlines reference-style catalogs that belong in separate files. The workflow is clear with embedded sanity checks, though they are not packaged as an explicit validate/fix/retry loop.

Suggestions

Trim the textbook sections — the VaR/CVaR definitions, the three-method advantages/disadvantages table, the subadditivity/Basel row, and kurtosis/skewness explanations — to one line each or drop them; keep only the module-specific conventions and gotchas.

Move the historical scenario table and the hypothetical STRESS_SCENARIOS catalog into a references/ file (e.g. references/scenarios.md) and link to it from the stress-testing section.

Promote the embedded sanity checks (cvar >= var, shape_xi stability across thresholds) into an explicit validation step in 'Analysis Steps' with a fix-and-retry instruction.

DimensionReasoningScore

Conciseness

Excellent module-specific guidance ('A loss is a positive number', the 2σ tail_type rule, the parametric-vs-historical direction table) sits alongside textbook material Claude already knows, e.g. 'VaR (Value at Risk) ... the maximum expected loss over a given horizon', the three-method advantages/disadvantages table, the VaR-vs-CVaR subadditivity/Basel comparison, and kurtosis/skewness explanations. Not 4 because these known-concept sections are a noticeable share of the body; not 2 because the bulk is genuinely non-obvious implementation knowledge, not padded filler.

3 / 5

Actionability

Concrete, copy-paste-ready calls throughout: 'historical_var(returns, confidence=0.99, horizon=10)', 'monte_carlo_gbm(s0=100.0, ..., seed=42)', 'fit_gpd_tail(returns, threshold_pct=5.0)', plus exact return shapes ('paths.shape # (10000, 253)') and failure behavior ('fewer than 2 exceedances raises'). The common cases are covered by specific examples; nothing is pseudocode.

5 / 5

Workflow Clarity

'Analysis Steps' gives a concrete 7-step sequence with numeric parameters ('compare three methods at both 95% and 99%', '10,000 paths', 'at least 3 historical scenarios + 2 hypothetical'), and there are real checkpoints ('If you ever compute a CVaR below its VaR, the tail mask is wrong', 'Check that shape_xi is stable across a few nearby threshold_pct values'). Not 5 because the checkpoints are embedded in method notes rather than an explicit validate-and-retry loop tied to the step sequence; not 3 because validation guidance is present, just not formatted as a loop.

4 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/), and the body is a single ~300-line file with good section headers. The historical scenario catalog, EVT theory, and method-comparison tables ('2008 financial crisis ... -65%', the GPD/POT explanation) are reference-style content that clearly belongs in a separate one-level-deep file. Not 2 because structure and navigation within the file are solid with consistent headers; not 4 because substantial separable content is inlined in SKILL.md itself.

3 / 5

Total

15

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-scoped description with strong domain keywords and comprehensive capability coverage, but it omits any 'Use when' trigger guidance, so a user (or Claude) deciding when to invoke it gets no explicit signal. Adding a trigger clause would lift it into the top tier.

Suggestions

Append a trigger clause such as: 'Use when the user asks about portfolio risk, drawdown, VaR/expected shortfall, tail risk, or wants to stress-test a backtest or allocation.'

Include common synonyms in the description — 'value at risk', 'expected shortfall (ES)', 'GPD/EVT tail fitting' — so natural phrasings match.

DimensionReasoningScore

Specificity

Quotes: 'VaR/CVaR/max drawdown calculation, Monte Carlo simulation, extreme-value tail-risk analysis, and historical scenario stress testing' — four concrete, distinct capabilities that comprehensively cover the skill's domain. Not below 4 because there are no coverage gaps; not applicable above since 5 is the top anchor and this matches it.

5 / 5

Completeness

Quotes: the description states what the skill does but contains no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. It has a clear, multi-part 'what' (score above 3 territory) but the 'when' is entirely absent, not just weakly implied — so 4 is unreachable without explicit trigger guidance.

3 / 5

Trigger Term Quality

Quotes: 'VaR', 'CVaR', 'max drawdown', 'Monte Carlo simulation', 'stress testing', 'tail-risk' — strong natural keywords a quant user would actually say. Not 5 because common synonyms are missing: 'value at risk' spelled out, 'expected shortfall'/'ES', 'backtest risk'. Not 3 because coverage is clearly good, not just 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

Quotes: 'VaR/CVaR', 'extreme-value tail-risk analysis', 'historical scenario stress testing' — a clear financial-risk niche with distinct technical triggers. Not 5 because the generic phrase 'stress testing' also denotes software/load testing, creating a minor collision risk with non-financial skills.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
HKUDS/Vibe-Trading
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.