CtrlK
BlogDocsLog inGet started
Tessl Logo

caveman-evidence-review

Read-only review of Caveman Cloud evidence: cost, Cave Score, workflows, traces, latency, errors, routing, savings. Use when asked what Caveman found or where LLM spend goes.

70

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable read-only review workflow with concrete tooling and clear guardrails. It is held back only by minor conciseness and progressive-disclosure refinements and the absence of explicit error-recovery feedback loops.

Suggestions

Tighten the report template and consolidate the per-step CLI-fallback blocks (e.g., one fallback reference) to reduce token bulk.

Consider offloading the trace-filter taxonomy or report template into a references/ file and linking to it from SKILL.md to improve progressive disclosure.

Add an explicit validate→adjust loop for Step 3 (e.g., if a cohort is too small to support a claim, widen the window and re-run) to push workflow clarity to the top anchor.

DimensionReasoningScore

Conciseness

The body is lean and assumes domain competence — short directives, no padding about what Caveman is — but the report template and repeated CLI-fallback blocks add some bulk that could be trimmed slightly, so it is not maximally token-efficient.

4 / 5

Actionability

It names specific MCP tools (caveman_context, caveman_report, caveman_plan, caveman_trace_search, caveman_trace_get) with concrete report types and a comprehensive trace-filter taxonomy, plus executable CLI fallback commands; this covers the common cases copy-paste ready.

5 / 5

Workflow Clarity

Five steps are clearly sequenced with explicit stop-conditions (stop if login/project missing; stop at strongest supported statement), but there are no validate→fix→retry feedback loops, leaving minor checkpoint gaps relative to the top anchor.

4 / 5

Progressive Disclosure

Content is well-organized into labeled sections (Hard rules, Steps 1–5) with no nested references and easy navigation, but at ~134 lines everything is inlined in SKILL.md with no offloaded detail, so it is good rather than exemplary structure.

4 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concise, and explicit about both capability and trigger conditions, scoped to a distinct niche. The only weakness is slightly limited action and synonym coverage.

DimensionReasoningScore

Specificity

Names the domain ("Caveman Cloud evidence") and enumerates several concrete review targets — cost, Cave Score, workflows, traces, latency, errors, routing, savings — but the action set is essentially a single verb (review), leaving minor coverage gaps versus a fully comprehensive action list.

4 / 5

Completeness

It explicitly answers both what (read-only review of the listed evidence types) and when ("Use when asked what Caveman found or where LLM spend goes") with concrete trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

"Use when asked what Caveman found or where LLM spend goes" supplies natural phrases a user would say, with good keyword coverage (Caveman, LLM spend); a few synonyms (e.g. AI spend, model costs) are missing, so it stops short of comprehensive.

4 / 5

Distinctiveness Conflict Risk

The Caveman Cloud niche and its specific triggers ("what Caveman found", "LLM spend") form a clear, distinct niche with minimal overlap risk against other skills.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
JuliusBrussee/caveman
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.