CtrlK
BlogDocsLog inGet started
Tessl Logo

orca-replay

Answers questions about a past agent run from its recording rather than from memory, and replays or forks that run. Use when asked why an earlier run did something, or to reproduce a failure.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable and well-sequenced with strong validation gates for destructive operations, and it avoids explaining concepts Claude already knows. The main room for improvement is trimming the repetition of side-effect/approval warnings across steps.

Suggestions

Consolidate the repeated external-side-effect and approval warnings from steps 4 and 5 into a single shared 'Approval gate' subsection that both steps reference, to tighten conciseness.

Consider moving the detailed 'What offline covers' and 'What a matching replay proves' rationale into a references/ file, keeping SKILL.md as a lean overview with one-level-deep pointers.

DimensionReasoningScore

Conciseness

The body is dense and assumes competence (no padding about what a recording or a worktree is), but the side-effect/approval warnings recur across steps 4 and 5 and could be consolidated, so a few tokens do not fully earn their place.

4 / 5

Actionability

Concrete tool invocations with arguments are given throughout (orca_graph with to:, orca_replay with worktree: true, orca_compare with from/verify) plus executable shell commands (orca record claude, orca export last -o run.html, orca scrub) and a tool argument table.

5 / 5

Workflow Clarity

A clearly numbered 1-5 workflow is sequenced with an explicit hard-gate validation checkpoint (preview-and-confirm before replay) and gated approval for destructive and external-side-effect operations, satisfying the feedback-loop expectation for destructive skills.

5 / 5

Progressive Disclosure

The skill is self-contained with well-organized section headers (When to Use, Workflow, Limitations, Tools) and no nested references, but at ~190 lines some of the safety rationale could live in a referenced file; structure is good rather than exemplary.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it states concrete capabilities, gives explicit 'Use when' trigger guidance, and occupies a distinct niche. The only weakness is that trigger synonyms and the broader action vocabulary (scrub, export, compare) are deferred to the body.

DimensionReasoningScore

Specificity

Names the domain and two concrete actions ('Answers questions about a past agent run from its recording' and 'replays or forks that run'), but does not enumerate the fuller action set (scrub, export, compare) that the body reveals, leaving minor coverage gaps.

4 / 5

Completeness

It explicitly answers both what ('Answers questions... from its recording... replays or forks that run') and when ('Use when asked why an earlier run did something, or to reproduce a failure') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Natural triggers are present ('why an earlier run did something', 'reproduce a failure'), matching phrases a user would actually say, but synonyms and variations (e.g. 'what changed this file', 'does it still reproduce') live in the body rather than the description.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche (recorded agent runs) with distinct triggers ('past agent run', 'its recording', 'reproduce a failure') that are unlikely to fire for unrelated skills.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
iflytek/skillhub
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.