CtrlK
BlogDocsLog inGet started
Tessl Logo

engineer-system-change

Evaluate and carry out non-trivial software-system changes from first principles. Use when assessing RFCs, issues, designs, features, refactors, migrations, dependency changes, or proposed fields, events, APIs, modules, and services whose need, consumers, system fit, validation, or rollback require scrutiny. Read the actual system, identify the concrete problem and named semantic consumers, choose the smallest sufficient solution, reject pseudo-requirements and speculative abstractions, and require evidence proportional to risk. Do not use for mechanical edits, source-code explanation, or a dedicated review of an already-complete diff.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tightly written decision-framework skill: a clear six-gate sequence with explicit validation checkpoints, a defined verdict vocabulary, and proportional output guidance, all in lean directive prose. The main costs are redundant restatement of verdict semantics across sections and the absence of any worked example to ground the abstract gates.

Suggestions

Consolidate the verdict semantics: gates 1 and 4 restate STOP/NEEDS_EVIDENCE conditions that the "Use Explicit Verdicts" section defines again; state each condition once and cross-reference it.

Add one compact worked example — e.g., a filled-in consumer-ledger entry or a two-line sample verdict with its blocking gate — to make gates 2 and the verdict mapping concrete.

Consider moving the consumer-question table and the consequence-dimension checklist to a reference file so SKILL.md stays a lean overview of the gates.

DimensionReasoningScore

Conciseness

The prose is dense and directive throughout — every line prescribes behavior and nothing explains concepts Claude already knows. However, verdict semantics are stated redundantly: gate 1 ("Return `STOP` only when... Return `NEEDS_EVIDENCE` when..."), gate 4 ("Evidence labels classify individual claims; verdicts classify the overall decision"), and the verdicts section each re-explain overlapping conditions, which could be consolidated. Fits the 4 anchor (efficient, minor instances that could be trimmed); not 5 because of this repetition, and not 3 since there is no genuinely unnecessary explanation.

4 / 5

Actionability

Highly concrete guidance for an instruction-only skill: an ordered solution-preference ladder ("1. No product or code change ... 7. Introduce a new subsystem"), a required-answer table for consumers, a defined verdict vocabulary, and a proportional output template ("For a simple change, report only: 1-5"). Not 5 because there are no worked examples — e.g., a sample consumer-ledger entry or one applied verdict — that would make the abstract gates copy-paste concrete; not 3 because the guidance is specific and executable as written, not high-level hints.

4 / 5

Workflow Clarity

The six numbered decision gates form a clear sequence with explicit validation checkpoints (return conditions at gates, the verdict vocabulary, and gate 6's "Map each important result claim to observed evidence... Verify negative boundaries and failure behavior, not only the happy path"). Since this skill governs potentially irreversible changes, the required rehearsal/rollback checks ("Require a production-like rehearsal plus executable containment or rollback for irreversible changes") satisfy the destructive-operation feedback-loop requirement. Not 4: validation checkpoints and error-recovery conditions (NEEDS_EVIDENCE loops, re-validate after fix) are explicit at every stage.

5 / 5

Progressive Disclosure

A single-file process skill with well-organized, clearly headed sections (task boundary, six gates, verdicts, proportional output) and no broken or buried references — appropriate for a cohesive workflow. Fits the 4 anchor (good structure, content appropriately placed, minor organization gaps); not 5 because at ~127 lines the consumer-question table and the consequence-dimension checklist are candidates for a reference file that would keep the overview leaner, and the simple-skill (<50 lines) exception does not apply.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person, concrete, and comprehensive on both capabilities and triggers, with rare negative-trigger guidance that sharply bounds the skill against adjacent skills. Every dimension sits at the top anchor with no padding or over-claims.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions — "Read the actual system, identify the concrete problem and named semantic consumers, choose the smallest sufficient solution, reject pseudo-requirements and speculative abstractions, and require evidence proportional to risk" — comprehensively covering the skill's capability surface. Matches the 5 anchor (multiple specific concrete actions, comprehensive coverage); not 4 because there are no noticeable coverage gaps.

5 / 5

Completeness

Explicitly answers both questions: what ("Evaluate and carry out non-trivial software-system changes... Read the actual system, identify... choose the smallest sufficient solution") and when ("Use when assessing RFCs, issues, designs, features, refactors, migrations, dependency changes..."). It even adds negative trigger guidance ("Do not use for mechanical edits, source-code explanation..."), exceeding the 5 anchor's requirements.

5 / 5

Trigger Term Quality

Comprehensive natural terms users would actually say: "RFCs, issues, designs, features, refactors, migrations, dependency changes" plus "fields, events, APIs, modules, and services". These are the natural vocabulary for change-engineering requests; not 4 because coverage includes both the artifact types and the assessment verbs with no notable missing synonyms.

5 / 5

Distinctiveness Conflict Risk

The "non-trivial... whose need, consumers, system fit, validation, or rollback require scrutiny" framing plus the explicit exclusion clause ("Do not use for mechanical edits, source-code explanation, or a dedicated review of an already-complete diff") carves a clear niche with minimal conflict risk against adjacent coding and review skills. Not 4: the exclusions leave no meaningful overlap ambiguity.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
bytedance/deer-flow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.