CtrlK
BlogDocsLog inGet started
Tessl Logo

behavior-preservation-checker

Compare runtime behavior between original and migrated repositories to detect behavioral differences, regressions, and semantic changes. Use when validating code migrations, refactorings, language ports, framework upgrades, or any transformation that should preserve behavior. Automatically compares test results, execution traces, API responses, and observable outputs between two repository versions. Provides actionable guidance for fixing deviations and ensuring behavioral equivalence.

81

2.62x
Quality

73%

Does it follow best practices?

Impact

97%

2.62x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/behavior-preservation-checker/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

53%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-structured and mostly actionable, but suffers from verbosity, six missing referenced scripts, and a lack of explicit validation feedback loops in its core workflow. Tightening the inlined methods/patterns into references and adding verify-before-proceed checkpoints would raise it substantially.

Suggestions

Add the six missing scripts referenced in the body (compare_api_outputs.py, isolate_difference.py, find_missing_functions.py, compare_api_contracts.py, benchmark_comparison.py, detect_state_issues.py) or remove those command examples, so the documented commands are actually executable.

Insert explicit validation checkpoints into the core workflow — e.g., after running comparisons, a 'validate report -> triage critical differences -> fix -> re-run' feedback loop — to satisfy the batch-operation validation requirement.

Move the four comparison methods and five actionable-guidance patterns into the existing reference files (comparison_techniques.md, difference_patterns.md), keeping SKILL.md a concise overview with one-level-deep links.

DimensionReasoningScore

Conciseness

The body is mostly efficient with concrete code examples and little 'what is X' padding, but it is noticeably verbose with repetition across four comparison methods and five patterns and some unnecessary explanatory sections (severity levels, basic equality/tolerance snippets). It fits the 'mostly efficient but could be tightened' anchor.

3 / 5

Actionability

Provides concrete, copy-paste-ready commands and code throughout, but six of the ten referenced scripts (compare_api_outputs.py, isolate_difference.py, find_missing_functions.py, compare_api_contracts.py, benchmark_comparison.py, detect_state_issues.py) do not exist in the bundle, so those examples are not actually executable — a significant gap keeping it below 4.

3 / 5

Workflow Clarity

The core workflow lists four sequenced steps, but Step 4 ('Fix Deviations') is vague ('Follow actionable guidance') and there are no explicit validation checkpoints or validate→fix→retry feedback loops. Per the rubric cap, missing validation for batch/comparison operations caps this at 3.

3 / 5

Progressive Disclosure

Good structure with two real reference files (comparison_techniques.md, difference_patterns.md) and scripts, clearly signaled in a Resources section. Not a 5 because substantial content that could live in the references (five detailed patterns, four comparison methods) is inlined in SKILL.md rather than split out.

4 / 5

Total

13

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and distinct, clearly stating both what the skill does and when to use it with concrete natural-language triggers. It is only slightly held back on trigger-term breadth, lacking a few colloquial synonyms.

DimensionReasoningScore

Specificity

Lists multiple concrete actions (comparing test results, execution traces, API responses, observable outputs) and provides actionable guidance for fixing deviations, giving comprehensive coverage of what the skill does.

5 / 5

Completeness

Explicitly answers both 'what' ('Compare runtime behavior...detect behavioral differences, regressions, and semantic changes') and 'when' ('Use when validating code migrations, refactorings, language ports, framework upgrades...') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural trigger phrases users would say ('code migrations, refactorings, language ports, framework upgrades') with good synonym coverage, though missing a few common variations. Not a 5 because it lacks the broader colloquial/extension-style terms seen in the top anchor.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — comparing runtime behavior between two repository versions — with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 6 missing

Warning

Total

15

/

16

Passed

Repository
ArabelaTso/Skills-4-SE
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.