CtrlK
BlogDocsLog inGet started
Tessl Logo

bdistill-behavioral-xray

X-ray any AI model's behavioral patterns — refusal boundaries, hallucination tendencies, reasoning style, formatting defaults. No API key needed.

52

Quality

58%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/bdistill-behavioral-xray/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill provides a clear overview of the bdistill Behavioral X-Ray tool with concrete CLI commands and a useful dimensions table. However, it suffers from moderate verbosity in motivational sections, lacks validation/verification steps in the workflow, and misses opportunities for progressive disclosure through supporting files. The 'When to Use This Skill' section and some explanatory text could be significantly trimmed.

Suggestions

Trim the 'When to Use This Skill' section to 2-3 bullets or remove it entirely — Claude can infer appropriate use cases from the overview

Add a validation/verification step after report generation (e.g., 'Verify the HTML report opens correctly and contains all 6 dimensions')

Include a brief example of actual output or a sample behavioral tag to make the output section more concrete and actionable

Create bundle files for detailed dimension descriptions and example reports, linking to them from the main SKILL.md

DimensionReasoningScore

Conciseness

The skill includes some unnecessary explanations (e.g., 'No API key needed' repeated, 'The AI agent probes itself' explanation, the 'When to Use This Skill' section with 5 bullet points that largely explain obvious use cases). The probe dimensions table and output section are reasonably efficient, but the overall content could be tightened.

3 / 5

Actionability

Provides concrete install commands, CLI invocations (/xray, /xray --dimensions refusal, /xray-report), and MCP natural language prompts. However, there's no example of actual output or how to interpret results, and the MCP setup instruction ('add bdistill-mcp as an MCP server in your project config') is vague for non-Claude Code tools.

4 / 5

Workflow Clarity

The two-step workflow (install → run) is clear but minimal. There are no validation checkpoints — no guidance on what to do if installation fails, if probes produce unexpected results, or how to verify the report was generated correctly. For a tool that generates reports, there should be a verification step.

3 / 5

Progressive Disclosure

The content is reasonably structured with clear sections, but everything is inlined in a single file with no bundle files. The probe dimensions table could link to detailed documentation for each dimension. The reference to '@bdistill-knowledge-extraction' and '/distill --adversarial' are mentioned but not linked or explained. No bundle files are provided to support progressive disclosure.

3 / 5

Total

13

/

20

Passed

Description

58%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description identifies a clear niche — behavioral analysis of AI models — and lists specific dimensions of analysis, giving it reasonable specificity and distinctiveness. However, it lacks an explicit 'Use when...' clause, which is critical for skill selection, and the trigger terms lean toward jargon ('X-ray') rather than natural user language. The 'No API key needed' constraint is a useful differentiator but doesn't compensate for the missing trigger guidance.

Suggestions

Add an explicit 'Use when...' clause with natural trigger phrases, e.g., 'Use when the user wants to test, probe, evaluate, or red-team an AI model's behavior.'

Replace or supplement the metaphorical 'X-ray' with natural user terms like 'test', 'evaluate', 'probe', 'benchmark', 'red-team', or 'analyze LLM behavior'.

Mention concrete outputs or methods (e.g., 'generates adversarial prompts', 'produces behavioral reports') to strengthen specificity.

DimensionReasoningScore

Specificity

Lists several specific actions: analyzing refusal boundaries, hallucination tendencies, reasoning style, and formatting defaults. These are concrete behavioral dimensions to examine. Minor gaps — doesn't mention specific methods or outputs (e.g., generating test prompts, producing reports).

4 / 5

Completeness

The 'what' is reasonably clear — analyzing AI model behavioral patterns across several dimensions. However, there is no explicit 'when' clause (no 'Use when...' or equivalent trigger guidance), which caps this at 3 per the rubric guidelines.

3 / 5

Trigger Term Quality

Includes some relevant terms like 'AI model', 'refusal boundaries', 'hallucination', 'reasoning style', but misses natural user phrases like 'test a model', 'probe', 'red team', 'benchmark', 'evaluate LLM', 'model behavior'. 'X-ray' is metaphorical rather than a natural search term.

3 / 5

Distinctiveness Conflict Risk

The focus on probing AI model behavioral patterns (refusal, hallucination, reasoning style) is a fairly distinct niche. The 'No API key needed' detail adds differentiation. Minor overlap risk with general AI evaluation or red-teaming skills, but the specific behavioral dimensions help distinguish it.

4 / 5

Total

14

/

20

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
administrakt0r/AI-Agents-Safe-Coding-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.