CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/interviewer-calibration-guide-author

Build-an-X workflow that produces an interviewer calibration guide for a QA hiring loop - takes the question bank and rubric (from sibling skills `interview-question-author` and `hiring-rubric-author`) plus 2-5 sample candidate transcripts/responses, and emits gold-standard model answers, common pitfalls, score-anchor examples per question, and a calibration-session script for the panel. Distinct from the question and rubric skills (which produce the questions and the scoring scaffold); this skill produces the **demonstration material** that brings two interviewers' scores into agreement. Use after the rubric exists and before the first real candidate - the calibration guide is what closes the inter-rater-reliability gap that the rubric alone cannot.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, substantive authoring skill: highly actionable worked-example templates, a clearly sequenced five-step workflow with explicit validation gates and feedback loops, and well-organized sections with one-level external references. Its only weakness is conceptual preamble in the Overview that restates rationale Claude largely already knows.

Suggestions

Tighten the Overview and 'third leg of the tripod' framing - remove rationale Claude can infer and lead with the skill's concrete purpose and output.

Trim the per-reference annotation sentences in the References section to a clause each; the URLs and sibling-skill names carry most of the signal.

Consider moving the long four-level worked-example block into a references/ template file so the body stays a lean overview, leaving the in-skill example as a single representative question.

DimensionReasoningScore

Conciseness

The worked examples, pitfall tables, and calibration script are load-bearing and concrete, but the Overview and tripod-framing prose restate inter-rater-reliability rationale and structured-interview concepts Claude largely already knows, which could be tightened - matching the score-2 'mostly efficient but includes some unnecessary explanation' anchor rather than the lean score-3 anchor.

2 / 3

Actionability

Provides fully executable, copy-paste-ready templates: four scored gold-standard answers per question with explicit 'Why score N' rationale tied to rubric anchors, a pitfall table with concrete corrections, a timed calibration-session script (90 min, ~10 min/question), and a structured output document with a numbered hand-off block - matching the score-3 anchor for fully executable, specific examples.

3 / 3

Workflow Clarity

The five steps are clearly sequenced with explicit validation checkpoints - the INSUFFICIENT_TRANSCRIPTS halt below 2 transcripts, 'Do not schedule real candidates before this session is complete', and a retro after the first 5 candidates - plus feedback loops for anchor refinement, matching the score-3 anchor for clear sequence with explicit validation and error-recovery loops.

3 / 3

Progressive Disclosure

No bundle files exist to structure; the body is organized into clear sections (Overview, When to use, Steps 1-5, Anti-patterns, Limitations, References) with one-level-deep, clearly signaled external references (sibling skills, ISTQB glossary, research links) and no nested/2+ level reference chains, matching the score-3 anchor for well-organized sections and easy navigation.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and clearly distinctive: it names concrete inputs and outputs, provides an explicit use-when trigger, and explicitly differentiates the skill from its two sibling skills. It uses third-person voice throughout with no first/second-person violations.

DimensionReasoningScore

Specificity

Lists multiple concrete actions - 'takes the question bank and rubric plus 2-5 sample candidate transcripts' and 'emits gold-standard model answers, common pitfalls, score-anchor examples per question, and a calibration-session script' - matching the score-3 anchor for multiple specific concrete actions.

3 / 3

Completeness

Explicitly answers both what (produces a calibration guide with gold-standard answers, pitfalls, score-anchors, session script) and when ('Use after the rubric exists and before the first real candidate'), matching the score-3 anchor for explicit triggers on both.

3 / 3

Trigger Term Quality

Covers natural domain terms a QA hiring manager would actually say - 'interviewer calibration', 'calibration guide', 'inter-rater reliability', 'hiring loop' - rather than abstract jargon, matching the score-3 anchor.

3 / 3

Distinctiveness Conflict Risk

Explicitly distinguishes itself from sibling skills - 'Distinct from the question and rubric skills... this skill produces the demonstration material' - establishing a clear niche unlikely to conflict, matching the score-3 anchor.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents