CtrlK
BlogDocsLog inGet started
Tessl Logo

calibrate-room

Run the ADR-151 per-room calibration pipeline — baseline → enroll → extract → train → a bank of small specialists (presence/posture/breathing/heartbeat/restlessness/anomaly).

62

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./harness/ruview/.claude/skills/calibrate-room/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a model of token efficiency: a terse, executable 4-step pipeline with an honesty/verification section that guards against over-claiming accuracy. Its only weaknesses are minor missing operational details (capture durations, enroll activity specifics, claim_check arguments) that keep actionability and workflow clarity just short of full marks.

DimensionReasoningScore

Conciseness

The ~30-line body is lean and assumes competence throughout — terse imperatives like 'Leave the room empty' and dense facts like 'Pure-Rust, edge-deployable (ADR-151)' — with no padding or explanation of concepts Claude already knows. Every token earns its place.

5 / 5

Actionability

Every step carries an executable command ('ruview_calibrate {step: "baseline"}') plus an invocation fallback ('installed wifi-densepose binary, else cargo run -p wifi-densepose-cli'). Not 5: minor gaps remain — no durations or parameters for the baseline capture or enrollment activities, and ruview_claim_check is referenced without arguments.

4 / 5

Workflow Clarity

A clear numbered 4-step sequence with commands, plus a final verification checkpoint ('tag presence/vitals accuracy MEASURED only with a held-out check — run ruview_claim_check'). Not 5: no intermediate validation between steps (e.g., confirming baseline quality before enrolling); the destructive/batch cap at 3 does not apply since operations are non-destructive.

4 / 5

Progressive Disclosure

Under 50 lines with no need for external references, and the rubric's simple-skill exception applies: well-organized sections (Sequence, Honesty) alone warrant a 5. The only pointer (to the 'room-watch' skill) is clearly signaled at one level deep.

5 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a concrete, well-scoped capability with a named step sequence and specific output types, but it entirely lacks 'when to use' trigger guidance and leans on internal jargon rather than natural user phrasing. Adding a 'Use when...' clause with user-sayable synonyms would raise both completeness and trigger term quality.

Suggestions

Append an explicit trigger clause, e.g. 'Use when the user asks to calibrate a room, set up room sensing, or mentions ADR-151 or the ruview_calibrate tool.'

Replace or gloss internal jargon ('a bank of small specialists') with natural synonyms users would say (e.g. 'train small per-room models') to improve trigger term quality.

Reconcile the step list with the body: the description lists 'extract' but the body's sequence is baseline → enroll → train-room → room-watch, which creates a small specificity gap.

DimensionReasoningScore

Specificity

Names multiple concrete pipeline actions ("baseline → enroll → extract → train") and concrete outputs ("presence/posture/breathing/heartbeat/restlessness/anomaly"), matching the 'several specific actions; minor gaps' anchor. Not 5 because coverage is compressed and the listed 'extract' step does not correspond to any step in the body.

4 / 5

Completeness

The 'what' is clear and specific (run the ADR-151 per-room calibration pipeline with named steps), but there is no 'Use when...' or equivalent explicit trigger guidance, which the judging guidelines cap at 3. Not 2 because the 'what' is concrete rather than vague.

3 / 5

Trigger Term Quality

Relevant keywords exist ("calibration", "per-room", "presence", "breathing", "heartbeat") but internal jargon ("ADR-151", "a bank of small specialists") crowds out natural user phrasings like 'calibrate the room' or 'set up room sensing'. Fits 'some relevant keywords but missing common variations or synonyms'; the good-coverage (4) anchor expects more natural terms users would actually say.

3 / 5

Distinctiveness Conflict Risk

A niche domain ('ADR-151 per-room calibration') with distinct triggers gives minimal conflict risk; minor overlap with the sibling 'room-watch' skill referenced in the body, and the missing 'when' clause leaves triggering somewhat broad — fits 4 better than 5.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ruvnet/RuView
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.