CtrlK
BlogDocsLog inGet started
Tessl Logo

model-right-sizer-budget-guard

The while-work-is-in-flight companion to `model-right-sizer-dryrun`: once `work_routing_map[]` is real and being dispatched as sub-agents, the runbook for keeping two things honest — the status ledger (`status`/`status_updated_at`/`status_note` per row, flipped at every real transition) and the token-budget guard (checking real spend against `budget.token_ceiling`; once `budget_threshold.py`'s `threshold_crossed()` trips at `warning_threshold_pct`, sending `format_budget_warning()`'s string into that unit's next turn). Checks at turn boundaries with whatever usage the dispatch mechanism reports — never a fabricated live ticker. Only applies to rows actually dispatched, never design-time-only `blueprint_rows[]`, and never replaces Pass B's usage report. Use when someone says "dispatch the work-routing map", "run the budget guard while units are in flight", "update the status ledger for unit X", or "did unit X cross its warning threshold".

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually rigorous operational runbook: exact function calls, explicit state-transition rules, a fully worked example, and honest epistemic constraints (never fabricate timestamps or spend). Its one real weakness is verbosity — several disciplines are stated three or four times across steps and the NOT-do section, and step 5 carries author's-log padding that could be cut roughly in half without losing any constraint.

Suggestions

State the fresh-timestamp rule once (e.g., in step 2) and reference it from steps 3–4 instead of restating it each time; the same applies to the never-fabricate rule, which currently appears in steps 2, 5, 7, and the NOT-do section.

Compress step 5's meta-narrative about the author checking the tool surface to a single-sentence constraint: 'Check spend at turn boundaries — no verified mid-turn usage signal exists; check more often only if your runtime positively confirms one.'

Move the full worked example into a references/ file (e.g., references/worked-example.md) and keep a two-line summary in the body, shortening SKILL.md toward the lean overview pattern.

DimensionReasoningScore

Conciseness

The content is operationally real but padded: the freshly-read-timestamp discipline is restated in steps 2, 3, and 4 ('the same discipline as steps 2–3, on every transition, no exceptions'), the never-fabricate rule appears in steps 2, 5, 7, and again under 'What this does NOT do', and step 5's author's-log narrative ('this skill's own author checked... Finding: no such signal was found') is unnecessary meta-commentary. It fits 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than the noticeably-verbose anchor, since no section explains concepts Claude already knows.

3 / 5

Actionability

Every step is concrete down to exact call signatures — 'budget_threshold.threshold_crossed(actual_tokens, budget.token_ceiling, budget.warning_threshold_pct or 0.7)' — and the worked example shows real JSON row states, computed percentages, and the verbatim warning string to send. As an instruction-only skill with fully actionable guidance (per the scoring notes), this matches the top anchor.

5 / 5

Workflow Clarity

Seven numbered steps in dispatch order, each with an explicit transition condition and validation checkpoint (check real spend after each dispatch returns), a genuine feedback loop (inject the warning verbatim into the next turn), and explicit error paths ('blocked' requires a concrete note; unknown spend stays unknown). This is the clear-sequence-with-explicit-validation-and-feedback-loops anchor.

5 / 5

Progressive Disclosure

Clean section structure (When this runs / What to do / Worked example / What this does NOT do / Related) with a well-signaled, one-level-deep Related section pointing to `eval/budget_threshold.py`, the sibling dryrun skill, the blueprint schema, and the agent file. No bundle files exist in this skill, and the ~240-line body inlines the full worked example where a references/ file could carry it — good structure with minor organization gaps rather than the ideal split.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense but genuinely specific description: concrete mechanics with exact function names, an explicit Use-when clause with natural trigger phrases, and explicit boundaries against sibling skills. The only gap is that all trigger terms are plugin-internal jargon with few synonyms a less-initiated user might say.

Suggestions

Add one or two plain-language trigger synonyms (e.g., 'keep an eye on token spend while sub-agents run', 'token budget check') alongside the plugin-vocabulary phrases.

Consider trimming the parenthetical mechanics (function names, field lists) slightly so the description reads faster during skill selection.

DimensionReasoningScore

Specificity

Both duties are enumerated down to exact field and function names — 'the status ledger (`status`/`status_updated_at`/`status_note` per row, flipped at every real transition)' and 'sending `format_budget_warning()`'s string into that unit's next turn' — matching the comprehensive-coverage anchor, not the minor-gaps one.

5 / 5

Completeness

It explicitly answers what (the runbook for the status ledger and the token-budget guard, with concrete mechanics) and when ('Use when someone says "dispatch the work-routing map", "run the budget guard while units are in flight"...') with concrete trigger phrases — the exact shape of the top anchor.

5 / 5

Trigger Term Quality

Quoted natural phrases ('dispatch the work-routing map', 'did unit X cross its warning threshold', 'update the status ledger for unit X') give good coverage, but every term is bound to this plugin's vocabulary and common user phrasings like 'token budget', 'spend guard', or 'cost while agents run' are missing, so it falls short of the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

It draws explicit boundaries against every neighbor — 'the while-work-is-in-flight companion to `model-right-sizer-dryrun`', 'never design-time-only `blueprint_rows[]`', 'never replaces Pass B's usage report' — and its trigger phrases are unique to this skill, so conflict risk is minimal.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 9 suspicious

Warning

Total

14

/

16

Passed

Repository
Cloudzero/cloudzero-claude-marketplace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.