github.com/Cloudzero/cloudzero-claude-marketplace
| Skill | Added | Review |
|---|---|---|
cost-anomaly-detection plugins/cost-analyst/skills/cost-anomaly-detection/SKILL.md Use when proactively scanning for cost anomalies, unusual spending, unexpected charges, or irregular patterns — during weekly reviews, after incidents, or when something looks off | 75 75 1.25x Agent success vs baseline Impact 100% 1.25xAverage score across 3 eval scenarios Securityby Passed No findings from the security scan Reviewed: Version: f539a8b | |
cost-comparison plugins/cost-analyst/skills/cost-comparison/SKILL.md Use when comparing costs between time periods, environments, accounts, regions, or teams to understand spending differences and identify inefficiencies | 57 57 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
cost-projection plugins/cost-analyst/skills/cost-projection/SKILL.md Project the monthly cost of an infrastructure definition (Terraform, CDK, CloudFormation, SAM) using CloudZero spend data. Reads IaC files, enumerates resources, and produces a line-item cost breakdown. | 60 60 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: f539a8b | |
cost-spike-investigation plugins/cost-analyst/skills/cost-spike-investigation/SKILL.md Use when a cost spike or unexpected increase has already been identified and you need to find which service, account, or resource is responsible | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
cost-trend-analysis plugins/cost-analyst/skills/cost-trend-analysis/SKILL.md Use when analyzing whether costs are growing, declining, or stable over time — for forecasting, budget planning, or understanding spending velocity | 57 57 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
custom-dimension-analysis plugins/cost-analyst/skills/custom-dimension-analysis/SKILL.md Use when analyzing costs by organization-specific dimensions like teams, products, business units, or applications for showback, chargeback, or business-aligned cost reporting | 55 55 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
diff-cost-projection plugins/cost-analyst/skills/diff-cost-projection/SKILL.md Analyze code diffs for infrastructure cost impact using CloudZero spend data. Detects Terraform, CDK, CloudFormation, SAM, K8s, scaling, and application code changes that affect cloud spending. | 67 67 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: f539a8b | |
model-right-sizer-audit plugins/model-right-sizer/skills/model-right-sizer-audit/SKILL.md One-shot, PER-CALL-SITE audit of a repo's EXISTING LLM calls — every place code invokes a model (an SDK/API call, a sub-agent dispatch, an agent frontmatter definition), decomposed by INTENT into the distinct jobs it does, never a grep hit on a model name and never one candidate per file. A skill pinned by one static `model:` key still gets its steps split by intent when severable under Claude Code's per-turn model binding. DELEGATES each call to `model-right-sizer-dryrun` (never re-implements its scoring) and merges results into ONE schema-conformant JSON blueprint committed at the TARGET REPO'S ROOT via a PR — never a markdown table. Distinct from `model-right-sizer-dryrun` (invoke directly for one hypothetical task) and `model-right-sizer-install` (the standing before/after mandate). Use when someone says "audit this repo's model calls", "right-size the models in <repo> per call site", "find every LLM call and right-size it", or "commit a model right-sizing blueprint for <repo>". | — | |
model-right-sizer-budget-guard plugins/model-right-sizer/skills/model-right-sizer-budget-guard/SKILL.md The while-work-is-in-flight companion to `model-right-sizer-dryrun`: once `work_routing_map[]` is real and being dispatched as sub-agents, the runbook for keeping two things honest — the status ledger (`status`/`status_updated_at`/`status_note` per row, flipped at every real transition) and the token-budget guard (checking real spend against `budget.token_ceiling`; once `budget_threshold.py`'s `threshold_crossed()` trips at `warning_threshold_pct`, sending `format_budget_warning()`'s string into that unit's next turn). Checks at turn boundaries with whatever usage the dispatch mechanism reports — never a fabricated live ticker. Only applies to rows actually dispatched, never design-time-only `blueprint_rows[]`, and never replaces Pass B's usage report. Use when someone says "dispatch the work-routing map", "run the budget guard while units are in flight", "update the status ledger for unit X", or "did unit X cross its warning threshold". | 70 70 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-dryrun plugins/model-right-sizer/skills/model-right-sizer-dryrun/SKILL.md Preview the model-routing MAP for an intent without building anything. You give it a free-text INTENT — "build a CLI that…", "add a feature that…", "refactor X into Y" — and it invokes the `model-right-sizer` agent in BLUEPRINT-ONLY mode: decompose the work, score each piece, and emit a single schema-conformant JSON blueprint (task→model→effort→budget→ schema→confidence, per `schemas/blueprint.schema.json`), then STOP. No build, no file edits, no after-the-fact usage report — it is the what-would-this-cost / how-would-this-route preview lever, safe to run against any idea, and the JSON it returns is what an orchestrator parses to route dispatch. Read-only. Use when someone says "dry-run the right-sizer on …", "what's the map for …", "how would you route …", or "show me the blueprint for … before I build it". | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-holdout-tuning plugins/model-right-sizer/skills/model-right-sizer-holdout-tuning/SKILL.md Tune model-right-sizer's wording knobs (`eval/tuning/knobs.py`) against a REAL, already-measured build's actual token spend — not the synthetic benchmark `model-right-sizer-prompt-tuning` searches, nor a fresh build per candidate. Picks a task from `overfitting_guard.py`'s `HOLDOUT_TASKS` registry (real actual/budgeted pairs already recorded), dispatches 3 INDEPENDENT BLIND dry-runs per candidate (no calibration-ledger access), averages the budget across draws (single-draw noise can flip within/over-budget calls), maps to real actuals, scores via `classify_budget_adherence` + `score_candidate`, diagnoses the miss pattern, proposes ONE wording change, and re-runs to check improvement. n stays fixed per task — flags rather than silently continues once squeezing looks like overfitting. Use when someone says "tune the knobs against this blueprint/build", "iterate the dry run with no prior context against the real actuals", "keep tuning until N%", or "how close does a blind estimate get to the actual cost". | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-install plugins/model-right-sizer/skills/model-right-sizer-install/SKILL.md Stamp a model-right-sizer mandate onto the CURRENT repo's agent-instructions file — `CLAUDE.md`, `AGENTS.md`, or both, whichever the repo has. Idempotent and append-only: an existing file is never overwritten, only the marker-delimited mandate block is inserted or refreshed. The mandate's "before" hook runs `model-right-sizer-dryrun` to produce a schema-conformant JSON blueprint for the orchestrator to route by. Also checks whether the `model-right-sizer` agent and `model-right-sizer-dryrun` skill are discoverable and, if either is missing, installs the `model-right-sizer` Claude Code plugin to fix that (falling back to manual instructions otherwise). Deliberately narrow and organization-agnostic — installs only the right-sizing mandate (plus its own dependencies, if absent), not any broader development process. Use when someone says "install model-right-sizer in this repo", "init this repo for model-right-sizer", "add the right-sizer mandate here", or "make this repo consult model-right-sizer every turn". | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-layer-ablation plugins/model-right-sizer/skills/model-right-sizer-layer-ablation/SKILL.md Empirically measure what each of model-right-sizer's four research-grounded citation layers (Token Economics, IBPO, BudgetThinker, Speculative Decoding) actually does to its blueprints — instead of trusting the citations alone. Renders layer-ablated variants (any of the 16 layer subsets), runs a fixed six-task benchmark through each variant's Pass A blueprint, and for a scoped subset actually executes the recommended build and scores whether real effort stayed within the predicted budget (wrapping `classify_budget_adherence`). Reports each layer's effect in ISOLATION vs. a zero-layer baseline, and every COMBINATION across the full 16-subset grid, so synergy or redundancy is visible, not assumed away. Read-mostly: writes only a scratch directory and a final report, never `agents/model-right-sizer.md`. Use when someone says "does the Token Economics layer actually change anything", "ablate the research layers", "run the layer-ablation study", or "audit model-right-sizer's citations empirically". | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-prompt-tuning plugins/model-right-sizer/skills/model-right-sizer-prompt-tuning/SKILL.md Tune the exact WORDING of model-right-sizer's four already-shipped research-grounded layers to maximize real-execution accuracy (effort stayed within budget, per `classify_budget_adherence`) — not whether to include a layer (see `model-right-sizer-layer-ablation`), but how an included layer should be phrased. Runs a discrete coordinate-ascent search (the finite-difference analog of gradient descent for prose) over four wording knobs in `eval/tuning/knobs.py`, each anchored at one spot in the shipped agent text plausibly moving the accuracy ratio: `token_ceiling` margin, how hard the effort dial leans down under difficulty-uncertainty, and the calibration/adherence knobs. Read-mostly: never edits the agent file directly, only proposes the winning wording as a diff to review. Use when someone says "tune model-right-sizer's wording for accuracy", "optimize the budget-ceiling wording", "run a gradient descent / hill-climbing search on the agent prompt", or "which wording maximizes budget-adherence accuracy". | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-release-report plugins/model-right-sizer/skills/model-right-sizer-release-report/SKILL.md Publish a per-version release report for `eval/token_ceiling_formula.py` every time `FORMULA_VERSION` bumps — the exact configuration (formula, signal weights, calibration constants), a ranked gap list for the next contributor, and what's settled and not worth re-relitigating. Why this is its own skill rather than folding into `model-right-sizer-research-report`: every claim must interweave WHY it matters, in the same breath as WHAT changed — it stays decision-support only if a reader can tell, without a second file, what breaks (wasted spend, false alarms, undetected overruns, a re-biased fleet of budgets) if a number or gap is wrong. Never a new-finding surface — synthesizes only from committed results files. Also the tool for BACKFILLING a report for a past version that shipped before this skill existed. Use when someone says "write the release report for this version", "version the token ceiling formula", "backfill a release report for v0.x", or after any `FORMULA_VERSION` bump. | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-research-report plugins/model-right-sizer/skills/model-right-sizer-research-report/SKILL.md Package up every result this plugin's tuning/validation research has produced — the layer-ablation study, the prompt-tuning coordinate-ascent passes and the dispatch-floor-awareness/held-out-task work, the averaged-vs-additive `token_ceiling_formula.py` pivot, and the real-work-signal validation experiments — into one condensed, research-paper-style EXECUTIVE report with real charts, built entirely from numbers already recorded in this repo's own dated results files (never invented or rounded up). Publishes a self-contained HTML report (loads the `dataviz` and `artifact-design` skills first) with an abstract, a key-findings table, a handful of figures, limitations stated as prominently as wins, and a reproducibility appendix pointing at the companion skills that can re-run each experiment. Use when someone says "write up all the tuning results", "executive summary of the research", "package the findings into a report", "research report with charts", or "summarize everything we've found so far for leadership". | 66 66 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
model-right-sizer-schema plugins/model-right-sizer/skills/model-right-sizer-schema/SKILL.md Prescribe a minimal output schema for ONE agent's handoff to its controller — the `model-right-sizer` agent's "Agent-to-agent message-schema design" lever, scoped to a single seam instead of a whole flow's blueprint. Given a target agent (a path to an existing agent `.md` file, or a description of one not yet written), returns a schema-conformant JSON prescription (`schemas/agent-schema.schema.json`) naming the reusable family the agent's reply fits, typed `in`/`out` fields, an exclusion list, and a ready-to-insert `## Agent-to-agent schema` markdown stamp — reproduced here in portable, organization-agnostic form. Offers to stamp that block directly into the target agent's file, idempotently. The point: an agent that used to hand its controller unscoped prose now hands it typed fields plus one bounded prose slot. Use when someone says "give this agent an output schema", "prescribe a schema for …", "minimize what this agent returns", or "stamp an agent-to-agent contract on …". | 64 64 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: f539a8b | |
model-right-sizer-signal-validation plugins/model-right-sizer/skills/model-right-sizer-signal-validation/SKILL.md Test whether a candidate real-work signal in `eval/token_ceiling_formula.py` (e.g. `context_ingestion_volume`, `investigative_uncertainty`, or new) deserves a nonzero default weight, via the blind multi-draw rating + correlation methodology this repo's research converged on. Ratings must come from genuinely independent sub-agent dispatches seeing ONLY a forward-looking task spec and the signal definitions — never a context holding the real actuals or this repo's retired write-ups, which would turn "blind rating" into transcribing the answer key. Dispatches 3+ independent draws per held-out task, computes per-signal CV and Pearson correlation (alone, and added to the signal sum — dilution, not weak correlation, is the dominant failure found twice), requiring replication on a SECOND held-out task before proposing a nonzero weight. Use when someone says "test this new signal", "does [signal] deserve a nonzero weight", "re-run the signal validation experiment", or "validate real-work signals against real data". | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b | |
optimize-triage plugins/cost-analyst/skills/optimize-triage/SKILL.md Fetch top unaddressed CloudZero Optimize recommendations, dispatch parallel research agents per item, apply SRE critique, and surface actionable findings with confidence verdicts and per-resource report files. Read-only research only. | 67 67 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: f539a8b | |
release .claude/skills/release/SKILL.md Cut a release of this repository (the CloudZero plugin marketplace itself — not a customer-facing plugin). Promotes every touched plugin's own `Unreleased` changelog section to a dated entry, synthesizes those changes into a new SemVer section at the top of the root `CHANGELOG.md`, bumps `.claude-plugin/marketplace.json`'s version, runs the full CI validation suite locally, tags the release, publishes a GitHub Release, and opens the companion docs PR against `Cloudzero/cloudzero-documentation`'s `v2.0` branch so customer-facing docs land in step with the code. This is a maintainer-only, repo-local skill (lives in `.claude/skills/`, not inside a `plugins/*/skills/` directory) — it is never installed by marketplace consumers. Use when told "cut a release", "release the marketplace", "ship a new version", "promote the changelog", or after merging PR(s) that need a version bump and a docs update — mirroring how v1.2.0 (the Model Right Sizer launch) shipped. | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: f539a8b |