CtrlK
BlogDocsLog inGet started
Tessl Logo

homelab-sre-agent

Applies Site Reliability Engineering practices to homelab operations including SLOs, toil reduction, and incident management.

56

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/homelab-sre-agent/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, token-efficient overview that correctly assumes SRE knowledge, with a clear incident lifecycle sequence. It falls short on actionability and workflow validation: no templates, commands, or examples make the responsibilities executable, and the incident workflow lacks checkpoints.

Suggestions

Add concrete artifacts — an SLO definition format and a blameless post-mortem template (natural fits for the empty references/ directory) — so directives like "Track SLO burn rate" and "Write blameless post-incident review" become executable.

Insert validation checkpoints into the incident workflow, e.g. "Verify the service is actually restored after mitigation before starting root-cause analysis" and a review step confirming action items are tracked to closure.

Trim the duplicated opening paragraph and the Invocation meta-section to tighten token efficiency.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence — it never explains what an SLO, toil, or post-mortem is, and uses terse bullets like "Track SLO burn rate and alert budget consumption". It is not 5 because the opening paragraph ("Brings SRE discipline to homelab operations... post-incident reviews") restates the frontmatter description, and the Invocation section is harness meta-information; not 3 because there is no padded or unnecessary explanation beyond those minor trims.

4 / 5

Actionability

There is some concrete guidance — the numbered incident lifecycle ("Detect", "Triage", "Mitigate", "Root cause", "Post-mortem", "Action items") and named automation tooling "(scripts, cron jobs, Ansible tasks)" — but the remaining directives are missing key executable details: "Track SLO burn rate" names no tool or method, "Write blameless post-incident review" provides no template, and "Track toil-hours saved" gives no mechanism. It is not 4 because no commands, examples, or templates make the guidance executable; not 2 because the lifecycle sequence and tooling pointers exceed high-level hints.

3 / 5

Workflow Clarity

The incident lifecycle is presented as a clear six-step numbered sequence (Detect → Triage → Mitigate → Root cause → Post-mortem → Action items), and the other sections are coherent. However, validation checkpoints are absent — e.g., nothing verifies the service is actually restored after "Mitigate" before root-cause work begins, and action items have no follow-through check. It is not 4 because checkpoints are missing entirely rather than a minor gap; not 2 because the sequence itself is well defined with few gaps.

3 / 5

Progressive Disclosure

The body is well under 50 lines with clean section organization, and the single external reference (Google SRE Book URL under "## References") is one level deep and clearly signaled; the empty references/ and scripts/ directories mean no buried or nested navigation. It is not 5 because there is no in-bundle reference structure at all — natural offload candidates like a post-mortem template or SLO format examples are absent rather than split out — and not 3 because the content present is appropriately placed and easy to navigate for this length.

4 / 5

Total

14

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, third-person description that names a clear niche and three concrete capability areas. Its main weakness is the missing "Use when..." trigger clause, which caps completeness, and its reliance on a single generic verb rather than per-area actions.

Suggestions

Add an explicit trigger clause, e.g. "Use when managing homelab reliability, defining SLOs or error budgets, reducing toil, or running incident post-mortems" — completeness is capped at 3 without one.

State a concrete action per capability area (e.g., "Define and track SLOs with burn-rate alerting, automate repetitive toil, run blameless post-incident reviews") instead of the generic "Applies ... practices".

Include user-natural synonyms such as "uptime", "error budget", and "postmortem" to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

The description enumerates three concrete capability areas — "SLOs, toil reduction, and incident management" — within a clearly named domain ("Site Reliability Engineering practices to homelab operations"), leaving only a minor coverage gap (the body's reliability reviews are unmentioned). It is not 5 because it relies on a single generic verb ("Applies") rather than stating what it does in each area, and not 3 because it goes beyond naming 1-2 actions with comprehensible, specific coverage.

4 / 5

Completeness

The "what" is clear (applies SRE practices: SLOs, toil reduction, incident management) but there is no "Use when..." clause or any equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. It is not 4 because the "when" is entirely absent rather than merely under-specified.

3 / 5

Trigger Term Quality

It includes several natural terms a homelab user would actually say — "Site Reliability Engineering", "homelab operations", "SLOs", "toil", "incident management". It is not 5 because common synonyms and variations are missing (e.g., "reliability", "uptime", "error budget", "post-mortem"), and not 3 because keyword coverage is genuinely good rather than partial.

4 / 5

Distinctiveness Conflict Risk

"Site Reliability Engineering practices to homelab operations" carves out a clear niche with domain-specific triggers unlikely to fire for unrelated skills. It is not 5 because "incident management" overlaps the detection/diagnosis scope of the companion homelab-monitoring-analyst skill, and not 3 because the domain scoping is far more specific than a generic overlap-prone description.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
pvnkmnk/AgenticSelfHostSkills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.