CtrlK
BlogDocsLog inGet started
Tessl Logo

ce-dogfood

Hands-off, diff-scoped browser QA of the active branch: maps user flows, drives a real browser, autonomously fixes small breakages with regression tests and commits, judges experience against product personas, and writes a durable dogfood report. Manual invocation only.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/ce-dogfood/SKILL.md

The canonical home for this skill is ce-dogfood in EveryInc/compound-engineering-plugin

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured orchestrator skill: the body carries only policy and boundaries, delegates phase mechanics to a clearly signaled reference file, and defines a complete workflow with validation checkpoints and error-recovery feedback loops. Its only weakness is minor redundancy in the 'read phases.md' and branch-slug instructions.

DimensionReasoningScore

Conciseness

The body is dense with environment-specific, non-obvious constraints (agent-browser CLI choice, docs_root resolution, branch-slug rules, terminal states) and explains nothing Claude already knows. Not 5 because of redundancy: the instruction to read 'references/phases.md' before Phase 0 appears nearly verbatim at both line 16 and line 62, and the branch-slug rule is stated in both SKILL.md and phases.md.

4 / 5

Actionability

Concrete, executable guidance throughout: 'command -v agent-browser >/dev/null 2>&1', 'mktemp -d "${TMPDIR:-/tmp}/ce-dogfood-XXXXXX"', 'git rev-parse --show-toplevel', and explicit phase order. Not 5 because the actual browser-driving and per-phase mechanics are deferred to references/phases.md, leaving minor execution gaps in the top-level file itself.

4 / 5

Workflow Clarity

The sequence 'Scope -> analyze the diff -> map the flows -> derive the matrix -> serve -> execute -> fix loop -> report' is stated as an invariant, with an explicit feedback loop ('A fix is not done until a regression test fails before it and passes after'), incremental validation checkpoints (create the report as soon as the matrix exists, update after each scenario/fix), and defined terminal states with resume behavior. This matches the clear-sequence-with-explicit-validation-and-feedback-loops anchor.

5 / 5

Progressive Disclosure

SKILL.md is a genuine overview: all phase detail is split into references/phases.md, which is strongly signaled ('Read references/phases.md before Phase 0 (Scope) and follow it'), and the bundle structure confirms one-level-deep navigation — phases.md reaches references/test-matrix-taxonomy.md, references/dogfood-report-template.md, and scripts/packs-resolve.py, all of which exist, with no deeper nesting. Not 4 because the split is clean and every reference path resolves to a real file.

5 / 5

Total

18

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A highly specific, distinctive description that comprehensively states what the skill does in natural terms. Its one real gap is the missing 'Use when...' trigger clause — 'Manual invocation only' constrains how it is invoked but never says when a user should reach for it.

Suggestions

Add an explicit trigger clause such as 'Use when the user asks to dogfood, QA, or end-to-end test the current branch or a PR in a real browser.'

Include natural synonym phrasings users would say (e.g., 'smoke test', 'test the branch', 'end-to-end QA') to broaden trigger coverage.

Consider stating the diff-scoped trigger more explicitly (e.g., 'use when the user wants only the branch's changes tested, not the whole app').

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'maps user flows, drives a real browser, autonomously fixes small breakages with regression tests and commits, judges experience against product personas, and writes a durable dogfood report' — with comprehensive coverage of the skill's capabilities. Not 4 because the coverage has no minor gaps: every listed action is specific and distinct.

5 / 5

Completeness

The 'what' is clear and concrete, but there is no 'Use when...' clause or equivalent explicit trigger guidance; 'Manual invocation only' states an invocation constraint, not when the skill applies, which caps completeness at 3 per the judging guideline. Not 4 because the 'when' is only weakly implied by 'of the active branch', and not 2 because the 'what' is fully explicit.

3 / 5

Trigger Term Quality

Good natural keywords — 'browser QA', 'dogfood report', 'regression tests', 'active branch' — phrased the way a user would ask for this skill. Not 5 because common synonyms and variations ('test the branch', 'smoke test', 'end-to-end test') are absent, leaving a few natural terms uncovered.

4 / 5

Distinctiveness Conflict Risk

'Hands-off, diff-scoped browser QA of the active branch' carves out a clear niche with distinct triggers (dogfooding, browser QA of a branch diff) that no generic test/commit skill would claim. Not 4 because overlap risk with neighboring skills is minimal — the persona-judging and durable-report framing is unique.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
crdant/compound-engineering-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.