CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-skills

github.com/mthines/agent-skills

SkillAddedReview
e2e-pr-stabilizer

skills/testing/e2e-pr-stabilizer/SKILL.md

Stabilizes or optimizes Playwright E2E tests on one PR via a local-first loop, then ratifies with a single CI run. Pulls Dash0 spans (`git.pull_request_link`) as the historical baseline, then captures every iteration's evidence locally with `--trace=on` (same OTel exporter, same trace schema). Validation is empirical, not predictive: before commit, every new locator must resolve against source (static grep) or the live app (`locator.count()`); after commit, the fixed test must pass three consecutive local runs before the single push. Modes: `stabilize` (default) heals flaky / failing tests; `optimize` is report-only and ranks slow-action wins by measured ms saved. Refuses `.skip`, `.fixme`, `waitForTimeout`, or any check-weakening edit. Use when a PR has flaky or failing E2E tests or when you want to find slow tests worth tightening. Triggers on "stabilize this PR", "fix flaky e2e", "heal playwright on PR", "ui-e2e is failing", "self-heal e2e", "optimize e2e", "/e2e-pr-stabilizer".

64

e2e-testing

skills/testing/e2e-testing/SKILL.md

Plans, generates, runs, and heals end-to-end tests using Playwright Test Agents (Planner, Generator, Healer) and the official `@playwright/mcp` server. Drives a spec-first feature-flow loop, proposes `data-testid` source diffs only when accessibility-tree locators fail, and stays token-aware via snapshot mode and `--last-failed` reruns. Use when adding E2E coverage, verifying a user journey, hardening a flaky flow, or wiring Playwright MCP into a repo. Triggers on "test this flow", "add e2e", "verify the user journey", "write e2e test", "feature test", "playwright agents", "/e2e-testing".

72

e2e-testing-mobile

skills/testing/e2e-testing-mobile/SKILL.md

Plans, generates, runs, and heals end-to-end tests for Expo and React Native mobile apps using Maestro (the 2026 standard for RN E2E, adopted by Meta, Microsoft, and DoorDash, and integrated with Expo via EAS Workflows). Drives a spec-first YAML-flow loop, proposes `testID` source diffs (never `accessibilityLabel` reuse), runs Maestro Cloud as an EAS job, and stays token-aware via `--shards`, `--retries`, and failure-only healing. Use for native flows in Expo / RN apps. Triggers on "test this RN flow", "add mobile e2e", "maestro flow", "expo e2e", "e2e for react native", "test the native app", "/e2e-testing-mobile". Defer to [`e2e-testing`](../e2e-testing/SKILL.md) for web flows and the WebView portion of hybrid apps.

69

eval-iterate

skills/quality/eval-iterate/SKILL.md

Iterates on a failing AI/LLM eval (an L2 suite, a golden-set / judge eval, or any eval gating a PR) until it is green AND confirmed, not just luckily passing once. Classifies the failure as a code bug, an eval-definition bug (stale golden item / criteria drift), judge-drift (silent grader version change), or flaky. Applies the minimal fix, then requires N consecutive confirming re-runs — 2 deterministic, 5 for anything a model call grades, since a flat "2" is statistically weak for a stochastic judge. Hard-capped at 5 fix iterations. Refuses to game the eval (Goodhart's Law: skip/delete/overwrite-in-place a case, loosen a threshold) without a second independent check plus `confidence(analysis) >= 90%` and a logged rationale. Composes `ai-engineering`, `confidence`, `critical`, `verify-behavior`. Use when an eval is failing on a PR and needs a real, non-gamed green. Triggers on "this eval is failing", "iterate on this eval", "fix this eval", "get this eval green", "optimize this eval", "/eval-iterate".

72

fix-bug

skills/workflow/fix-bug/SKILL.md

Resolves a single bug from any starting evidence — Dash0 telemetry (span, log, web event, RUM error link), a stack trace, an error message, a code pointer (file:line), a screen recording, a Linear ticket URL, or a free-text symptom. Classifies the input, triages complexity (Phase 0.5) to pick a fast lane or a full holistic-analysis lane, runs a pre-flight sweep, locks a failing reproduction (delegating to /tdd, /e2e-testing, or /e2e-testing-mobile by layer), delegates root-cause analysis to the isolated rca-investigator agent on complex bugs, gates on confidence(analysis), and at >= 92 % hands off without human confirmation — fast lane via aw-create-plan + aw-executor, standard lane via aw-planner + aw-executor, both under a CEGIS refinement contract. A bug-fix-verifier agent grades the PR before undrafting; for telemetry-sourced bugs an optional Phase 8 polls the originating signal post-deploy. --analyse-only stops at the proposal; --force-holistic skips the fast lane. Triggers on "/fix-bug".

65

github-actions-author

skills/delivery/github-actions-author/SKILL.md

Authors fast, cheap, maintainable GitHub Actions workflows applying 2026 best practices: caching with `hashFiles` + `restore-keys`, parallelization via matrix + artifacts, reusability (composite actions for steps, reusable workflows for jobs), security (SHA-pinned actions, least-privilege `GITHUB_TOKEN`, concurrency), trackable errors (named steps, step summaries, annotations, and stdout/stderr that always reaches the run log so agents can act on failures), and feedback for comment-triggered runs (👀 acknowledgement reaction on start, 🚀/👎 outcome reaction plus a run-linked comment at the end). Two modes: `scaffold` (default) generates workflow YAML; `review` audits an existing workflow against the same rules. Use when creating CI/CD pipelines, optimizing slow workflows, deduping copy-pasted YAML across repos, or auditing workflow security. Triggers on "github action", "github workflow", "ci pipeline", "create workflow", "speed up ci", "review my workflow", "/github-actions-author".

67

handoff

skills/authoring/handoff/SKILL.md

Distills the current session's task, context, and plan into a clean, copy-pasteable markdown handoff document for another agent or LLM to pick up. Captures the goal, current state, remaining plan, and the key files, decisions, and gotchas — and leaves out conversational noise, tool-call transcripts, and dead ends. Writes to `.agent/{branch}/handoff.md` and copies it to the clipboard. Use when continuing work in a fresh session, passing a task to a teammate's agent, or briefing a different model. Invoke with /handoff; add a focus phrase to scope the handoff, or --brief for a condensed version.

68

holistic-analysis

skills/analysis/holistic-analysis/SKILL.md

Forces a full step-back re-analysis when a fix or refactor is not working: traces the whole execution path end-to-end — entry point to exit, every block, every contract boundary, and the full data flow — instead of patching in isolation. Three modes: `fix` (default) for bugs and broken behavior, `refactor` for restructuring, and `review` for PR validation (returns intent-match and system-fit findings for the `pr-reviewer` agent to consume — never run standalone for routine review work). Trigger ONLY once at least one isolated fix attempt has already failed, or on an explicit request for a full execution-path analysis: "step back", "think holistically", "analyze the whole thing", "zoom out", "look at the bigger picture", and "rethink this" qualify in that context only. Never for trivial one-line fixes or first-attempt debugging. Also triggers on "/holistic", "/step-back", "/rethink", "/zoom-out".

67

ideate

skills/analysis/ideate/SKILL.md

Generates, stress-tests, and iteratively evolves ideas for a stated problem — product concepts, features, solution options, strategies — using research-grounded divergent/convergent agent loops: parallel persona generators (nominal-group simulation), independent judges scoring novelty, feasibility, impact, and fit on separate axes, and bounded recombination rounds gated by /confidence. Auto-triages run depth (quick in-context vs deep multi-agent; override with quick|deep). Use when brainstorming, exploring solution options, or pressure-testing a concept. Triggers on "brainstorm", "give me ideas", "help me come up with", "ideate on", "what could we build", "/ideate".

74

implement-suggestion

skills/workflow/implement-suggestion/SKILL.md

Implements review-comment suggestions across one or more PRs. Multi-PR mode (default when $ARGUMENTS holds PR URLs; empty $ARGUMENTS auto-detects the active PR) resolves a worktree per PR, fetches every actionable comment from human teammates AND AI review bots (claude[bot], coderabbitai[bot], …), validates each through /critical + /confidence, builds a structured suggestion-pack, and dispatches a worker subagent that applies each approved change as its own commit, pushes to the existing branch, and resolves the addressed thread — so every handled comment ends resolved and the PR is left clean. Free-text mode applies a single pasted suggestion in place. --watch re-applies on one PR until the reviewers go quiet (max 5 iterations); --resolve-all also replies to and resolves every non-fix thread it can honestly close. Triggers on "implement suggestion", "apply review comments", "/implement-suggestion".

66

interview

skills/analysis/interview/SKILL.md

Aligns on the scope of a request before any plan or implementation begins — the requirements-elicitation interview a senior engineer runs before touching code. Restates the request and diffs it against the user's words, researches the codebase so questions are specific, surfaces unknowns, non-goals, edge cases, and success criteria, then runs a batched clarifying-question interview only when research cannot resolve the ambiguity (adaptive — stays silent when the request is already crisp). Produces a confirmed brief.md artifact that downstream planning consumes, plus a readiness verdict (ready / ready-with-assumptions / blocked). Convergent and pre-plan: hands option-generation to /ideate, plan review to /critical, and scoring to /confidence. Use before autonomous work, before planning, or when a request feels underspecified. Triggers on "align on scope", "interview me", "clarify the request", "scope this", "scope alignment", "before we plan", "/interview".

66

jev-assert

skills/quality/jev-assert/SKILL.md

Verifies a UI expectation semantically by asking TypeSafe's Jev model whether a user-observable outcome holds in a page's captured TEXT state (accessibility tree / page text — never a screenshot), then maps Jev's typed answer and probability onto this repo's verify-behavior receipt vocabulary (confirms, contradicts, ambiguous, null, unobtainable). Driver-agnostic: consumes text state from either the Playwright (aw-tester) or the claude-in-chrome (aw-tester-chrome) runner, so the Chrome-vs-Playwright choice is orthogonal to it. Use it for a semantic UI assertion where an exact locator or string match is brittle — "did the user see a success state", "is this the right screen". Delegates the API contract to the typesafe@typesafe-ai skill (TYPESAFE_API_KEY). Triggers on "assert semantically", "does the page show", "verify this outcome with jev", "semantic UI assertion", "jev-assert", "/jev-assert".

66

measurable

skills/quality/measurable/SKILL.md

Ensures every delivery ships with the telemetry needed to prove its impact and catch its own regressions: RUM/analytics events for user-facing web changes (delegates to rum-tracking), OpenTelemetry traces, metrics, and structured logs for new or changed API endpoints, and explicit error/warning signal paths so failures surface instead of going silent. Knows OpenTelemetry Weaver — validating a signal against a semantic-convention registry, and treating a renamed metric or attribute as the breaking change it is. Modes: guide (default), implement, audit, setup. `setup` interviews the project once, recording its telemetry stack, schema (Weaver registry), per-package instrumentation approach for monorepos, and regression-detection expectations as a committed Observability Profile. Triggers on "is this measurable", "add telemetry", "instrument this endpoint", "check observability coverage", "will we know if this regresses", "does this break the telemetry contract", "/measurable".

68

observe-run

skills/quality/observe-run/SKILL.md

Runs a command and reads the telemetry that run just emitted, returning a `verify-behavior` receipt (confirms / contradicts / ambiguous / null) that grades a behavioral assertion — span count, parent/child structure, duration, error-path status, fan-out, attribute cardinality, ordering — against the observed spans, never against the diff read back. Inputs are a command plus an expectation set. Walks two cheapest-first rungs (an in-memory/file exporter, then `dash0 -X otlp proxy --agent-mode` plus `dash0 spans query`), stamps two-layer run identity, and self-skips when the repo has no Observability Profile dev target. Vendor-neutral: OTLP is the contract, Dash0 one implementation of the read. Use when a fan-out count, a retry that fired twice, an error swallowed into a 200, an N+1, or an unbounded cardinality needs proof from a real run rather than a reading of the code. Triggers only on an explicit ask — "observe this run's telemetry", "prove this behavior with a trace", "/observe-run".

64

optimize-approach

skills/quality/optimize-approach/SKILL.md

Reviews whether a change takes the most optimal approach for its stated intent and, when it does not, researches the code, validates a concretely better approach via holistic analysis, and either proposes it (report mode) or applies it behind a confidence gate (apply mode). Judges four axes — codebase-fit, simplicity, performance, robustness — at the approach level, deferring line-level and failure-mode findings to code-quality, critical, and holistic-review. Stays silent when the approach is already optimal (quiet early-exit). A `plan` mode reviews a drafted plan's approach at plan time (aw-planner Phase 1) — the cheapest moment to switch. Called by the pr-reviewer agent, the polish skill, and aw-planner as a default-on lens; also runnable standalone, where `--deep` also judges every `critical deep` lens's alternative (end-of-feature check). Triggers on "is this the best approach", "better way to do this", "is this optimal", "optimize this approach", "rethink the approach", "/optimize-approach".

67

optimize-claude-md

skills/authoring/optimize-claude-md/SKILL.md

Audits CLAUDE.md files (root, nested, `.claude/rules/*.md`) for context bloat and emits ranked suggestions across two levers — (1) shrink inventory entries, (2) flag rarely-used agent-invokable skills that should become slash-only to drop their description from the always-on available-skills list. Triggers on Claude Code's "Large CLAUDE.md will impact performance" warning (> 40k chars), inventory entries duplicating harness-loaded skill descriptions, "CLAUDE.md is too big", "shrink CLAUDE.md", "optimize CLAUDE.md", "/optimize-claude-md". Three modes — `audit` (read-only ranked report + slash-conversion candidates), `trim` (interactive one-line hook + diff approval), `extract` (moves sections to linked files preserving content). Composes with `docs` (Placement Resolver) and `create-skill` (invocation matrix). Hard rules: refuses files < 10k chars; never deletes silently; never edits any skill's canonical `SKILL.md` frontmatter — routes to `/create-skill`.

66

persistent-memory

skills/authoring/persistent-memory/SKILL.md

Persists context across conversations as plain markdown so every future session can enrich a topic-scoped memory (e.g. `project-acme`). Four operations: `write` (extract candidates, resolve as ADD / UPDATE / DELETE / NOOP per Mem0), `read` (load a ≤ 200-line INDEX, fetch detail on demand), `consolidate` (sleep-style merge + prune), `forget` (delete or redact with audit). Three storage tiers: home (`~/.agent-memory/<scope>/`, default), project-local (gitignored), project-shared (committed). Enforces a never-store list (secrets, keys, financial and identity numbers) and a consent preview before every write. `rules/scaling-tiers.md` covers scaling to SQLite FTS, vector DB, and managed memory, plus the LoreKit backend the self-improvement loops run on: scope mapping, the `loop::<skill>-lessons` tag and key convention, and the shared lesson schema. Triggers on "remember this", "save to memory", "recall memory", "what do you remember about", "consolidate memory", "forget that", "/persistent-memory".

67

playwright-trace-analyzer

skills/analysis/playwright-trace-analyzer/SKILL.md

Analyzes Playwright E2E `trace.zip` archives (and bare trace JSONL when unpacked). Extracts the action timeline, network waterfall, console errors, and DOM-snapshot anchors, then identifies the highest-impact problems (flaky waits, slow selectors, network bottlenecks, hung actions, unhandled console errors, navigation churn) and proposes concrete test or app fixes ranked by measured impact. Auto-detects whether the input is a `trace.zip`, a directory of unpacked trace files, or a single `trace.trace` / `trace.network` JSONL stream. Iterates via the `/confidence` skill — if root-cause certainty is below 90%, it digs deeper before recommending a fix. Use when handed a Playwright trace, asked "why is this test flaky?", "why did the test time out?", or asked to optimise an E2E suite with evidence. Triggers on "analyze trace", "playwright trace", "e2e trace", "test flake", "why did playwright fail", "playwright timing", "/playwright-trace-analyzer".

74

polish

skills/quality/polish/SKILL.md

Re-runnable pre-PR quality gate for the current branch. Composes two existing passes over the branch diff: a broad pr-reviewer pass (read-only review via the branch's open PR, which `pr-reviewer` requires) and a code-quality simplify pass (applies Class M mechanical refactors behind a confidence ≥ 90 % gate, reverting on failure). Run bare for the full review + simplify works; scope it with `review`, `simplify`, or the light `quick` mechanical pass. Commits each pass separately for traceability (`--no-commit` to skip). Use standalone any time mid-development to clean a branch. Note that `/create-pr` delegates its quality step to `review-loop` (the bounded convergence loop) post-draft, and the `review-loop` skill calls only `polish simplify` — never full `polish`. Triggers on "polish my branch", "clean this up before the PR", "review and simplify", "tidy up", "prep my branch", "/polish".

68

profile-optimizer

skills/analysis/profile-optimizer/SKILL.md

Analyzes React DevTools Profiler exports, Chrome DevTools Performance traces, and Chrome heap snapshots / heap-timelines / heap-profiles. Identifies the highest-impact bottlenecks (long tasks, expensive renders, layout thrash, wasted memoisation, blocking scripts, retained memory, leaks) and proposes concrete code fixes ranked by measured impact. Auto-detects the input format (React `.json` profile, Chrome trace `.json` / `.cpuprofile`, or `.heapsnapshot` / `.heaptimeline` / `.heapprofile`). Iterates via the `/confidence` skill — if root-cause certainty is below 90%, it digs deeper before recommending a fix. Use when handed a profile file, asked "why is this slow?", "why is memory growing?", or asked to optimise a hot path with evidence. Triggers on "analyze profile", "react profiler", "chrome performance", "optimize from profile", "profile this", "why is this slow", "memory leak", "heap snapshot", "/profile-optimizer".

69