CtrlK
BlogDocsLog inGet started
Tessl Logo

talk-birgitta-closing-keynote

Answers questions about, retrieves safe excerpts from, explains concepts from, and summarizes key arguments in Birgitta Böckeler's talk "State of Play: AI Coding Assistants" (AI Native Dev conference, 2026). Use when the user asks about the last 12 months in AI coding assistants, model-task fit, LLM statelessness, context window and attention trade-offs, coding harnesses, harness engineering/context engineering, skills/MCP/sub-agents/plugins/hooks, guide-and-sensor feedback loops, background agents and swarms, review bottlenecks, AI coding costs, cognitive surrender, or risk-based supervision of coding agents.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong instruction-only reference skill body: lean, well-structured routing procedures with explicit grounding and safety rules, and no wasted tokens. The main residual weaknesses are small — the quote.md/quotes.md ambiguity, no worked citation example, and implicit verification before attributing quotes.

Suggestions

Resolve the file-name ambiguity: determine whether the pre-extracted quote file is quote.md or quotes.md and reference only that file, removing the 'or' hedge in the Factual Q&A workflow.

Add one short worked example of a supported answer (a one-sentence answer plus a verbatim excerpt with an L#### line-range citation) so the citation format is unambiguous.

In the Factual Q&A workflow, add an explicit checkpoint step to re-verify each chosen excerpt exists verbatim in transcript.md before including it in the answer.

DimensionReasoningScore

Conciseness

The ~60-line body is lean routing guidance with no padding and no explanation of concepts Claude already knows — every section either carries talk-specific facts (the intro's frame, the glossary anchors like "Inferential vs computational sensors") or executable procedure. It matches the anchor 'Lean and efficient; assumes Claude's competence; every token earns its place'; score 4 would require identifiable over-explanation to trim, and there is none.

5 / 5

Actionability

Guidance is largely executable: exact file routing ("read `outline.md` to locate the relevant section, then read that section of `transcript.md`"), a concrete citation convention ("Cite by `L####` source line range"), and ready-made fallback phrasing ("She does not appear to address this in the talk"). It falls short of the fully copy-paste-ready score-5 anchor because of the hedge "Check `quote.md` or `quotes.md`" — the skill itself is unsure which file exists — and no worked example of an answer with a cited excerpt.

4 / 5

Workflow Clarity

All four workflows (Q&A, explain, apply, draft) are clearly sequenced numbered steps with an explicit not-found checkpoint in rule 3. It sits at the score-4 anchor ('clear sequence with most checkpoints present; minor validation gaps') rather than 5 because verification is implicit — no step confirms a chosen quote actually exists in `transcript.md` before answering, and the apply/draft workflows lack a final self-check. It is above score 3 since sequences are complete and the failure path is explicitly handled.

4 / 5

Progressive Disclosure

The body is a well-organized overview that routes by task type to one-level-deep, clearly signaled references (`outline.md`, `transcript.md`, `quote.md`/`quotes.md`), which approaches the score-5 anchor. It drops to 4 due to minor organization gaps: the ambiguous "`quote.md` or `quotes.md`" pair, and the referenced bundle files are not present in the provided bundle to verify, so navigation cannot be confirmed fully clean.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model reference-skill description: third-person voice, four concrete actions, an explicit 'Use when' clause dense with natural trigger terms, and an unmistakable niche. Every token is a functional trigger rather than padding, so verbosity concerns do not apply.

DimensionReasoningScore

Specificity

The description lists four concrete, distinct actions — "Answers questions about, retrieves safe excerpts from, explains concepts from, and summarizes key arguments in" a named talk — which comprehensively covers the use modes of a reference skill. It matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage'; it is not score 4 because there are no noticeable coverage gaps, and not below because nothing is generic.

5 / 5

Completeness

It explicitly answers both questions: 'what' via the four named actions against the talk "State of Play: AI Coding Assistants", and 'when' via an explicit "Use when the user asks about..." clause with concrete trigger phrases. This is the clear score-5 anchor (mirrors the PDF good example); the 'Use when' presence also avoids the completeness cap of 3.

5 / 5

Trigger Term Quality

The 'Use when' clause gives extensive natural phrasing users would actually say: "model-task fit", "LLM statelessness", "context window and attention trade-offs", "coding harnesses", "skills/MCP/sub-agents/plugins/hooks", "background agents and swarms", "review bottlenecks", "cognitive surrender", plus synonyms like "harness engineering/context engineering". This matches the comprehensive synonym coverage of the score-5 anchor; a user who saw the talk would naturally use these exact terms.

5 / 5

Distinctiveness Conflict Risk

It names a unique niche — a specific speaker's ("Birgitta Böckeler") specific talk at a named conference and year — so it would not trigger for any other skill. This matches the score-5 anchor 'Clear niche with distinct triggers; minimal conflict risk'; no neighboring anchor fits better despite re-checking.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
jscraik/Agent-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.