CtrlK
BlogDocsLog inGet started
Tessl Logo

talk-douglas-training-ai-on-your-own-code

Answers questions about Brian Douglas's talk on training AI on your own code. Use when a user asks about Brian Douglas's pipeline for capturing agent sessions, extracting skills from traces, fine-tuning small local models, tapes/steros tooling, SFT vs DPO decisions, or wants to apply his agent telemetry and training data approach to their own work with Claude Code, QLoRA, or parallel agents.

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, well-organized reference-skill template with a clear grounding workflow and sensible safety rules for untrusted source material. Its critical weakness is that all three files it instructs Claude to read (outline.md, transcript.md, quote.md) are absent from the bundle, making the entire progressive-disclosure structure unusable as shipped.

Suggestions

Include the referenced bundle files (outline.md, transcript.md, quote.md) under references/ and update the body's paths to match (e.g., references/outline.md), so the read-order instructions resolve to real files.

Resolve the ordering ambiguity between the grounding rules (read transcript.md sections first) and the Key quotes section (check quote.md first) by stating one explicit sequence, e.g., locate in outline.md → check quote.md for pre-extracted evidence → read the transcript section for context.

Replace the placeholder good-response example ("[safe excerpts from transcript.md]") with one concrete worked example showing a real question answered with an actual short quote.

DimensionReasoningScore

Conciseness

The ~40-line body is lean with no padding and no explanation of concepts Claude already knows; it moves straight to grounding rules ("read `outline.md` to locate the relevant section, then read that section of `transcript.md`") and a compact good/bad response example. Every section earns its place, matching the 'lean and efficient' anchor 5 rather than anchor 4's 'minor instances of over-explanation'.

5 / 5

Actionability

For an instruction-only skill the guidance is concrete and executable: a defined read order (outline.md → transcript.md section), "check `quote.md` first for strong citable evidence", explicit quoting rules, and a failure path ("If a claim isn't in `transcript.md`, say so explicitly"). It stops short of anchor 5 because the good-response example uses a placeholder ("[safe excerpts from transcript.md]") rather than a concrete worked example, and handling of partially-covered questions is not specified.

4 / 5

Workflow Clarity

The multi-step answering process is clearly sequenced (locate section in outline.md → read that transcript section → check quote.md for citable evidence → respond with quotes or state the topic wasn't covered), including a verification checkpoint for unsourced claims. It falls below anchor 5 because the ordering is slightly ambiguous — the grounding rules say to read transcript.md sections first while the Key quotes section says to check `quote.md` first "before searching the full `transcript.md`" — and there is no explicit recovery loop for ambiguous outline entries beyond the multi-section rule.

4 / 5

Progressive Disclosure

The body is structured around one-level references (`outline.md`, `transcript.md`, `quote.md`) that are clearly signaled, but the bundle contains no such files — no references/, scripts/, or assets/ directories exist, so every referenced path fails to resolve. Per the guideline to score against the actual bundle structure, the navigation the skill depends on is broken; this fits anchor 2's minimal-structure failure rather than anchor 1, since the SKILL.md itself is well organized rather than a monolithic wall of text.

2 / 5

Total

15

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does and when to use it, with specific, natural trigger terms tied to a uniquely named source. Keyword coverage is good but misses a few common user phrasings, keeping trigger term quality just short of the top anchor.

DimensionReasoningScore

Specificity

"Answers questions about Brian Douglas's talk on training AI on your own code" names the domain and the concrete scope of coverage ("capturing agent sessions, extracting skills from traces, fine-tuning small local models, tapes/steros tooling, SFT vs DPO decisions"). It lists several specific capabilities with only minor gaps (e.g., citing quoted excerpts from the transcript is implied but not stated), matching anchor 4 rather than 5, which requires a fully comprehensive multi-action inventory.

4 / 5

Completeness

It explicitly answers "what" ("Answers questions about Brian Douglas's talk on training AI on your own code") and "when" with a concrete "Use when a user asks about..." clause enumerating specific trigger topics. Both halves are explicit and concrete, matching the anchor 5 example pattern exactly; anchor 4 would require a weaker or less specific 'when' clause.

5 / 5

Trigger Term Quality

Natural trigger phrases are present and specific: "Brian Douglas", "training AI on your own code", "SFT vs DPO", "fine-tuning small local models", "Claude Code", "QLoRA", "parallel agents", "tapes/steros". Good synonym coverage (fine-tuning small local models / QLoRA), but a few natural phrasings a user might say (e.g., "that talk", "Brian Douglas's talk", "training a model on my codebase") are only partially covered, so it fits anchor 4 rather than the comprehensive anchor 5.

4 / 5

Distinctiveness Conflict Risk

The description is anchored to a named person, a named talk, and niche tooling ("tapes/steros", "SFT vs DPO", "QLoRA"), creating a clear niche with distinct triggers and minimal conflict risk with generic skills. It is far more distinguishable than anchor 4's "minor overlap risk with closely related skills" case.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
jscraik/Agent-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.