Evaluate a vendor's productivity claim about an AI coding assistant against the evidence checklist from the talk. Use when someone quotes a percentage gain and asks whether the team should adopt the tool.
72
90%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
The canonical home for this skill is evaluate-ai-assistant-claims in jbaruch/shownotes-template
When a productivity claim about an AI coding assistant lands in front of you ("40% faster!"), do not argue with the number. Ask what it measured.
Reply with a short table:
| Claim | What it measured | Applies to us? | Pilot? |
|---|---|---|---|
| (quote it) | task type, baseline, reviewer | yes / partly / no | yes / no |
Claim: "55% faster with the assistant"
Measured: time to first passing test on a toy HTTP server task
Applies: partly (boilerplate-heavy service); not to our data pipelines
Pilot: yes, 2 weeks, 3 volunteers, track PR cycle timef8c9b77
Canonical home
since Sep 12, 2026
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.