CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-ai-assistant-claims

Evaluate a vendor's productivity claim about an AI coding assistant against the evidence checklist from the talk. Use when someone quotes a percentage gain and asks whether the team should adopt the tool.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Evaluate AI Coding Assistant Claims

When a productivity claim about an AI coding assistant lands in front of you ("40% faster!"), do not argue with the number. Ask what it measured.

Steps

  1. Identify the task type behind the claim: boilerplate, tests, refactors, greenfield, or debugging. Reported gains cluster in the first two.
  2. Ask for the baseline. "Faster than what?" A junior on an unfamiliar codebase and a senior on their own service are different denominators.
  3. Check who reviewed the output and how long that took. Time saved typing is often time spent reading.
  4. Look for the "vibe coding" tell: code merged without anyone being able to explain it. Treat that as a risk finding, not a productivity finding.
  5. Recommend a two-week pilot on one real backlog slice, measured with the team's own before-and-after numbers.

Output format

Reply with a short table:

ClaimWhat it measuredApplies to us?Pilot?
(quote it)task type, baseline, revieweryes / partly / noyes / no

Example

Claim:     "55% faster with the assistant"
Measured:  time to first passing test on a toy HTTP server task
Applies:   partly (boilerplate-heavy service); not to our data pipelines
Pilot:     yes, 2 weeks, 3 volunteers, track PR cycle time
Repository
jbaruch/shownotes-template
Last updated
First committed

Also appears in

jbaruch/shownotes
In sync

since Sep 12, 2026

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.