Flag concrete false-positive proof introduced by changed specs, not test helper or channel preferences. Advisory only; never gates Warden clearance.
58
66%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./.warden/skills/spec-provenance-review/SKILL.mdReview changes under evals/specs/** and evals/worlds/** for one question:
does this diff let a spec pass while the specific behavior it claims to test
is broken?
Test code has a different purpose from production code. Review the validity
of its evidence, not production hardening, abstraction, style, or preferred
helper usage. Channel conventions in evals/README.md are authoring guidance;
a channel mismatch alone is not a finding.
Report a MEDIUM (advisory) finding only when ALL of these hold:
Examples worth reporting:
Do not report:
evaluateOnSurface, document.body.innerText,
or probe.* merely because user.see/user.notSee could be used instead.
For example, opening /pricing and asserting new prices in rendered body
text is acceptable pricing evidence; the title saying "visitors see" does
not by itself require a different helper. Report only if the implementation
shows that the asserted text does not prove the specific claimed outcome.seed.*, direct API calls, or browser evaluation used to arrange state,
including setup between actions; report only when setup substitutes for
the behavior actually under test.agent.* in specs testing the agent, control rail, or voice.// TODO(primitive): comments or helper migration suggestions.Every finding is medium advisory; never report high or low, and never
turn helper-style policy into a finding. Use one finding per root cause, group
related locations, quote the claimed behavior, identify changed-code causality,
explain the reachable concrete failure that would still pass, address contrary
evidence, suggest the smallest fix, and state Clear when: with an observable
condition. If that evidence is missing, report nothing.
417244c
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.