Find flaky and repeatedly failing tests from recent CI history and file a ticket for each one, before people learn to ignore them.
73
92%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Tessl evals compare success rates of agents with and without our optimized context