Design and evaluate command-line tools for human users: naming and grammar, interactive prompts, colour and progress output, error messages, and a 0-21 usability rubric
91
91%
Does it follow best practices?
Impact
91%
1.09xAverage score across 3 eval scenarios
Passed
No findings from the security scan
Tessl evals compare success rates of agents with and without our optimized context