Ensure thorough validation, error recovery, and transparent reasoning in research tasks with multiple tool calls
40
38%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Fix and improve this skill with Tessl
tessl review fix ./bundled/skills/comprehensive-research-agent/SKILL.mdThis skill addresses common failures in multi-step research tasks: unhandled tool errors, missing validation, opaque reasoning, and premature conclusions. It provides structured protocols for source validation, error recovery, and thinking transparency that significantly improves research quality and reliability.
After (Pattern): 'Search returned 15 results on context engineering. Evaluating relevance: Liu et al. (2024) appears most authoritative on 'lost in the middle' phenomenon; Anthropic documentation likely has current context window specs; Patel (2023) covers RAG best practices. Ranking these as top 3 priorities. Reading top result first. If the primary source fails (URL error), will try backup search for correct documentation URL and note the gap in final report.'
After (Pattern): Tool returned 404 for Anthropic URL. Thinking: 'Primary source failed. Fallback: search for alternative Anthropic documentation URL or find archived version. If unavailable, note context window data from secondary sources only and add disclaimer about verification status.' Then: 'Cross-validated Claude context window: Anthropic blog (successfully read) and two developer documentation sources agree on 200K. Confident in this claim.' Source tracking table shows: Anthropic URL (failed, backup used), Blog (success), Dev docs (success).
Complex research tasks with multiple tools (6+) and multi-step reasoning chains typically achieve scores in the 65-75 range. This is not a limitation of the prompt but reflects:
Focus on relative improvement and pattern elimination rather than absolute scores. A 5-10% improvement from optimization is significant for complex tasks.
Generated: 2026-01-11 Source: Reasoning Trace Optimizer Optimization Iterations: 10 Best Score Achieved: 72/100 (iteration 4) Final Score: 70.0/100 Score Improvement: 67.6 → 70.0 (+3.6%)
f627ab5
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.