Use when the user asks about Amit Kushwaha's AI Native DevCon talk on benchmarking agent-era systems, measuring performance beyond single LLM calls, inference, workflow complexity, tool use, and real-world workloads.
61
73%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./Plugins/aidevcon/skills/talk-kushwaha-benchmarking-agent-era/SKILL.mdAmit Kushwaha argues that agent-era benchmarks must measure workflows, tool calls, context length, inference behavior, and real-world complexity rather than only single-turn model output.
outline.md first to locate the relevant section or concept.quote.md for short supporting excerpts, then verify against transcript.md when precision matters.Answer from the bundled files. Use short excerpts only when they clarify the answer, and cite the transcript line IDs when available.
When the user asks how to apply the talk, identify the matching concept from the outline, summarize the relevant transcript evidence, and adapt it to the user's context. Mark anything beyond the talk as your own recommendation.
When comparing this talk with another AI Native DevCon session, ground this talk's side in outline.md and quote.md before drawing connections.
a3808a9
Also appears in
on Jun 8, 2026
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.