Review Agent-Written Code Against Your Team's StandardsLearn More
Book a Demo
CareersDocs
Log inBook a Demo

ARTICLE

Skills Everywhere Need A Context Pipeline

Why reusable agent skills need a tested, versioned context pipeline that can travel across repositories, tools, and rapidly changing coding harnesses.

John Groetzinger

John Groetzinger

·17 Sept 2026·16 min read

This post reflects my own experience and opinions, not an official Cisco position.

When I talk about putting skills everywhere, I am not talking about scattering more markdown files across every repository and hoping agents magically improve. That is how you get context drift, duplicated work, and a shared folder nobody trusts. The harder problem is how to package context once, evaluate it, keep it in sync, and distribute it to the humans and agents that need it.

That was the frame for my talk, "Skills Everywhere." I wanted to talk about the part of agent adoption that starts to hurt inside a larger organization: not whether one person can write a useful skill, but what happens when many teams start depending on skills as shared infrastructure.

My view has changed over the last year. I spent a lot of time jumping between harnesses, chasing the newest model, and assuming that when something broke the model was probably the problem. More recently, skills changed that. Once I could carry useful context across Claude Code, Devin, GitHub Copilot CLI, and other harnesses, the bottleneck became much clearer: the model was often capable enough, but the context was not engineered well enough.

So the argument is simple. You do not always need a smarter model to get business value. Very often, you need smarter context engineering. Skills are one of the ways to make that context reusable, testable, and transferable.

Use This Talk As Agent Context

Tessl has turned my AI Native DevCon talk into a skill your agent can use as context. You can also watch the full recording.

DevCon NYC
Register to get the early birds discount

Why Did Skills Change My Model Strategy?

The first shift for me was realizing that skills can make cheaper, faster models viable for more work. If a workflow is repetitive and the skill carries the right context, I do not always need the most expensive model in the stack.

That matters in enterprise environments because agentic fan-out changes cost very quickly. A single prompt can spawn many sub-agents. A harness can turn one request into a large amount of model usage. If every workflow defaults to the highest-end model, cost becomes a serious constraint.

For the last few months, I have used a medium reasoning model for a lot of my day-to-day work. I still reach for higher-end models when I need complex planning or very large context, but I challenge engineers to start with the medium tier and move up only when the task actually requires it.

The reason that works is not model magic. It is context. If the skill tells the agent how we structure projects, which tools we prefer, what patterns are allowed, and what success looks like, the workflow becomes more deterministic. If I can get the same result with a cheaper, faster model, I should not pay for the slower one just because it feels safer.

That is why evals are part of the story from the beginning. Without evals, you are guessing whether the skill improved the outcome. With evals, you can ask whether the cheaper model still produces the correct behavior for the workflow that matters.

What Counts As A Skill?

A skill can be simple. At the bare minimum, it starts with a SKILL.md that gives an agent instructions. But it can also include rules, supporting markdown, scripts, scaffolding, examples, and evaluations.

One simple skill I built for my team defines repository standards. We have technical debt, we have people creating new projects, and we have models trained on patterns that are not always current. If I ask an agent to create a new Python project and it gives me requirements.txt with a pile of pip install assumptions, that is not what I want. I want our projects to use uv. I can put that into a skill so engineers do not have to remember every setup rule by hand.

The important part is not that this is a clever prompt. The important part is that it becomes shared context. A new project should not depend on whether the developer remembered to paste the right instruction into a chat window. The agent should be able to load the same standard each time.

That is also why skills need evaluations. Define the minimum behavior that should always work. As you improve the skill, the old behavior should not silently break. If a skill is carrying a standard, an operational process, or a knowledge pattern, it has to be tested like something the team depends on.

How Do You Turn Existing Knowledge Into Skills?

The first story I told came from Cisco TAC, the Technical Assistance Center. I spent many years there, working on high-pressure customer problems for financial institutions, hospitals, critical infrastructure, and other enterprise environments where downtime matters.

TAC has a strong knowledge-base culture. When an engineer solves a difficult case, they write the experience down so the next engineer and the next customer do not have to suffer through the same three-week investigation. That knowledge is valuable because it is curated, maintained, and grounded in real incidents.

The obvious AI move is to let an agent search the knowledge base. We tried versions of that. It produced some wins, but it also produced hallucinations and answers that did not make sense because not every article is current and not every document should be consumed by an agent.

The better pattern was not to blindly feed everything into the agent. The better pattern was to pipeline selected, high-quality articles into skills. Engineers can pick the articles that are actively maintained and valuable, convert them into agent-usable context, and build evals around the outcomes we expect.

Think about a support case. A customer opens a case with a problem description. Three weeks later, after logs and investigation and diagnosis, the team solves it. The question is: how do we shorten that path from initial description to likely solution? A good skill can carry the relevant diagnostic context so the agent can help earlier.

But the skill should not drift away from the source article. If the article changes, the skill needs to change. If the change is minor and the evals pass, the pipeline can release the updated skill automatically. If the change introduces a new topic or materially changes behavior, that is where a human should review it and probably add a new eval.

That is the context pipeline I care about: source knowledge, generated skill, evaluation, release, and sync.

Stop Writing Every Skill By Hand

One point I made strongly in the talk is that the skill is not for you. The skill is for your agent.

That does not mean humans do not matter. It means the most valuable human time should go into source quality, expected outcomes, and evaluations. Humans should decide what good looks like. Humans should validate whether the agent is helping. But humans do not need to hand-craft every word of every skill.

If you already have a strong knowledge base, use the model to transform that knowledge into the shape it needs. Let the model organize the information for its own use. Then validate the output through evals and review gates.

This matters because manual skill writing does not scale. Large organizations already have too much knowledge for each team to rewrite it into prompts by hand. If every workflow depends on one person maintaining a private prompt, you have not built organizational capability. You have built a local habit.

The pipeline approach also keeps human documentation alive. I do not think human docs automatically die in this world. In places with strong documentation culture, the opposite can happen. The docs become more important because they ground the agents. But that only works if the agent-facing context stays aligned with the human-facing source.

How Did We Use A Skill To Roll Out Evals?

The second story was developer-focused. We were building a new AI-native platform at Cisco: a multi-agent orchestration system with a semantic router, specialist agents, embedded AI actions, and a unified customer interface. Different teams own different agents, and more agents are being added over time.

That architecture only works if we can evaluate it. We needed to know whether a user question went to the right agent, whether the answer was good, whether the right tools were called, whether parameters made sense, and whether latency and token usage stayed within expectations.

I was asked to help decide whether the system was good enough to ship. That meant we needed an eval framework teams could actually use. We used tools such as LangChain, LangGraph, and LangSmith-style observability, but the usual dataset and eval patterns were not enough for how I wanted teams to compare environments.

One concrete example was dataset format. Many eval datasets are JSON files, and JSON often becomes one huge line. That is painful for agents to edit. We moved toward JSONL so each example lives on its own line and a coding agent can make precise updates without dragging an enormous file into context.

The bigger issue was rollout. I had to explain this approach to multiple globally distributed teams with very little time. A meeting was not going to do it. A one-hour presentation about testing was not going to make every team produce consistent datasets, CI integration, environment-aware invocation, and shared metrics.

So I turned the eval framework into a skill.

The skill contained the instructions, examples, scripts, example evals, dataset schema, and patterns I had already worked through with one team. I had an agent help distill the working approach into a reusable skill for other coding agents. Then I rolled it out gradually: one team, feedback, another team, then more.

The result was much more consistent than asking each team to invent its own eval framework. Everyone could start from the same pattern. Their coding agents could do much of the setup. The teams still owned their agents and their data, but the framework gave them a shared language.

That is what skills are good at: distributing a way of working.

How Do Humans Stay In Sync?

The eval framework created a second problem. Managers and stakeholders wanted to know what had changed, but not everyone lives in GitHub. Some people live in Confluence, Jira, Webex, or management workflows.

The wrong answer would have been to write a separate Confluence page by hand. It would drift. I would forget to update it. The repository would say one thing, the wiki would say another, and nobody would trust either.

The better answer was to keep the source in one place. The skill lives in a repository. The human README lives in the repository. The agent-facing context and the human-facing explanation can be updated together. Then a deterministic script can sync the README into Confluence.

That does not require an LLM to edit Confluence directly. Markdown can become HTML. A heading can become a heading. The important part is that the manager reads the same source of truth the agent is using, translated into the platform they actually open.

That small workflow solved a real organizational problem. It kept the repository authoritative while making the information accessible outside Git.

The Cultural Shift Is Asking, Is This A Skill?

The cultural shift I am trying to build is the reflex to ask: is this a skill?

When a new engineer asks where the DevOps information is, the answer should not only be a link to a wiki page. The better question is whether we have a skill that explains it so the engineer's agent can use it. If we do not, maybe we should create one together.

That does not mean every thought becomes a skill. It means repeated explanations, standards, workflows, onboarding knowledge, eval patterns, and operational procedures should be candidates for reusable context.

It also means avoiding skill explosion. I do not want 15 engineers creating 15 slightly different skills for the same workflow. That wastes tokens, creates confusion, and makes quality hard to judge. Shared skills need ownership, versioning, evals, and a clear path for contribution.

In Q&A, I compared this to a shared library. People can submit pull requests, but they should not be able to break behavior others depend on. Evals answer that question. If the skill claims to support a behavior, write an eval for it. When the skill changes, run the eval.

Start With One Repeated Explanation

You do not need to begin by building a large internal platform. We got there through years of work and a lot of friction. The starting point can be much smaller.

Pick one concept or workflow you explain over and over. Turn it into a skill. Decide who the audience is. Decide where the source of truth lives. Decide where it needs to be distributed. Add a basic eval that defines what the skill must help the agent do.

Then version it honestly. A 0.0.x skill is experimental. Use it yourself. Share it with a few people. Get feedback. When it becomes more reliable, move it forward. When it reaches 1.0, people should be able to trust that it does what it says with minimal friction.

The durable investment is not the harness you happen to use this month. Agent architectures will keep changing. Models will keep changing. The context your organization knows, tests, and maintains is the part worth compounding.

That was the argument I brought to AI Native DevCon London: skills everywhere only works if skills become part of a context pipeline. Build the pipeline now, because everything around it will keep moving.

COPY & SHARE

John Groetzinger

John Groetzinger

John Groetzinger is a Principal Engineer on the Cisco Customer Experience Engineering team with deep roots in network security. He spent 12 years as a technical leader in Cisco security TAC, where he built automation that started at a startup, survived the acquisition into Cisco, was adopted into the security product line, and won industry awards over a decade later. That security engineering background shapes how he approaches agent development today — designing agent architectures and evaluation frameworks, building internal agent platforms, and helping engineering teams across Cisco adopt agentic development. He is an AWS Certified Solutions Architect and has been practicing AI-native development since 2023.

READING

·

0%

IN THIS POST

Use This Talk As Agent ContextWhy Did Skills Change My Model Strategy?What Counts As A Skill?How Do You Turn Existing Knowledge Into Skills?Stop Writing Every Skill By HandHow Did We Use A Skill To Roll Out Evals?How Do Humans Stay In Sync?The Cultural Shift Is Asking, Is This A Skill?Start With One Repeated Explanation

COPY & SHARE

John Groetzinger

John Groetzinger

John Groetzinger is a Principal Engineer on the Cisco Customer Experience Engineering team with deep roots in network security. He spent 12 years as a technical leader in Cisco security TAC, where he built automation that started at a startup, survived the acquisition into Cisco, was adopted into the security product line, and won industry awards over a decade later. That security engineering background shapes how he approaches agent development today — designing agent architectures and evaluation frameworks, building internal agent platforms, and helping engineering teams across Cisco adopt agentic development. He is an AWS Certified Solutions Architect and has been practicing AI-native development since 2023.

YOUR NEXT READ

Double your coding agent’s chances of writing secure code with the CodeGuard Skill

Enhance AI coding agents with the CodeGuard Skill to improve secure code generation by applying Cisco's security rules, covering 23 categories and multiple languages.

Simon Maple
John Groetzinger

Simon Maple, John Groetzinger

·12 Feb 2026·10 min read
Read article