ARTICLE
Engineering a Context-Driven Software Factory
Explore how context-driven software factories empower agents to automate tasks requiring judgment, enhancing engineering efficiency and reliability.

Patrick Debois

How context lets agents take on more work, and the engineering needed to make it reliable.
During an API migration, an agent discovers that a caller depends on undocumented error behavior. The ticket says to adopt the new interface, but the compatibility policy says to preserve existing behavior. It could add a wrapper, change the caller, split the migration or ask an engineer. The ticket alone does not settle the choice.
A software factory combines agents, tools and engineering practices to take work from intent to verified output. In this case, it also takes on a judgment an engineer would have made. With enough context about the intended outcome and the constraints, an agent can investigate and choose an approach without a programmed branch for every possibility. That makes some previously impractical automation worth attempting.
For an engineering team, the implications are practical:
- More work becomes possible: agents can exercise judgment where fixed workflows run out of cases.
- Authority needs evidence: permissions and checks must match the decisions being delegated.
- Experience becomes context: code, commits and PRs influence the next run, for better or worse.
The opportunity: automate work that needs judgment
The compatibility policy tells the agent what must be preserved. Previous decisions and examples help it work out which tradeoffs the team accepts. The model uses this context to choose an approach suited to the system it is changing.
The team can describe how to make a choice without specifying every action. Instructions such as “preserve existing behavior,” “split the change if releases need coordination” and “escalate if compatibility cannot be maintained” give the agent room to investigate while making the team’s expectations explicit.
The agent can then attempt cases the workflow author did not anticipate. It may still interpret the guidance incorrectly, which makes the next question important: which decisions can it act on, and how will the team check the result?
The boundary: define the factory’s authority
The agent needs room to choose an approach, but some constraints should be enforced independently of its interpretation. Consider a commerce API with order transitions and customer lookups. Migrating it involves three kinds of boundary:
| Boundary | Example | How it is enforced |
|---|---|---|
| Application behavior: what the software may do. | An order can move from confirmed to shipped, but a shipped order cannot be cancelled. | A runtime validator rejects forbidden transitions. Tests verify that the migrated API preserves those rules. |
| Workflow progression: when work may advance. | A migration cannot move from implemented to ready for release until compatibility checks pass. Failed checks return it for revision. | The harness checks the required evidence before allowing the next workflow state. |
| Agent permissions: which actions it may take. | The agent can edit the adapter and run tests, but cannot change production orders or release software. | Scoped credentials and tool permissions restrict access. The team can grant additional authority explicitly. |
Compiled context is a useful analogy for deriving deterministic checks from requirements. It does not imply automatic translation or complete coverage of the original meaning. A state model describes valid transitions, and a validator enforces them. Tests and workflow gates can check other requirements, while a linter catches some implementation patterns. The team must validate what each check covers, record its limits and review it when the source requirement changes.
The team chooses which boundaries this deployment needs and reviews changes to the rules and checks separately from the agent’s implementation. Inside those boundaries, the agent can investigate callers and choose an adapter design. Compatibility still includes details beyond state transitions: if the old API returns 404 for a missing customer and the new one returns an empty 200, a caller that relies on 404 needs that behavior preserved too.
How one correction helps the next task
During review of the API migration, an engineer catches that the adapter passes through the new API’s 200 response when a customer is missing. The caller expects 404. The agent corrects the adapter and adds a regression test. The team also updates its migration guidance: check how each caller handles a missing customer before replacing the old API.
Accepted work becomes evidence for future decisions. A later agent can copy the adapter, read its commit message and PR discussion, and use the regression test to understand the expected behavior. A maintenance agent can check callers already migrated. Each caller still needs assessment, but the same correction can help several workflows. We are producing future context every time we accept a change.
That gives us a feedback loop, though recording a fix does not mean the factory has learned from it. We should see later migrations handle the missing-customer case correctly before a reviewer has to raise it. Better guidance, examples and checks can produce that improvement without retraining the model.
The risk: a local mistake can become factory policy
Context poisoning can be accidental. Suppose review misses the adapter passing through 200 for a missing customer. The change lands, and the next agent copies it as an approved migration example. A PR description claiming that compatibility was preserved reinforces the mistake. After several migrations, callers that depend on 404 are broken, and the repetition makes the adapter look like an established convention. Deliberate poisoning can exploit the same trust, through external material containing instructions the factory should treat as untrusted input.
Context decay. Now suppose the adapter was correct, but the team later approves a new contract: migrated callers should handle an empty result with 200. The old 404 guidance and regression test remain. An agent following them may undo the intended behavior, and the stale test may reject a correct implementation. A linter or workflow rule derived from the old requirement can keep enforcing it just as reliably. The requirement, guidance and dependent checks need to change together.
Both problems can spread through shared context. An implementation agent copies the adapter, a reviewer checks against the outdated guidance, and a maintenance agent spreads the same behavior to other callers. Their agreement is misleading: all three are working from the same flawed source. To investigate that failure, the team needs to know where the rule came from, which callers it applies to and which checks depend on it.
Engineering practices that make delegation dependable
1. Package shared guidance as a versioned skill
Give implementation, review and maintenance agents the same approved standards through a shared skill. Include relevant examples, verification commands and escalation conditions. Let agents propose skill updates as reviewable changes. Check whether a lesson applies beyond the implementation that prompted it before sharing it with other projects.
Record what context each run received, including the skill revision and retrieved code, commits and PR discussions. Distinguish approved instructions from historical evidence and untrusted input. When an agent copies a bad pattern, the trace should help establish whether it received misleading context, missed the policy or misapplied it.
2. Implement authority through tool permissions
Configure the agent’s environment to match the authority it has been given. Use an isolated checkout, scoped credentials and explicit tool permissions. If the deployment requires approval for an action, have the harness check for that approval before enabling it. Give the agent a way to report a conflict, request a decision and pause the affected work.
Test that permissions hold. Try actions the agent is not allowed to take, including proceeding without approval or altering acceptance criteria. Confirm that the harness blocks them and records the outcome. Within those boundaries, leave the agent free to choose its investigation and implementation steps.
3. Turn failures into agent evals
Keep the regression test for the software, and add an eval for the task that produced it. The adapter test checks the response. The agent eval checks whether the factory can perform a migration while respecting the contract. Inspect the resulting code, behavior and tool actions. The agent’s claim that it succeeded is not enough.
Vary the task to challenge its judgment:
- A feasible fix: does it preserve the required behavior?
- A misleading old PR: does it follow approved guidance instead of copying the example?
- A conflict outside its authority: does it escalate rather than make an unauthorized choice?
- A changed requirement: does it apply the new contract only where it belongs?
Run these tasks when changing skills, retrieval, models, coding agents or the harness. Compare outcomes, inappropriate actions, escalation, cost and review effort across repeated runs. Keep evaluation criteria outside the implementation agent’s control. A second agent agreeing from the same flawed source is not sufficient validation.
4. Track dependencies and version each run
Link derived checks to the requirements they encode. Maintain a mapping from source revisions to generated tests, linters, workflow rules and cached evaluations. When a requirement changes, flag those dependencies for review and invalidate results that relied on the old version. Keep the scope explicit: changing the contract for some callers does not make the old checks wrong for every caller.
Record the context and check revisions alongside the model identifier, available model revision, settings, coding agent, tools and harness for each run. Evaluate updates on a small set of tasks before wider rollout, and keep the previous configuration available for rollback. Where immutable model versions or exact replay are unavailable, record the limitation.
What makes a factory ready to operate
Skills, loops and factories connect through shared context. Skills capture how work should be done, loops put that guidance to use with feedback, and a factory connects those workflows. Keeping them useful requires four ongoing commitments:
- Maintain shared guidance. Keep intent, standards, examples and escalation conditions in versioned skills that agents can retrieve and teams can review.
- Provide tools and verification. Support the work with scoped tools, tests, linters, validators and observability. Improve repository structure and setup wherever they make changes difficult to understand or verify.
- Close the feedback loop. Review outcomes and proposed lessons, update context and derived checks, and use task-level evals to establish whether those changes improve later runs.
- Share learning across workflows. Connect implementation, review and maintenance through guidance and checks they can all use. Verify that lessons apply in each setting instead of assuming they transfer unchanged.
Measure time to accepted work, review effort, repeated corrections, escaped defects and total cost. If engineers spend more time correcting the extra output, check whether the factory is saving them any work.
Context lets us delegate work we could not practically specify step by step. The code and history left behind will influence how the factory handles the next task. We are responsible for what the factory produces and for what future agents learn from it. The result should be fewer repeated mistakes and less work for the team to redo.
COPY & SHARE

Patrick Debois
Patrick is a pioneer, credited with coining the term DevOps, co-authoring the DevOps Handbook, and launching the very first DevOpsDays back in 2009. Since then, he’s been shaping the tech industry and brings depth across development, security, operations, and GenAI.
READING
·
0%
IN THIS POST
COPY & SHARE

Patrick Debois
Patrick is a pioneer, credited with coining the term DevOps, co-authoring the DevOps Handbook, and launching the very first DevOpsDays back in 2009. Since then, he’s been shaping the tech industry and brings depth across development, security, operations, and GenAI.
YOUR NEXT READ
AI DevCon NYC: From Coding Agents to Software Factories
Curating a conference is an interesting way to see where an industry is heading, because the program tends to reveal which questions have already become accepted and which ones people are only star...

Patrick Debois



