ARTICLE
Collective Intelligence Starts With Attributed Mistakes
How attributed mistakes, provenance, and measured corrections help agent teams turn isolated failures into durable collective intelligence.

Edouard Maleix

What did your agents learn yesterday that your team still knows today?
I opened my AI Native DevCon talk with that question.
A few months ago, repeated agent mistakes were annoying. Now they are getting expensive.
On Monday, an agent updated our REST API. It regenerated the OpenAPI specification and the TypeScript client, but missed the Go client.
On Tuesday, I corrected it. The agent acknowledged the mistake and promised that it would not happen again.
On Wednesday, a fresh session missed the Go client again.
CI would have caught the stale client before merge. But we had already paid for the lesson in review time, an extra agent iteration, and another interruption. The team had retained nothing.
As agents move upstream into specifications, reviews, planning, and architecture, these missed lessons stop being local. The harder question is no longer how to get more output from a model. It is how to stop paying for the same confusion twice.
Teams need a knowledge factory, not another static rulebook nobody trusts. It captures interruptions from real work, preserves where they came from, turns them into reusable guidance, tests whether that guidance helps, and retires it when it stops helping.
Attribution is what makes that factory trustworthy. It preserves who learned what, which event produced the lesson, how the lesson was transformed, and whether reusing it later made a difference.
Use This Talk as Agent Context
Tessl has turned my AI Native DevCon talk into a skill your agent can use as context. You can also watch the full recording.

Automation Makes Forgetting More Expensive
When an agent only helps with implementation, a missed lesson may cost one coding session. You correct it, swear at your terminal, and move on.
Once agents also write specs, review code, plan changes, and shape technical decisions, the same lesson has more places to leak into. A bad assumption can travel through the development lifecycle with excellent formatting and no obvious interruption to expose it.
Most teams respond by adding instructions: another line in AGENTS.md, another skill, another project note, another comment explaining how things are usually done. It feels like progress because words were written down. But words written down are not the same thing as reusable knowledge.
The rules accumulate until the repository starts looking like it is managed by configuration sediment. Nobody knows which instruction came from a real failure, whether it still applies, or whether it appeared when the agent needed it.
A team has collective intelligence when the next agent does not have to rediscover what the previous one already learned the hard way. For that to happen, the lesson must survive the session, the team must be able to trust where it came from, and it must return at the right moment.
Without those properties, we do not have collective intelligence. We have automation folklore.
Attribution Starts With a Separate Actor
In the talk, I showed a pull request I opened for platformatic/mcp, a Fastify plugin for building MCP servers.
Claude Code produced most of the work, but the evidence told a different story. The commits carried my name and GPG signature, and Claude Code opened the pull request through my GitHub account. Everything looked clean, but it did not tell the truth about authorship.
Identity is the first building block. A separate identity names the agent as author, lets it sign its changes, and limits the authority it receives.
It does not shift responsibility away from the human operator. It separates authorship, authority, and responsibility. If an agent acts through my identity, we cannot tell which actor produced the change. If it records a lesson through my identity, we cannot tell who encountered the failure or whose later behaviour improved.
Most debates about AI authorship are trust debates in disguise: who made this change, why did they make it, and what should a reviewer rely on?
Identity begins to answer the first question. It does not answer the second. For that, the work needs somewhere to preserve what happened.
Evidence Comes Before Guidance
I previously argued that before teams can evaluate agent context, they need to generate trustworthy evidence.
This distinction matters because evidence and guidance are not the same artifact.
An incident, a correction, the reasoning attached to a change, or a review that catches a broken assumption belongs on the evidence side. A rule, a summary, a context pack, or a skill is derived from that evidence. If we collapse the two too early, we get polished advice with a weak chain back to reality.
In the workflow I presented, evidence lands in a diary. The unit inside that diary is an entry: a signed, attributed record of a decision, incident, correction, or useful discovery.
An entry is not yet team knowledge, and it is not automatically a new rule. It preserves the event before a later summary cleans away the details that made it meaningful.
For example, when an agent wrote a database query that bypassed our transaction wrapper, the entry linked to the code review that explained the problem and the commit that fixed it. A later session can retrieve the trap, the correction, and the evidence behind both.
Capture should be easy, but not indiscriminate. The goal is not to preserve every conversation, token, or tool call as sacred memory. It is to retain the incidents, decisions, corrections, and unresolved gaps that may deserve to shape future work.
Entries are raw material. The next problem is deciding which ones belong together.
Curation Turns Evidence Into a Pack
Once enough entries exist, the team can discover patterns across them: repeated incidents in one scope, related corrections, or several failures caused by the same missing practice.
Selected entries become a pack: a small, deliberately ordered bundle assembled for a particular task area. A pack is not a bag containing everything the diary remembers. It is closer to a small exhibition whose pieces make a useful argument together.
The raw pack is still not necessarily what the agent reads. An LLM renders it to fit the agent's context budget—as Markdown, repository guidance, or a skill loaded on demand.
That transformation is useful and dangerous.
Rendering trims repetition and increases density, but every summary risks flattening qualifications, removing contradictions, or inventing a cleaner story than the evidence supports. Compression can make context easier to inject while making it harder to verify.
This is why attribution must survive rendering. Each part of the generated guidance should point back to the entries it came from. The useful version of a rule is not simply:
Always regenerate the Go client.
It is:
Regenerate the Go client when this API changes because we missed it before; here is the incident, here is the correction, and here is the change that fixed it.
Now a teammate can inspect why the guidance exists. If the code generator changes six months later, the team can revisit the source instead of maintaining a mysterious rule forever.
The chain now looks like this:
agent identity → signed change and attributed entry → curated pack → rendered guidance with source references
Attribution is not decorative metadata attached at the end. It is what keeps the transformation auditable from beginning to end.
Knowledge Has to Earn Its Place
Capturing and rendering context still does not prove that the result deserves to influence future work.
As I wrote in Coding Agents Need a Knowledge Factory, Not Just a Knowledge Base, a serious system must test its output. Useful guidance should earn its place, while stale or ineffective guidance should eventually decay.
For a rendered pack, I care about three questions:
- Did the rendering stay faithful to the source entries?
- Did the result help on a real task?
- Did the guidance appear when it was needed?
The talk focused on the first two.
Fidelity asks whether the rendered guidance remained true to its sources. Did it preserve the important conditions and contradictions, or did summarisation quietly turn a narrow lesson into a universal rule?
Usefulness asks a different question. Even a perfectly faithful summary can be useless. Does loading it actually help the agent avoid the failure?
The useful measurement is the delta. Run the same task with the same agent in a controlled environment, once without the pack and once with it. Then compare the outcomes against explicit criteria.
In the Go client scenario, the baseline run failed with a score of 67. With the relevant pack loaded, the run passed with 95. The pack did not make the model smarter in general. It returned one specific correction at the moment the agent had to make the same decision again.
The scenario matters as much as the score. It came from a real incident, not from reverse-engineering a tidy benchmark out of already-correct code. A real mistake already knows what failure looks like because somebody paid for it.
LLM judges do not make this automatic or objective. Before trusting a judge, a human needs to label an initial batch and tune the criteria until the model catches what the human would catch. Otherwise the eval produces a precise-looking number for a poorly specified judgement, which is one of our industry's more reliable forms of theatre.
Evals provide evidence for keeping, revising, or removing guidance. They do not absolve humans from deciding what good means.
This is where the deeper value of attribution appears. When a pack improves a task, attribution lets us trace the gain back through the render, the selected entries, the original incidents, and their authors. Without that chain, the eval tells us only that some context helped once.
The selection loop remains experimental. Agent knowledge does not yet promote, demote, or repair itself safely. Attribution gives teams the provenance they need to keep, revise, or retire guidance responsibly.
The system can test condensed guidance without hiding its source. Useful guidance earns trust for now—not forever, because models, repositories, and practices change.
That is the difference between memory and a knowledge factory. Memory retains things. A factory transforms them, checks the output, and rejects what no longer holds up.
More Autonomy, Not Less Responsibility
Attribution also supports more autonomous workflows: agents can claim tasks while other agents or humans evaluate the results. Their histories may reveal useful specialisations—coder, critic, or domain expert.
How those identities grow into roles—and how autonomous agents choose, review, and trust one another—is a story for another article. I am already working on it as a series, because the topic is too large for one post.
A separate identity does not turn an agent into the person responsible for the organisation deploying it. Persistent memory does not make every remembered lesson true. An evaluation score does not decide what the team should value.
Humans still set goals and decide what good means. Agents contribute continuity, repetition, and recall. Attribution makes the collaboration inspectable.
Follow the Experiment
MoltNet is the open-source infrastructure and experiment behind this work, and LeGreffier is the coding-agent workflow I am building on top of it. If you want to explore the argument further, start with how agents generate trustworthy evidence, then read why coding agents need a knowledge factory rather than another static knowledge base.
You can also watch the AI Native DevCon talk, load Tessl's agent-ready version of the talk, or start with the MoltNet documentation.
What Did Your Agents Learn Yesterday?
The question I opened with is still the one I care about:
What did your agents learn yesterday that your team still knows today?
If the answer is hidden in a closed session, the team did not learn it. If it became an unattributed rule nobody can explain, the team cannot trust it. If it was rendered into guidance but never tested, the team does not know whether it helped.
Collective intelligence begins when a mistake can travel without losing its story: who encountered it, what happened, which correction followed, how the lesson was transformed, and whether bringing it back improved the next attempt.
One attributed mistake is small. Enough attributed mistakes, deliberately shaped and tested, become knowledge the whole team can use.
If you try any of this, I would rather hear what broke than what worked. That interruption is the raw material.
COPY & SHARE

Edouard Maleix
Edouard Maleix is a freelance consultant based in Vienna, helping startups scale past their MVP — system design, application security, dev productivity, and AI integration. He's currently building an open source project: MoltNet, a platform to turn AI agents' experience into proven, reusable context.
READING
·
0%
IN THIS POST
COPY & SHARE

Edouard Maleix
Edouard Maleix is a freelance consultant based in Vienna, helping startups scale past their MVP — system design, application security, dev productivity, and AI integration. He's currently building an open source project: MoltNet, a platform to turn AI agents' experience into proven, reusable context.