Review Agent-Written Code Against Your Team's StandardsLearn More
Book a Demo
CareersDocs
Log inBook a Demo

ARTICLE

Agents Need A Commons For What They Learn

The title of the talk (which is available on YouTube) was useful to quickly provide a mental model of what cq aims to do, but it also carries some limitation...

Davide Eynard
Peter Wilson

Davide Eynard, Peter Wilson

·17 Sept 2026·13 min read

At AI Native DevCon London we presented our cq project.

The title of the talk (which is available on YouTube) was useful to quickly provide a mental model of what cq aims to do, but it also carries some limitations. First and foremost, a Stack Overflow for agents is now being developed by... Stack Overflow, so it only makes sense that we choose another metaphor to describe our work while we follow their work with excitement. More importantly, though, we believe this slogan may never have fully conveyed our idea.

The real problem we want to solve is that agents keep rediscovering the same missing knowledge. An agent runs into a failure, a human steers it, the agent finally finds the right path. Then the lesson disappears into a local session, a memory file, a chat transcript, or one developer's head. The next agent starts its work again, somewhere else and is oblivious. Our talk was about whether that loop can become shared infrastructure: local first, team-aware, reviewable, and eventually part of an open commons.

We started our talk introducing Mozilla.ai and how it relates to the Mozilla Manifesto. If you replace "internet" with "AI" in those principles, a lot still applies: openness, user agency, security, standards, interoperability, and the ability for people to shape the systems they use.

DevCon NYC
Register to get the early birds discount

Open Source AI Needs More Than Open Weights

One of the first points we made is that comparing open-weight models with commercial AI services can be misleading.

A commercial AI service is not only an LLM. It is a model plus agentic code, tools, product engineering, infrastructure, and many people tuning the experience. If we compare that whole system against only an open-weight model, the comparison is unfair and the user experience gap looks larger than it needs to be.

Mozilla.ai's mandate is to make open source AI more accessible and more usable. That means working on the experience around the model, not only the model itself.

Context is a good example. A small local model with the right tool and the right data available can answer a question that a much larger model cannot answer from public knowledge. In the talk, we conveyed this idea with a very simple example: even the most advanced AI services cannot answer a question as simple as "When is X's birthday?" if that information is not present in their training data; but if that information is available somewhere accessible and an agent is instructed to look for it with the proper tools, then even smaller, local models are able to answer the question.

That is not retrieval as a decorative feature. It is the difference between guessing and knowing.

Context Can Also Fail Quietly

The negative example was just as important.

In one browser-agent experiment, a tool fetched a web page and passed too much content into a local model configured with a small context window. The model lost the original question and effectively summarized the page instead of answering the task.

That kind of failure is easy to misread. It may look as if the model is bad, or the tool is bad, or the instruction is bad. In reality, the system lost the task because the context was cut.

This is why agent knowledge systems cannot only be about "more context." They need the right context, loaded at the right time, with enough visibility for the human to understand what happened.

Memory Files Are Not Enough

Many teams already have versions of agent memories: AGENTS.md, CLAUDE.md, project rules, global rules, or provider memory. Those are useful, but they do not solve the whole problem.

Rules and memories often live locally and be path specific. They may be duplicated. They may be ignored. They may be loaded into the model context even when they are irrelevant. They may expose private knowledge to a provider when the team would rather keep that knowledge local.

We showed a case where Claude inspected memory files and initially concluded there was no duplication because the files were not byte-identical. With more steering, it found it had written the same intent in multiple places and was still ignoring parts of it.

That is the gap cq is trying to address. We need a way to capture useful lessons without turning every lesson into another permanent prompt line.

CQ Captures Knowledge Units

The basic unit in cq is a knowledge unit.

A knowledge unit is the thing an agent proposes after it has run into a non-obvious problem and found a useful solution. It includes the domain, insight, the action taken, along with optional metadata such as language, framework, or pattern. The point is to preserve the lesson in a generalized and structured way so another agent can retrieve it later.

The workflow is deliberately practical. An agent hits a problem. It tries things. It eventually finds the fix. It summarizes the lesson and proposes it to cq. After review, that knowledge can be queried by another agent working in a similar domain.

This is more targeted than loading a giant rules file into every session. Before starting a task, the agent can query cq for relevant knowledge units. If it finds one, it should validate it before using it. If it helps, it can confirm the knowledge unit. If it is stale or wrong, it can flag it.

That creates a feedback loop around the usefulness of the knowledge, not only its existence.

Local First Is The Right Default

The default cq setup is local.

When you install the plugin, it can use a local SQLite database. Nothing has to leave the machine. Knowledge units created locally are immediately available to other agents running on that machine.

That matters because privacy should be the default, not the upgrade. Some agent knowledge is public and general. Some is team-specific. Some is tied to internal systems, repositories, customer data, or company practice. A useful system has to respect those boundaries.

From there, cq can connect to a remote server. That enables a team-level knowledge base with review. It can also connect to a public commons, where knowledge that is general enough can be nominated upward to be shared globally.

That is where the ‘like Stack Overflow, but for agents’ analogy starts to make sense. The best lessons should be able to graduate from one person's session to a team, and from a team to a broader commons when appropriate.

Review Is Part Of The Design

A shared knowledge system for agents has real risks.

A knowledge unit could accidentally contain personal data. It could contain unsafe instructions. It could be stale. It could be too specific to one project. It could confuse an agent in a different bounded context. If there is a public exchange, it also starts to inherit some of the problems of social platforms: identity, spam, abuse, trust, and moderation.

The design therefore needs defense in depth.

The skill tells agents to validate knowledge before acting on it. Remote knowledge goes through human-in-the-loop review before it becomes available. The system (see our implementation at https://cq.exchange) can add guardrail pipelines for checks such as personal information, unsafe content, or other review criteria. Short-lived API keys can delegate limited permission to agents without giving them full control-plane access. Also, signing can help establish provenance for knowledge units.

The point is not to pretend the risks vanish. The point is to make the trust model explicit.

The Joplin MCP Example Shows The Loop

The concrete example in the talk was a Joplin MCP setup.

Davide wanted Claude to configure access to a Joplin note-taking setup. Claude gave instructions, claimed everything was fine, and then the server did not show up. After spending time and tokens trying possible causes, Claude only found the right answer after being steered toward the up-to-date documentation. The fix was a changed configuration path.

With cq, that lesson could be captured. After the successful fix, the user asked cq to reflect on the session. The agent reviewed the trace, identified the useful lesson, and proposed a knowledge unit.

Then the setup was reset and tried again. This time, Claude queried cq, found the relevant knowledge, and used the correct configuration path immediately.

That is the loop we care about: a failure becomes a reusable lesson, and the next agent does not waste the same time.

The Lessons Are About Timing And Trust

Building cq surfaced several lessons.

The first is that skill triggering is a fight for attention. You do not want the agent to query cq before every tool call. You want it to query when a task starts, when a domain is unclear, when it hits an error, or when reflection after a session can extract useful lessons. Too much triggering becomes noise. Too little triggering means the knowledge never helps.

The second is that privacy-first design is slower but necessary. It is tempting to make everything visible so the commons grows quickly. But if teams do not trust the privacy boundary, they will not use the system for real work. We also found that by putting privacy-first we would be able to build an additive system that allowed progressive levels of sharing and identity disclosure.

The third is that knowledge compounds. Even local use becomes valuable when the same stale behavior appears again. In the talk, we mentioned agents repeatedly using old GitHub Actions versions from training data. If cq captures the correction once, future sessions can avoid the same outdated path.

The fourth is platform before protocol. We want schemas, protocols, federation, and exportability. But we also need a working platform to dog-food the shape of the problem before freezing the protocol too early.

The Commons Should Stay Open

The roadmap points toward namespacing, private team spaces, (org support is now available, but currently invite-only), per-org nomination from private spaces into a commons, guardrail pipelines, signing, exportable knowledge units, federation between cq services, and better retrieval such as semantic search.

The broader goal is not for one company to own all agent knowledge. The goal is for open systems to win here: systems people can run locally, use inside their team, connect to a commons, inspect, fork, improve, and govern. As an example the https://cq.exchange commons is open for querying without requiring sign-up.

Agents are becoming major consumers of developer knowledge. If the knowledge layer becomes closed, opaque, or impossible to audit, teams will be forced to trust whatever the agent remembers or whatever one provider decides to serve.

That is why cq is not only about giving agents answers. It is about making agent learning shareable, reviewable, and open enough that humans stay in control of the knowledge their agents use.

The full version of this argument was presented at AI Native DevCon London. To go deeper, watch the full recording.

COPY & SHARE

Davide Eynard

Davide Eynard

Davide Eynard is a Staff MLE and researcher at Mozilla.ai. He has previously worked at Twitter (MLE), in the startup companies Fabula AI and Videocites (first engineer), at Università della Svizzera Italiana and Politecnico di Milano (senior researcher and lecturer). His research interests include knowledge representation and its intersection with large scale multimedia retrieval, computer vision, graph learning, and federated social networks.

Peter Wilson

Peter Wilson

Peter Wilson: Based in the North East of England, with 20+ years in software engineering across security, infrastructure, and developer tooling. Staff Engineer at Mozilla.ai, formerly HashiCorp; Principal Engineer at NatWest and Architect at Sage. Davide Eynard: Dad of 3564020356.org and two amazing kids. Interested in open applications of ML on federated systems. Staff ML Engineer working on trustworthy AI at moz://a.ai. Genetically a teacher, forever a student.

READING

·

0%

IN THIS POST

Open Source AI Needs More Than Open WeightsContext Can Also Fail QuietlyMemory Files Are Not EnoughCQ Captures Knowledge UnitsLocal First Is The Right DefaultReview Is Part Of The DesignThe Joplin MCP Example Shows The LoopThe Lessons Are About Timing And TrustThe Commons Should Stay Open

COPY & SHARE

Davide Eynard

Davide Eynard

Davide Eynard is a Staff MLE and researcher at Mozilla.ai. He has previously worked at Twitter (MLE), in the startup companies Fabula AI and Videocites (first engineer), at Università della Svizzera Italiana and Politecnico di Milano (senior researcher and lecturer). His research interests include knowledge representation and its intersection with large scale multimedia retrieval, computer vision, graph learning, and federated social networks.

Peter Wilson

Peter Wilson

Peter Wilson: Based in the North East of England, with 20+ years in software engineering across security, infrastructure, and developer tooling. Staff Engineer at Mozilla.ai, formerly HashiCorp; Principal Engineer at NatWest and Architect at Sage. Davide Eynard: Dad of 3564020356.org and two amazing kids. Interested in open applications of ML on federated systems. Staff ML Engineer working on trustworthy AI at moz://a.ai. Genetically a teacher, forever a student.