Review Agent-Written Code Against Your Team's StandardsLearn More
Book a Demo
CareersDocs
Log inBook a Demo

ARTICLE

Spec-Driven Development Works When Specs Are System Use Cases

The story that led me into spec-driven development started with a volunteer management system.

Simon Martinelli

Simon Martinelli

·17 Sept 2026·12 min read

The story that led me into spec-driven development started with a volunteer management system.

I live in a small town in Switzerland and help a sports club that runs track and field competitions. Since 1997, I have written software for them because ranking lists, competition administration, and event logistics otherwise take a lot of manual work. In summer 2024, someone from the club told me their volunteer management system was outdated and asked whether I could rebuild it with AI.

I did, and it was relatively quick. Then someone from a music festival heard about it and asked for a similar system. That is when the problem appeared: I had not really specified what I was implementing. The requirements had to change, the assumptions moved, and I needed a better way to use AI for business applications.

My talk, "Lessons from Spec-driven Development," was about the process I arrived at: using system use cases, an entity or domain model, architecture constraints, skills, guardrails, and risk-based review to generate software from durable specs.

Use This Talk As Agent Context

Tessl has turned my AI Native DevCon talk into a skill your agent can use as context. You can also watch the full recording.

DevCon NYC
Register to get the early birds discount

Why A Process-Centric Approach?

I mostly work on Java business applications for enterprises: insurance, wholesale, retail, government, and large internal systems. I do not usually build developer tools or greenfield products. That shapes how I think about specs.

Many spec-driven tools are developer-centric. They take a product requirements document, generate a plan, generate tasks, and then implement those tasks. That can work, but it does not cover the whole software development lifecycle in the way many enterprise projects need.

My approach is process-centric. I call it the AI Unified Process. It starts with vision and requirements, then uses system use cases and a domain model as the core specification artifacts. Those specs sit alongside architecture guidance, skills, MCP servers, tests, and review practices.

The important difference is that I skip the plan-and-task phase. I use the system use case as the input to code generation.

Why System Use Cases?

System use cases are not a new idea. Ivar Jacobson introduced use cases decades ago, and I used them earlier in my career at Swiss Railways as a communication specification between business and engineering.

They work well with AI because they are structured and widely understood. A use case has actors, preconditions, a main success scenario, alternative flows, postconditions, and acceptance criteria. It describes behavior in enough detail for both stakeholders and agents.

Compared with user stories, I find use cases better suited to AI generation. A user story is often too small. One story may represent one flow inside a larger use case. If the goal is to generate coherent business application behavior, I want the larger behavioral unit.

The use case does not describe everything. A UI may come from Figma. An API may come from an API spec. But the use case defines the behavior, and that is the part I want the agent to implement consistently.

Greenfield And Modernization Need Different Flows

In a greenfield project, the process can start with business analysts, product owners, requirements engineers, or software engineers creating the entity model and use cases. Those specs are reviewed against a definition of done. Then the engine generates code and tests using skills, MCP servers, guidelines, and guardrails.

Modernization is more common in my work. There, I reverse-engineer use cases, entity models, tests, and documentation from existing systems. Business people review the extracted use cases. Then we generate a new implementation from the behavior rather than lifting and shifting old code.

That distinction is important. Modernization is not simply translating COBOL to Java or one stack to another. We tried direct transformations decades ago, and they were not enough. Real modernization means rethinking how people work with the software, preserving necessary behavior, and adding missing features or improved flows.

The spec becomes the stable middle layer. The old system informs it. The business validates it. The new implementation is generated from it.

Reviews Should Be Based On Risk

Review is still part of the process, but the amount of review should depend on risk.

In an ERP system, not every module has the same criticality. If inventory management fails temporarily, the business impact may be manageable. If order management fails, the company may lose money immediately. Those areas should not have the same review path.

This is not special to AI. It is basic risk management. AI just makes it more visible because generation can happen quickly. The faster the implementation path, the clearer the review policy needs to be.

For APIs, I often prefer test-driven generation: start with tests and use them to drive the code. For UIs, the order can differ because you may need the UI shape before creating meaningful tests. The process should match the system and the risk.

PetClinic Shows The Pattern

I used Spring PetClinic as a small demo because it has the shape of a business application: UI, data, owners, pets, vets, visits, and simple workflows.

I reverse-engineered PetClinic into a use case diagram and an entity model. The diagram identifies actors, roles, and modules such as login, doctors, owner management, pet management, and visit management. The entity model captures the data relationships.

Then I created use cases. For example, "list doctors" includes the actor, preconditions, scenario, API behavior, postconditions, and acceptance criteria. With the right skills and guardrails, the agent can implement from that use case.

In these projects, I do not prompt everything manually. We have skills for specification, implementation, testing, and stack-specific behavior. The skills are iterated and shared, although sharing across different agents and tools is still a practical problem.

Architecture Has To Fit AI Work

Architecture has a large effect on how well this works.

Many organizations adopted microservices in a naive way and ended up with distributed big balls of mud. An insurance company with hundreds of microservices and hundreds of micro frontends creates a difficult context problem for AI. The code an agent needs may be scattered across too many repositories and too many services.

Moving all the way back to a huge monolith is not the answer either. A very large monolith can create a different context problem.

The architecture style I like here is the self-contained system. You split the application into verticals that include UI, business logic, and database in one place, usually one repository or project. The vertical is large enough to own meaningful behavior but small enough for the agent to reason over.

This also lets different verticals use different stacks where appropriate. Inventory might use one frontend framework while order management uses another. Each vertical can have skills matching its technology and business domain.

If an organization can standardize on one stack, that simplifies the skills. In the real enterprise world, many customers use multiple frontend and backend frameworks, including in-house frameworks, so the harness has to reflect that reality.

Teams And Flow Change Too

Spec-driven development also changes team shape and process.

If use cases are the work items and generation is fast, two-week sprints can become too slow. The work moves toward continuous flow, with use cases tracked more like Kanban items.

Team size can shrink as well. In some modernization work, one or two developers can handle work that previously needed five to seven, although I still prefer two people for knowledge exchange and because working alone can become boring or risky.

The role of requirements engineering grows. If implementation takes minutes or hours, the bottleneck moves left. Product owners and requirements engineers have more work to do because the quality of the spec determines the quality of the generated system.

That is one of the largest shifts. AI accelerates coding first, but the sustainable advantage comes when requirements, specs, tests, and architecture improve too.

Determinism Comes From The Harness

The goal is not magical generation. The goal is near-deterministic generation from good specs and a good harness.

I recommend not letting AI create the whole project from scratch. Use the official project generator, framework CLI, or current stack tooling. If you are using Spring Boot, start with the proper Spring Boot setup. Save tokens and avoid stale defaults.

Then define the rules. Do not put everything into one giant instruction file. Larger system prompts can create more hallucination. Put architecture guidance, stack-specific rules, and workflows into the right skills and context. Use MCP servers for large documentation sources such as in-house frameworks.

If you do that well, you can delete generated code, run the process again, and get roughly the same result. The source of truth is the spec plus the harness, not one lucky output.

That does not remove bugs. It gives you a better way to decide where to add review, tests, and reflection. The code is still there because today we need source code for review and history. In the future, programming languages may even become more AI-friendly and less human-readable, but for now we work with the tools we have.

The main lesson is that specs are not enough on their own. They need the harness: skills, guardrails, architecture, tests, MCP servers, and risk-based review. But when that system exists, system use cases give agents a durable, business-readable way to generate software.

The full version of this argument was presented at AI Native DevCon London. To go deeper, watch the full recording.

COPY & SHARE

Simon Martinelli

Simon Martinelli

Simon Martinelli is a Java Champion, Vaadin Champion, and Oracle ACE Pro, with over three decades of experience as a software architect, developer, consultant, and trainer. As the owner of Martinelli LLC, he specializes in optimizing full-stack development with Java using AI and has a deep focus on modern architectures and software modernization. He frequently shares his expertise by speaking at international conferences, writing articles, and maintaining his blog, Keep IT Simple: https://martinelli.ch. His passion for teaching is reflected in his work as a lecturer at two universities in Switzerland.

READING

·

0%

IN THIS POST

Use This Talk As Agent ContextWhy A Process-Centric Approach?Why System Use Cases?Greenfield And Modernization Need Different FlowsReviews Should Be Based On RiskPetClinic Shows The PatternArchitecture Has To Fit AI WorkTeams And Flow Change TooDeterminism Comes From The Harness

COPY & SHARE

Simon Martinelli

Simon Martinelli

Simon Martinelli is a Java Champion, Vaadin Champion, and Oracle ACE Pro, with over three decades of experience as a software architect, developer, consultant, and trainer. As the owner of Martinelli LLC, he specializes in optimizing full-stack development with Java using AI and has a deep focus on modern architectures and software modernization. He frequently shares his expertise by speaking at international conferences, writing articles, and maintaining his blog, Keep IT Simple: https://martinelli.ch. His passion for teaching is reflected in his work as a lecturer at two universities in Switzerland.

Spec-Driven Development Works When Specs Are System Use Cases