Introduction
When software engineering emerged in the mid-20th century, shaped by pioneers such as Alan Turing, John von Neumann, Grace Hopper, Edsger Dijkstra, and later Fred Brooks, software was largely about logic, mathematics, and precise instructions for a machine. Programs were smaller, the gap between the implementation and the knowledge required to understand was relatively short, and much of the truth about a system could be found directly in its source code. The code was the documentation.
Then came higher-level programming languages, like COBOL, Lisp, C, and many others. Software became easier to create and much easier to grow. Big tech companies started with a great plan to make money from ever-growing, complicated software. Understanding a system could no longer mean simply reading its source code.
We started documenting the knowledge around the code. Requirements specifications described what a system was supposed to do. Change requests captured how it should evolve. Architecture diagrams explained its structure. UML attempted to provide a common visual language for describing software. Processes such as Waterfall tried to establish an orderly progression from requirements, through design, to implementation. For a while, we believed that the truth about a software system could be captured in specifications and models. The detailed specification and diagrams were the documentation.
Then Agile arrived. One of the best-known statements from the Agile Manifesto is “working software over comprehensive documentation.” It did not say that documentation had no value. The manifesto explicitly acknowledged value in both. But the emphasis changed. Instead of attempting to specify a system completely before building it, teams increasingly relied on working software, short feedback loops, collaboration, and incremental design. Somewhere along the way, however, less comprehensive documentation was too often interpreted as little or no documentation. The source of truth became distributed across code, Jira tickets, pull requests, wiki pages, Slack conversations, and the heads of engineers. The scattered notes and ‘tribal knowledge’ were the documentation.
For a long time, we could make that work. But, today, the scale is different. Modern digital systems can consist of hundreds or thousands of services and hundreds of millions of lines of code, developed over many years by teams that continuously change. No single engineer can fully understand such a system. We navigate it by constructing a local understanding: reading the code, looking at an architecture diagram, searching old tickets, asking another engineer, or simply knowing from experience that this particular component should not be touched that way.
The underlying problem throughout those years was not the documentation itself, but where we believed the source of truth lived, and how someone understood the system enough to make changes safely. Or in other words: where is the context of the system?
For decades, we could afford to leave a surprising amount of that context implicit. In conversations, historical decisions, organizational conventions, and the heads of experienced engineers.
And then we invited AI agents to write the code. Thus, does it still matter to create and maintain documentation in the era where ‘coding is solved’?
Impact of no documentation on software development
The documentation problem existed long before LLMs.
Documentation has always been part of how developers reconstruct a mental model of a system. Its importance to program comprehension and maintenance has been studied for decades, and the documentation was always a part of the overall cost of the system development [3]. Hence, there has always been temptation to neglect documentation, or marginalize at best. From a cost-cutting and financial perspective, it seems reasonable. However, I’m pretty sure that any software engineer who worked on an undocumented system, had a hard time understanding the context. Sure, we can read the code. But, why is it there? Why was this particular solution chosen? Is this strange behaviour intentional or accidental? Which assumptions must remain true? What depends on it? What happens if we change it? When that context is not explicit, developers have to reconstruct it.
There is also the statement ‘Why do we need the documentation? Nobody reads it anyway’. From my experience, if the goal of documentation is just to exist, not to inform, then it is true. But well-structured, concise, scoped documentation, easy to find and search through will become a knowledge base about the system.
The goal of documentation is to inform, and clearly communicate how and why the system works [4]. It should capture the context necessary to understand the system without forcing every engineer to rediscover that context independently.
Take for example the SMD electronics and microcontrollers industry. I cannot imagine working with a microcontroller without its datasheet. I could do the signal reverse engineering, and eventually discover the ‘API’. But why would I want to?
The datasheet defines the contract that allows the user to safely connect and use the hardware. But for some reason, in software engineering the documentation seems like a burden.
Removing documentation does not remove the need for system knowledge. It moves the cost of acquiring that knowledge somewhere else. It is hidden in every question and discovery, every time someone wants to know the system or part of it. The team cannot say ‘here is the knowledge base’, they can eventually say ‘Mark knows the most about this part’.
Every new developer has to reconstruct some of it. Every unfamiliar component requires investigation. Every major change requires rediscovery of assumptions and dependencies. When experienced engineers leave, part of the system's accumulated context may leave with them. And when the reconstructed mental model is wrong, the cost is no longer just engineering time. It becomes a risk.
What happens when system context is missing?
There are no quantitative measures of the impact a lack of documentation has on overall system development [3]. However, based on my experience and observation I’ve pinpointed the following effects of missing context.
There is no clear source of truth
Different people develop different mental models of the same system. Code says one thing, an old wiki page says another, and an engineer remembers that there was an important reason for doing something differently, but nobody remembers exactly what it was.
System risks and bottlenecks become difficult to identify
Architecture is shaped by constraints and trade-offs that are not necessarily visible from individual pieces of code. Without explicit knowledge about those decisions, engineers may not know how a component behaves without understanding why it has certain performance, availability, security, consistency, or scalability characteristics.
High-risk changes lead to decision paralysis
The less we understand about the consequences of a change, the harder it becomes to approve it. A seemingly simple modification may affect an unknown consumer, violate an undocumented invariant, or invalidate an architectural assumption made years ago. Eventually, the safest answer becomes: “Don't touch it.”
Development becomes an exercise in reconstruction
Before implementing a change, engineers have to rediscover enough of the system to determine where the change belongs and what it might affect. Some amount of investigation is inevitable in software engineering. But without documentation, the same architectural and domain knowledge is reconstructed repeatedly by different people. Rediscovering the knowledge is a waste of time and resources.
Business decisions become less informed
Missing system knowledge is not exclusively an engineering problem. Can this system support ten times the traffic? Can we enter another market? Can this component be replaced? How expensive would it be to separate this domain? What would happen if we changed this business process? Without an explicit understanding of boundaries, constraints, dependencies, and architectural decisions, estimates increasingly depend on assumptions.
The gap between engineering and business grows
Business understands the intent and desired outcomes. Engineers understand the implementation. The knowledge connecting the two is where documentation can provide a common language. Without it, both sides can be correct within their own mental models while still misunderstanding each other.
Bad documentation has a cost too
Having documentation is not automatically better than having none.
Documentation that is outdated, contradictory, impossible to navigate, or detached from the system can create false confidence. When there is no documentation, an engineer knows that investigation is necessary. When authoritative-looking documentation is wrong, they may make decisions based on information they assume to be true.
The goal therefore cannot simply be more documentation. The goal is documentation that communicates the relevant system knowledge clearly, has an identifiable scope and source of truth, and evolves together with the system.
Which raises the more important question: What does good documentation actually look like?
What is a good documentation
Software engineering is an evolving field. There is no universally accepted model describing what the documentation of a software system as a whole should look like. Perhaps there shouldn't be.
A small application and a distributed platform consisting of hundreds of services do not need the same shape of documentation. The goal should not be to follow a documentation standard or produce a prescribed set of artifacts. The goal is to preserve enough context for someone unfamiliar with a particular part of the system to understand it and change it safely.
The crucial realization is that documentation is a living and evolving part of the software system. It evolves together with the system and stops evolving only when the system itself stops changing.
Even a system in maintenance mode continues to accumulate knowledge. Dependency upgrades, operational changes, incidents, discovered constraints, deprecations, and security updates can all change what we know about the system and how it should be operated.
What follows is not an attempt to define a golden documentation pattern. It is a structure I have arrived at through working with systems and teams of very different sizes. Sometimes where documentation was a unicorn that everybody believed existed, but nobody had actually seen, and other times where documentation existed in abundance but had become an unsearchable dumping ground.
The goal of this structure is to make the context necessary to understand and change a system explicit, scoped, and discoverable.
Properties of good documentation
A well-defined documentation is characterized by documentation quality attributes[3].
However, we need to be cautious in strictly adhering to all of them, since not all of them are required. It depends on the complexity of the system.
Following is a list of opinionated properties that, I think, may be applied to any software documentation.
Intent
Why does this system/component/service/feature exist?
What business or user problem does it solve? What outcomes is it responsible for?
Intent provides the highest-level context for evaluating whether a proposed change actually belongs in the system.
Quality attributes
What characteristics matter beyond functional correctness?
Availability, performance, scalability, security, consistency, recoverability, maintainability, cost, and other quality attributes frequently explain architectural decisions that cannot be inferred from the code alone. The driving force behind the architecture design decisions.
Boundaries
What does this system own, and what does it not own?
Boundaries define responsibilities between systems, services, domains, teams, and external dependencies.
Example: IoT system owns the algorithms that govern statistics, monitoring, predictions and thresholds. It also may own the hardware and drivers. But it outsources the MQTT server, IdP, payments.
Constraints
What must remain true?
Constraints capture technical, business, regulatory, organizational, and operational restrictions within which a solution must operate.
Example: The financial system operates in a highly regulated industry. The IDP must conform to security standards.
Contracts
How does the system interact with the outside world?
APIs, events, schemas, protocols, compatibility guarantees, and other contracts establish expectations that a local code change must not accidentally violate.
Domain concepts
What does the system do? What is the domain expertise that we are trying to build?
Domain concepts establish the vocabulary, entities, relationships, rules, and invariants necessary to reason about the problem rather than merely about its implementation. This is the ‘core business’ of the system.
Example: Operating systems that govern hardware and software interaction are composed of multiple interconnected domains.
Architecture Decision Records
Why does the system look the way it does? Why are those particular decisions taken, and why do they match the system?
ADRs preserve significant decisions, their context, alternatives, and trade-offs. They are valuable to anyone wanting to get answers on why some parts of the system are working the way they work, or why this particular library or software was selected.
Features
What behaviour does the system currently provide?
Feature documentation connects user or business intent with system behaviour and, ideally, with the domain concepts, contracts, constraints, and architectural decisions involved in implementing it. These artifacts should not exist as independent islands of documentation. They form a graph of system knowledge.
A feature exists within a domain. It operates within boundaries and constraints. It exposes or consumes contracts. Its implementation is shaped by quality attributes and architectural decisions.
That graph defines the context of a change. The graph may be simple with a small impact on the system, or complicated with a long list of impacts, including the high-risk changes.
Bellow is the sample representation of knowledge graph built out of the Feature definition. Notice that all the rectangles (nodes) represent documents and arrows (edges) represent references to the documents. How it is technically done it is up to the team, and depends on the company standards. From my experience markdown format works really well. It is simple, tool agnostic, and easy to track history with any VCS.

Infrastructure and technical stack
How does the system is structured, and deployed? What is the technical stack?
Infrastructure provides high level overview of the whole system. Well defined, allows to quickly familiarize with the system. With included tech stack it provides the blueprint.
Characteristics of good documentation
Regardless of its form, useful documentation has several important characteristics. Following is an opinionated list, based on experience.
- Easy to find. Documentation that exists but cannot be discovered when needed might as well not exist.
- Has single source of truth. The same knowledge should not be independently maintained in five different places, slowly diverging into five different versions of reality.
- Well scoped. It should be clear whether a statement applies to the whole system, a domain, a service, a feature, or a particular interface.
- Referenceable. Concepts, decisions, constraints, contracts, and features should be identifiable and linkable rather than buried inside large documents.
- Collaborative and reviewable. Important changes to our understanding of the system deserve review just as changes to its implementation do.
- Versioned. Historical changes allow to search through the past and figure out reasoning behind.
And above all, it is living. Documentation is not an artifact produced at the beginning of a project and archived when implementation starts. It is part of the system and should evolve together with it. A change to the system can change its code, but it can also change a contract, introduce a domain concept, invalidate an architectural decision, establish a new constraint, or alter an existing feature. If those changes are not reflected in the documentation, then the documentation and the system have diverged. Thus, it is no longer the documentation, but history.
Do not over-engineer the documentation
For a small application, the documentation might mean a README describing its intent, a handful of important constraints, and a few architectural decisions.
For a large distributed system, the same goal may require explicitly defined domains and boundaries, contracts between systems, quality attributes, domain concepts, constraints, features, and hundreds of architectural decisions, organized so that the relevant subset can be discovered when needed. Thus, it is up to us, the system architects / designers / developers, to make a reasonable decision.
Context is king. Structure is queen. Complexity decides how much of each you need.
The [1] publication shows that access to sufficiently rich, relevant context improves repository comprehension by agents. Performance generally improves as more documentation is retrieved. The authors interpret this as deeper documentation inquiry leading to better repository comprehension. However, this does not mean the more, the better.
An AI agent starts with the smallest context necessary to understand the change. It expands the context when the agent discovers a dependency, unanswered question, ambiguity, or missing relationship. The same principle applies to software engineers. When working on a change, a developer does not need to “load” the entire system into their mental context. They need enough context to understand the change, its boundaries, and its consequences.
The structure
Sample structure template that can be used as a baseline. Notice the structure includes DDD aspects. However, as mentioned above, not all elements of documentation must be applied to every software system. Thus, the section DOMAIN / BOUNDED CONTEXT may be completely dropped if not applicable. It might be replaced with description of components/libraries/services or whatever suites best for the particular system.
Sample feature description
How to use documentation in AI era
We now have all this knowledge. How does an agent know which parts it needs for a particular change?
In this context documentation is no longer only something a developer reads when trying to understand the system. It becomes part of the context from which an agent constructs its understanding of the system before making a change.
The agent doesn't need all the context. It needs the right context.
In the publication [1] researchers tested GPT-4.1 and Gemini 2.5 Pro with no documentation and with documentation produced by several approaches.
Without documentation, performance was poor. Across all six metrics, every documentation method improved over the no-documentation baseline, with absolute improvements ranging from roughly 5% to 38% depending on the task and metric. Notice that this publication is based on outdated models, and current situation,as of 2026, might be different.
Evidence from [1] and [2] research suggests that coding agents perform better when they can access relevant, well-scoped system knowledge. Simply adding more documentation does not help, and can even hurt. The challenge is to structure documentation so that an agent can discover the right context for the task without flooding its context window with irrelevant information.
An LLM may be excellent at modifying a method when given the method and an instruction. But real changes often propagate: changing a method signature may require changing callers, constructors, inherited methods, tests, or other files. The repository is interconnected, and it is generally too large to simply place the whole thing into the prompt. With this in mind, the [5] shows that the context takes an important role in understanding the change the AI agent is about to implement. The proposed solution, CodePlan, treats repository-level coding as a planning problem rather than a sequence of isolated prompts. The system maintains a dependency graph and continually updates a plan as the LLM changes the repository. This experiment shows that without temporal context, the model fails to modify code because it doesn't know that an earlier change created a new requirement. Without spatial context, it sometimes invents missing pieces because it doesn't know that the relevant functionality already exists elsewhere.
Missing documentation can create an analogous problem when the missing information is architectural or domain context that cannot be recovered from source-code dependencies. This is especially important for LLMs in automating repository level coding tasks.
CodePlan models relationships between pieces of source code, then follows those relationships to understand the change. I believe that this principle may be applied to system documentation as well. Making explicit relationships gives a consistent knowledge base. The documentation becomes a traversable graph, in the way that AST and Call Graph helps to understand the code.
Validity is not the same as intent.
The code changes might be completely valid, compile, pass the tests, and shipped to production. But does it fully implement the intended change? How can the AI agent know that? As already stated, the source code doesn’t always give the right context, then the missing part is the documentation bringing in the answer to the ‘why’ question.
So how should documentation be used by an AI agent?
There is no single best practice as for now. However, I usually start with a Feature/Change Request/Bug Description in an already established system. The description written with ‘properties of good documentation’ in mind should ground the instructions for AI agent.
For Features including multiple domains, and affecting multiple components my practice is to explicitly ask AI agent for a development plan, and instruct to offload the smaller tasks with a specific context to the subagents.
The following is an example of such a starting point. It provides an entry point into the system knowledge relevant to the change. All relevant system knowledge should be referenced. This example is simplified enough just to give the idea. In real-life production cases the relevant information may be, and should be, more precise and detailed.
Example feature description
Notice that all related components of the Change Request definition should be references to those particular components. For example: Item Collection should be a reference to Item Collection definition that lies somewhere in the Item domain.
With such a definition the agent can then use a workflow similar to following steps.

Notice that context discovery is a loop, not a preprocessing step. Planning or validation may expose missing knowledge, causing the agent to traverse the system documentation again before proceeding.
With this approach, the agent is provided with the right context to make changes. A well-defined change request and relevant system context reduce the amount of information the agent has to infer, and therefore reduce one important source of incorrect assumptions. Every explicit constraint is one less constraint the agent has to invent.
A self check might be performed using the Acceptance Criteria and User Journey. An Intent, Constraints, User Journey should be reflected in the new/updated documentation. Thus closing the change loop.
With all that said we can instruct AI agent to provide the implementation plan with consideration of splitting the independent tasks into subagents. All provided with a specific context. This way the context maybe narrowed down to implementing domain model changes, implementing APIs, or running tests. This means running smaller tasks with smaller contexts but still orchestrated by the main agent.
The 7 principles of context-aware agentic development
Following principles are based on my experience with documentation and working with AI agents.
- Start with intent, not code.
- Documentation is system context.
- The agent needs the right context, not all context.
- Changes must be understood through their relationships.
- Use context to scope, traverse, and plan before implementing.
- Validity is not the same as intent.
- Close the loop: update system knowledge together with the code.
Summary
Software engineering has always depended on an enormous amount of undocumented context. Humans have compensated for it through experience, conversations, organizational knowledge, and mental models built over months or years of working with a system. AI agents make the cost of that missing context much more visible.
They also expose something important about the claim that “coding is solved.”
Writing code was never the whole problem.
Some of the hardest parts of software engineering happen before a line of code is written: discovering what actually needs to be built, understanding the domain, deciding where the boundaries of a system should be, identifying its constraints, and making trade-offs that will shape it for years. The code is the implementation of those decisions. It cannot explain decisions that were never made explicit.
Give an AI agent a well-defined problem, clear boundaries, relevant architectural decisions, domain knowledge, and explicit constraints, and its ability to implement a solution can be remarkable. Give it a vague task and an undocumented system, and it has to fill in the gaps. Human engineers do this too. The difference is that experienced engineers develop a mental model that helps them recognize when something does not make sense. They ask colleagues, search through historical decisions, remember previous incidents, or challenge a requirement. An agent does not automatically possess that institutional memory. When context is missing, it can make a perfectly plausible assumption, and then confidently implement it.
This is where documentation becomes much more relevant in the AI era. It is because we need documentation that captures the context that cannot be reliably recovered from the code: intent, boundaries, constraints, contracts, domain concepts, architectural decisions, and the reasons behind them. It needs to be scoped, discoverable, and maintained as part of the system rather than as an artifact created once and forgotten. In that sense, documentation is becoming part of the software itself. And in fact, it has always been.
If we want AI agents to make “coding is solved” even remotely true, then business owners, product managers, architects, and engineers have a new responsibility. We have to become much better at making the context in which software is built explicit.
The better we describe what we are building, why we are building it, and within which boundaries it must operate, the less an agent has to guess.
And perhaps that is the irony of the AI era: the more capable we make machines at writing code, the more important it becomes for humans to clearly explain what that code is supposed to mean.
References
- Evaluating Repository-level Software Documentation via Question Answering and Feature-Driven Development, Harbin Institute of Technology, Shenzhen, China 2026
- Recent studies: Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
- Garousi et al., Cost, benefits and quality of software development documentation: A systematic mapping, Journal of Systems and Software, 2015.
- A. Forward, Software Documentation – Building and Maintaining Artefacts of Communication, Ottawa‐Carleton Institute for Computer Science, University of Ottawa, 2002
- CodePlan: Repository-level Coding using LLMs and Planning, Microsoft Research, India, 2024
Reviewed by:
Szymon Winiarz
Bartłomiej Turos
Michał Ostruszka
Adam Warski
Marcin Baraniecki




