“No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.”
In his essay Agents Don’t Need Memory. They Need Documentation., developer Kevin Liao argues that the current approach to memory in the AI agent tooling ecosystem is fundamentally flawed. While many solutions layer background daemons, rerankers, and automated summarization routines on top of RAG databases, real-world development teams still struggle with the management nightmares created by black-box memory stores.
The RAG Lottery Cannot Fix Five Core Engineering Traps
Most agent memory plugins on the market follow an identical blueprint: parse session transcripts, generate memory snippets, insert them into a vector database, and retrieve the top five matches on every prompt to inject into context. Liao characterizes this setup as a lottery over RAG snippets. In chunked vector records, the original intent behind the code and the state of the surrounding environment are easily lost.
Figure: Isolated memory snippets that are difficult to trace. Source: liao.gg
Similarity search merely calculates the distance between two snippets in embedding space; it cannot determine which rule is actually valid in the current codebase. Treating past records as current truth poses severe risks in fast-evolving repositories. If authentication logic changes, obsolete authentication snippets lingering in the database will feed stale instructions directly to the agent. Trapped in a black box of 10,000 SQLite embeddings, developers have no easy way to discern which memories are outdated or which have never even been retrieved. Even when agents are handed a global search tool, they lack the meta-awareness to know what knowledge they are missing or when to search for it.
Reshaping the Workflow in Four Distinct Stages
To escape the pitfalls of black-box memory, Liao proposes Document-based Memory. This paradigm strips away background daemons and vector interfaces in favor of structured, plain Markdown documentation. A single AGENTS.md file is rarely enough to sustain complex projects; the system must encompass project instructions, technical specifications, architectural decision records, and exploratory research.
Figure: The shift from “prompt → build → forget” to “prompt → consult → build → update”. Source: liao.gg
This shift transforms the development cycle from a unidirectional prompt → build → forget loop into prompt → consult → build → update. Upon receiving a task, the agent first consults the documentation for the relevant module; after completing the build, it writes back newly established constraints and updated states into the project repository. Memory is no longer an opaque retrieval database bolted onto the side, but a transparent workspace that can be read, edited, and shared across the team.
Production Practices Settle on a Three-Tier Directory Layout
As developers move toward plain-text memory systems, practical directory conventions have begun to crystallize across the community. Rather than dumping every conversation into a flat folder, teams are adopting a tiered architecture: .agents/plans/ for execution roadmaps, .agents/notes/ for active investigation logs, and .agents/knowledge/ for validated, persistent conclusions.
Under this taxonomy, notes represent scratchpad thinking, and only insights that have been tested and verified are promoted to knowledge. When knowledge files within a specific domain expand, developers introduce local INDEX.md catalogs to guide navigation. Once the storage medium returns to plain text, decades of established principles in directory design and engineering governance can directly govern AI agents.
Community Debate: Can Text Protocols Stop Code Hallucinations?
Some developers remain skeptical about whether plain-text documentation alone provides sufficient guardrails. In production, practitioners report that even with explicit prompt instructions like “use jq instead of writing a python script to parse json,” models frequently disregard directives and churn out ad-hoc Python scripts when handling complex JSON structures.
For probabilistic LLMs, text instructions have inherent limits. These engineers advocate pairing written protocols with compiler-level enforcement. By turning linting rules and static analysis failures into structured compiler feedback returned to the model, teams can enforce hard boundaries and steer agents reliably back onto the intended path.
Engineering Commonsense Drives Toolchains Back to Plain Text
Whether maintaining static Markdown indexes or enforcing rigid linter checks, both approaches converge on a fundamental objective: modern developer toolchains must remain transparent, deterministic, and auditable. Many memory plugins reduce context management to a simplistic retrieval recall problem, attempting to paper over state management flaws by piling on more moving parts.
The true foundation of software engineering knowledge has always been human-readable, version-controlled plain text. What agents need is not a probabilistic memory black box, but a living documentation workspace they can consult before work begins and update when the job is done.
Reference Links:
- Agents Don’t Need Memory. They Need Documentation.
- Discussion on Hacker News
- Operator Memory Repository