Agent Knowledge Management Standards: Storage and Retrieval vs. Context Engineering
Who this is forEngineers and no-code builders who keep persistent file- or wiki-based knowledge for AI agents and want shared vocabulary for storage, retrieval, and context management.
I recently posted a two-axis framework for agent knowledge in a thread: one axis covers storage (how, when, single source of truth, periodic review) and the other covers retrieval (when and how fast). When I checked it against the industry literature, I found that the two axes overlap with established frameworks, but not one-to-one. Anyone building agents on files, wikis, or memory stores runs into the same questions, so a shared vocabulary helps. This article maps the storage and retrieval axes to published standards, shows which parts of context engineering fall outside them, and explains where the information management analogy holds and where it breaks down.
One-line summary
The two-axis framework I posted in the thread covers the persistent knowledge layer: files, wikis, and memory stores. The industry standard, LangChain’s four context engineering strategies (Write, Select, Compress, and Isolate), covers the context window at every step. The two frames work at different levels, so they do not map one-to-one. Still, Write and Select are the bridge operations between the storage layer and the window layer, so storage roughly corresponds to Write and retrieval roughly corresponds to Select. Compress and Isolate exist only inside the window layer, which is why a storage-centric view does not show them. For storage format, the standard reference is the CoALA memory taxonomy. For digestion, the standard is Letta’s sleep-time compute. For retrieval speed, the standard is progressive disclosure combined with hybrid loading. My intuition that digging far enough leads to management information systems (MIS) holds only for the storage and retrieval axes. These two axes are agent versions of concepts from information management, such as DIKW, the records lifecycle, and MDM/SSOT. Compression and isolation come from a different lineage. They are runtime techniques that respond to the finite context window of an LLM, and they trace back to operating system memory management and cognitive science work on working memory.
Core diagram

Core data
1. The four context engineering strategies (LangChain, accessed July 19, 2026)
LangChain defines context engineering as filling the context window with exactly the right information at each step of an agent’s trajectory. The four strategies are:
| Strategy | Definition | Representative techniques | Mapping to the thread’s two axes |
|---|---|---|---|
| Write | Store outside the context window | Scratchpad, cross-session memory (ChatGPT, Cursor, and Windsurf auto-memory; Reflexion) | ≈ Storage axis (bridge from window to store) |
| Select | Pull into the window when needed | Memory selection, tool retrieval (RAG improves tool selection accuracy 3x), grep/AST/knowledge graphs | ≈ Retrieval axis (bridge from store to window) |
| Compress | Keep only the needed tokens in the window | Summarization (Claude Code auto-compact), trimming | No counterpart; window layer only |
| Isolate | Split context to prevent interference | Multi-agent (separate window per sub-agent), sandboxing, state schema | No counterpart; window layer only |
This is not an exact one-to-one term mapping. The two axes take a lifecycle view of persistent storage, while the four strategies take a per-step view of window management. Three points do not line up cleanly. First, the single source of truth (SSOT) on the storage axis is a governance concern (master data management, or MDM), not a Write operation. Second, digestion is consolidation, which corresponds to sleep-time compute, not Write. Third, the initial load on the retrieval axis is pre-loading during prompt construction, not Select, which is runtime retrieval.
2. Industry standards for the storage axis
How to store (md, json, or db) maps to the CoALA memory taxonomy. CoALA, from Princeton in 2023 (arXiv:2309.02427), draws on cognitive science (Tulving 1972, Squire 1987, Baddeley and Hitch 1974) to define four memory types. This taxonomy is the shared foundation for Letta, Mem0, and LangChain. In our MacBook agent environment, the types map as follows:
| Memory type | What it holds | Counterpart in our MacBook agent environment |
|---|---|---|
| Working (task) | Current task context and intermediate reasoning | The context window itself, TaskList, scratchpad |
| Episodic (episode) | What happened when X was tried | Session transcripts, sessions/, work logs, git history |
| Semantic (meaning) | Factual and conceptual knowledge | wiki/, researches/, memory/ fact files, environment facts in CLAUDE.md |
| Procedural (procedure) | Skills and procedures | skills/SKILL.md, hooks, rules/, scripts |
Format standard. Google’s OKF v0.1 (June 12, 2026) formalizes an open spec that requires Markdown with YAML frontmatter, a mandatory type field, and reserved files index.md and log.md. The spec names as its lineage Karpathy’s LLM wiki, the AGENTS.md and CLAUDE.md conventions, and metadata-as-code. Anthropic also states that folder hierarchy, naming conventions, and timestamps are all signals. In other words, file names and directory structure are themselves the search interface.
When to store (hooks and handoffs) maps to structured note-taking. Anthropic names structured note-taking as one of three techniques for long-running tasks. The agent periodically writes notes to memory outside the context window and reloads them later. The Claude platform’s memory tool, which is file-based and in public beta, productizes this pattern at the API level. Combining memory with context editing reportedly produced a 39% performance gain in an internal agentic search evaluation.
Periodic review (digest and backfill) maps to sleep-time compute. Letta introduced sleep-time compute in 2025. During idle time between user interactions, a background agent rewrites memory. It consolidates fragmented memories, finds patterns across sessions, removes duplicates, and archives stale information. Letta reports about 5x less test-time compute to reach the same accuracy on benchmarks, and a 2.5x lower average cost for repeated queries over the same context. These figures are Letta’s own reporting, and independent reproduction has not been confirmed, so cite the source when quoting them. Our weekly compile and monthly lint, run by a cron job on a Mac mini, implement this pattern directly.
3. Industry standards for the retrieval axis
When to retrieve (init, plan, and explore) maps to hybrid loading. Anthropic recommends a hybrid approach. Information that is always useful, such as CLAUDE.md, is pre-loaded at init. Everything else is retrieved just in time with glob and grep. The principle is to avoid loading all data into context and instead keep lightweight identifiers, such as file paths and links, and load content dynamically at runtime.
How fast (frontmatter and index) maps to progressive disclosure. Progressive disclosure is a design principle in which the agent gradually discovers relevant context through exploration. Agent Skills is the canonical implementation. It loads in three stages: one line of SKILL.md frontmatter, then the body, then reference files. The index.md file is specified in OKF as a reserved filename for progressive disclosure, described as a table of contents. Karpathy, OKF, and our wiki-query skill share the same conclusion: at personal scale, about 400,000 words, agentic search that reads the index and then reads only the candidate pages is enough, and embedding-based RAG is not required.
4. Does the information management lineage hold? Only for storage and retrieval
The storage and retrieval axes (the persistent knowledge layer) do have an information management lineage.
- DIKW pyramid (Ackoff and Zeleny, 1980s): data, information, knowledge, and wisdom. In an agent harness, the raw/ folder (source material), the wiki/ and researches/ folders (refined knowledge), and the output/ and decision layers correspond directly to these levels.
- SSOT and MDM: Keep master data in one place and everything else as a pointer. Our structure, where AGENT_GUIDE.md is the SSOT and CLAUDE.local.md is a pointer, is a typical MDM pattern.
- Records lifecycle: Creation, use, then appraisal and disposition (archiving). The monthly governance decisions to archive, defer, or reactivate content are records appraisal in practice.
- The OKF spec itself names data-team metadata-as-code in its lineage.
Compression and isolation (the window layer) do not belong to the information management lineage. Both respond to a constraint specific to LLMs: the context window is finite, and performance caps out as it grows (context rot). The closest lineages are not MIS. Compress relates to memory hierarchy and cache management in operating systems and computer architecture, and to the limits of working memory in cognitive science (Baddeley and Hitch). Isolate relates to process isolation and sandboxing in operating systems. This is the same reasoning CoALA uses when it takes its memory taxonomy from cognitive psychology (Tulving and Squire) rather than from management science.
Insights
-
The two axes and the industry’s four strategies are frames at different levels, so their overlap is the interesting part. The two axes take the view of persistent storage, and the four strategies take the view of the window. Write and Select are bridge operations between the two layers, so storage roughly matches Write and retrieval roughly matches Select. Compress and Isolate exist only inside the window and are structurally invisible to someone who thinks in terms of storage. The accurate hook for content is not “I found two strategies and the standard has four.” It is “someone looking at storage sees only two. The other two live inside the window.”
-
The answer to the storage format question (md versus json versus db) depends on the memory type. Using the CoALA taxonomy: semantic memory fits Markdown with frontmatter, which is readable by both humans and agents. Episodic memory fits append-only logs such as JSONL or session files. Procedural memory fits executable files such as skills, hooks, and scripts. Working memory is not stored at all; it stays in the context window. The single-format debate arises from not distinguishing memory types, so it is a false problem.
-
At personal scale, a file system with an index is settling in as the standard, rather than embedding-based RAG. Karpathy, OKF, and Anthropic (file-based memory tool and progressive disclosure) point in the same direction. The main value of embeddings, finding content even when it is phrased differently, is largely replaced when an agent retries grep with synonyms and with Korean and English spellings. Retries are nearly free locally. Grep actually breaks down in three situations: large external corpora with uncontrolled naming, such as crawled or transcribed data; runtime products such as chatbots that cannot tolerate the latency of multi-round exploration; and indexes that become too large to read, beyond roughly 400,000 words. Personal knowledge layers fall into none of these. When search starts to miss, the first response is to improve naming, frontmatter, and the index, not to introduce RAG.
-
Periodic digest review is compute economics, not a nice-to-have. The 5x and 2.5x figures for sleep-time compute rest on the claim that precomputing consolidated memory structurally lowers the cost at query time. Our weekly compile, monthly lint, and monthly governance already implement this, so it can serve as a real-world case study in the content.
-
The MIS intuition is a valid content direction only for the storage and retrieval axes. The angle that 1970s information management concepts (DIKW, the records lifecycle, and MDM) are being revived in the persistent knowledge layers of agents is a differentiated topic for an AI education creator. However, grouping compression and isolation under information management is an overgeneralization. Those belong to operating system and cognitive science lineages, and should be presented separately for accuracy.
Bottom line
The storage and retrieval axes correspond to established standards: the CoALA memory taxonomy for storage format, sleep-time compute for periodic review, and progressive disclosure with hybrid loading for retrieval speed. They also inherit concepts from information management, including DIKW, the records lifecycle, and MDM/SSOT. Context engineering adds two strategies, Compress and Isolate, that have no counterpart in a storage-centric view. These address the finite context window and come from operating system and cognitive science lineages instead. The evidence supports a clear division: treat storage and retrieval as an information management problem, and treat compression and isolation as window-level runtime techniques.
Sources
Primary sources
- Effective context engineering for AI agents, Anthropic Engineering: compaction, structured note-taking, the three sub-agent patterns, just-in-time retrieval, progressive disclosure, the hybrid CLAUDE.md approach, and filesystem metadata as a signal (accessed July 19, 2026)
- Context Engineering for Agents, LangChain Blog: definition and techniques for the four strategies, Write, Select, Compress, and Isolate (accessed July 19, 2026)
- Cognitive Architectures for Language Agents (CoALA), arXiv:2309.02427: the working, episodic, semantic, and procedural memory taxonomy (accessed July 19, 2026)
- Sleep-time Compute, Letta Blog and Sleep-time agents, Letta Docs: background memory consolidation and the 5x and 2.5x figures, which are Letta’s own reporting; independent reproduction is not confirmed, so cite the source when quoting (accessed July 19, 2026)
- Equipping agents for the real world with Agent Skills, Anthropic: progressive disclosure as the core design principle of Skills (accessed July 19, 2026)
Secondary sources
- Anthropic’s Memory Tool Reframes How We Build Agents, S3P Studios: context for the 39% figure on memory tool and context editing (accessed July 19, 2026)
- Types of AI Agent Memory, Atlan: summary of CoALA as the shared taxonomy for Letta, Mem0, and LangChain (accessed July 19, 2026)
- DIKW pyramid, Wikipedia and ISKO Encyclopedia: DIKW hierarchy: the DIKW lineage (accessed July 19, 2026)
Frequently asked questions
- Does the storage and retrieval framing cover the whole agent context problem?
- No. Storage and retrieval cover persistent knowledge layers. LangChain's Compress and Isolate strategies manage the context window within each step, and a storage-centric view has no counterpart for them.
- Which memory types should use markdown versus append-only logs?
- Semantic memory fits markdown with frontmatter, and episodic memory fits append-only logs such as JSONL or session files. Procedural memory fits executable files like skills, hooks, and scripts, and working memory stays in the context window instead of being stored.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›