Building a forgetting-capable memory bank for AI Agents
Title: Building a Memory Bank That Forgets for AI Agents
I tinkered with an AI Agent memory bank over the weekend and hit quite a few pitfalls. An AI Agent is basically an assistant that looks up info, remembers things, and executes tasks on its own. It started when I saw an article on Machine Learning Mastery discussing agent memory design—how AI Agents should design their memory. The article said current memory systems are fragile, like removing the forgetting function from the human brain. My testing showed: in demos, it remembers everything; in real business, it gets dragged down by outdated facts.
I digested the core of the article into three things: Short-term memory shouldn't be too long (don't stuff the context); Context is the range of conversation AI can view at once. Long-term memory shouldn't be too messy; retrieval must be verifiable. From an asset allocation perspective, a memory bank isn't a warehouse; it's a balance sheet. Every record has costs: storage, retrieval, misuse, expiration. Over these past few days, I built a minimal version: just using local folders and an assistant that can read files, getting the flow working first.
Preparation phase.
1. Open the terminal (command line window on your computer), type mkdir ~/agent-memory, then type cd ~/agent-memory. You'll see the command prompt shorten, indicating you've entered the directory.
2. Type touch raw.md facts.md decisions.md rules.md to create four empty files. raw.md holds raw conversations, facts.md holds reusable facts, decisions.md holds conclusions, rules.md holds remember/delete rules.
3. Open rules.md in an editor and write three rules: Only record information that affects future actions; Every fact needs a source and validity period; In case of conflict, new entries overwrite old ones, marking old ones as deprecated. Expected result: You have an instruction manual constraining the AI.
Getting started phase.
4. Paste a real conversation into raw.md. Could be meetings, customer service, or project discussions. I've been testing STT (Speech-to-Text) these last two days, dumping records in, and found that if you shove the whole segment into the memory bank, it becomes a mess in three days.
5. Open the input box in Claude Code and paste: Read raw.md, extract no more than 5 reusable facts. Format: Fact, Source, Validity Period, Confidence, Recommended for Writing. Do not extract chitchat or emotions. Confidence means how sure it is. It will return bulleted text formatted in Markdown. Expectation: Not writing directly to the file, but giving you candidates first.
6. Human review. Keep only entries that can change future decisions. For example, "Customer budget cap is within a certain range" is worth remembering. Copy reviewed entries into facts.md.
7. Add status to each fact. Template: - Fact: . ; Source: corresponding paragraph in raw.md; Validity: 30 days; Status: Current. Change to Status: Deprecated after expiration. This action looks dumb, but it's the foundation of the memory system.
8. Before the next conversation, input: Read facts.md and decisions.md, use only entries with Status: Current. Cite sources when referencing facts. Expected result: It no longer stuffs entire history into context, but retrieves as needed.
9. At the end of each task, write final conclusions into decisions.md. Format: Conclusion, Basis, Unconfirmed Items. This step turns chat into an evidence chain.
Pitfall phase.
The first pitfall is treating raw records as memory. Raw records are like surveillance footage; memory is like meeting minutes. Storing everything equals storing nothing.
The second pitfall is no expiration. TTL is the built-in expiration time for each memory.
A memory bank without TTL, decay, or overwrite mechanisms looks impressive in demos but degrades with the speed of change in the real world.
I've had similar feelings since starting to use Google Search these last few days: AI Overviews give summaries, but the original blue links (the actual search result links) require scrolling down to find yourself. Same for memory banks; next to fact cards, you must be able to return to raw.md.
The third pitfall is relying solely on semantic retrieval (searching by meaning, not keywords). Turning sentences into vector coordinates means converting text into strings of numbers and searching by similarity. It only finds similarities; it doesn't judge correctness.
The fourth pitfall is ungraded permissions. When my clients use AI resume tools, mixing HR permissions causes auto-updates to scramble the candidate pool. Same for memory systems: who can write, edit, or deprecate must be defined upfront.
Conclusion here: AI Agent memory isn't about remembering more, but knowing when not to trust itself. From a business model perspective, memory isn't a feature; it's a long-term asset. From a risk-reward ratio, a system that forgets beats one that remembers forever but frequently cites outdated facts. Previously I thought RayNeo iO's value lay in audio data; now my thinking is more specific: the more data, the more boundaries needed, otherwise pollution spreads faster.
Looking ahead, what truly differentiates isn't chatting, but who can distill conversations into auditable, expirable, reusable workflow data. Managed memory services (services where others store your memory) will increase, but the advantage hard for others to copy isn't the ability to store, but the judgment of which information deserves to enter the next decision.
After learning this, try turning facts.md into a small retrieval library: filter by keywords first, then try vector retrieval; or connect decisions.md to your usual terminal so every task completion auto-appends a conclusion. Get the four steps of Recording, Reviewing, Expiring, and Citing running smoothly first.
📌 Compiled from Machine Learning Mastery, original article: https://machinelearningmastery.com/ai-agent-memory-design-what-works-and-what-doesnt/
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier