You split memory into four stores, and you keep them apart on purpose. Policy is slow, global, and written by you. Client facts are per-account and verified on a schedule. One-off instructions live for one task and then get archived. Working memory is the scratch pad the agent writes to while it runs.
That separation is the core of any AI agent operating system. When all of it sits in one folder, your agent treats a Tuesday note as company law, and you spend your mornings re-explaining your business.
Why Does One Memory Bucket Break Everything?
Because everything in a single store carries the same weight.
You paste in your refund policy. You paste in notes from a call with a client. You write "do not mention the price increase until March" and you drop it next to both. The agent reads it all, and now a note that is true for six weeks sits at the same altitude as a rule that is true forever. Nothing in the file tells it which is which.
There is a second problem, and it bites harder. Context bloat. Every token you load is attention your agent is not spending on the actual task. I watched a founder load 40 pages of client history into every single run. The agent got slower, more expensive, and dumber. It was busy, not useful.
That is the failure I hear about most from founders one to three years in. The agent is not broken. The memory is just undifferentiated.
What Are the Four Layers of AI Agent Memory?
Four layers, four owners, four lifespans.
Policy. Global rules that apply to every client, every task, every time. Tone, escalation triggers, what the agent is never allowed to promise, legal boundaries. Slow changing. Versioned. Owned by you.
Client facts. Everything true about one account. Who they are, what they bought, how they like to be talked to, what annoyed them in March. One file per client, updated on a schedule, with a last-verified date at the top.
One-off instructions. "Use the short version of the proposal this time." "Do not CC Dana." Task scoped. Expires when the task ends.
Working memory. The run log. What the agent tried, what it found, what it handed back. Mostly for you, occasionally for the next run.
I keep policy and client files as plain Markdown in Obsidian, synced to a private GitHub repo so I get version history for free. Retrieval runs through Supabase with pgvector. You do not need that stack to start. Notion plus Claude Projects will carry you a long way, and this is context engineering for founders before it is anything technical.
How Do You Build the Policy Layer So It Stops Drifting?
Write it like a rulebook, not an essay.
One rule per line. Number them. Keep the whole thing under a page, and if it creeps past a page, you have two documents pretending to be one. Mine has 22 rules. Rule 4 is "never quote a delivery date without checking the calendar." Rule 11 is "if a client asks for a discount, reply that you will check and stop there."
Two things make this layer hold. First, a header on every policy file: "These are global rules. They apply to all clients unless a client file explicitly overrides them." That one sentence kills a pile of weird outputs. Second, a version date and a review cadence. I read the whole file on the first Monday of the month. Fifteen minutes. If a rule has not mattered in 90 days, it gets cut or pushed down into a client file.
Tools that work here: Obsidian or Notion for the writing, GitHub for the versioning, and a small n8n workflow that pushes the file into the system prompt at the start of every run. The goal is not to write a manual for a human. It is to teach AI your business once, in a form it can actually check against.
How Do You Keep Client Facts From Eating Your Context Window?
You stop loading all of them.
Every client gets a folder. Inside: facts.md, history.md, open-items.md. The agent retrieves by client ID at the start of a task, not by searching across everything. In Supabase you do this with a metadata filter on client_id. Pinecone does the same thing. If you are running folders in Obsidian, it is just a path.
Then put a last-verified date at the top of every facts file. Anything older than 60 days gets flagged for a human pass. Facts rot. A client changes their billing contact, nobody updates the file, and your agent sends a proposal to someone who left in October.
I run a 20-minute cleanup on Fridays. Five clients, about four minutes each. It is boring, and it is the difference between an agent that knows your business and one that guesses.
How Do You Make One-Off Instructions Expire?
You give them their own file with a clock on it.
Session instructions go into a session.md that lives for one task. Timestamp at the top. When the task closes, an n8n or Zapier trigger moves the file into an archive folder, and nothing in the archive gets loaded again unless you go get it.
Precedence has to be written down, because agentic workflows will hit this conflict within a week. Session overrides client facts. Client facts override policy. That order lets a one-off instruction beat a global rule when it needs to, which is exactly what you want and exactly what will make you nervous the first time it happens.
So you log it. Every time an agent uses a session instruction that contradicts a policy rule, it writes one line to an override log. I read that log weekly. Twice it has caught an agent quietly bending a rule I wrote on purpose, and both times the fix was one sentence in the policy file.
FAQ
Can I run this with just Notion and Claude Projects?
Yes. Three pages: policy, clients, and a session page you clear at the end of each task. Claude Projects will hold them as project knowledge. It starts to strain around 15 or 20 clients, and that is when you move retrieval into Supabase or Pinecone.
How often should client facts be updated?
After every meaningful interaction, plus a scheduled pass every 60 days. The date at the top of the file is what makes it work. Unscheduled memory is how you end up with an agent that is confidently wrong about a client you have known for two years.
What happens when a one-off instruction conflicts with a company policy?
The session layer wins for that task, and the conflict gets logged. The log is the point. You want the override to be possible and visible, not impossible or invisible.
I wrote the full playbook for this, and it walks through the exact architecture I use, step by step. You can get it at https://a.mastermindshq.business/ai-os-book
