You stop onboarding the client into a chat and start onboarding them into a folder. Five plain markdown files, a stable structure, and every agent reads them before it does anything. The brief stops being something you paste and becomes something the system loads.
I have been running this pattern across 26 ventures and 45 AI employees. The gap between an agent that is busy and an agent that is useful comes down to almost nothing else.
What does it actually mean to onboard a client into an AI agent operating system?
Most onboarding looks the same everywhere. A Google Doc, a Slack channel, a kickoff call. Then the client sends their first real question, you open a fresh chat window, paste 900 words of background, and explain the business again. The model does great work for about six messages. Then the context window fills up and you are re-explaining who the client's customer is.
Client context engineering is the fix, and it is less glamorous than it sounds. It means the client's business lives in named files with a stable structure: who they sell to, what they charge, what they refuse to do, who signs off, what already failed. An AI agent operating system is just the layer that decides which of those files an agent needs to read to finish a specific task.
Once that exists, trying to teach AI your business stops being a metaphor and becomes a maintenance job. That shift is the whole thing.
What goes in the client context folder?
Five files. Not twenty. I have watched people build 40 page brand bibles that no agent ever reads cleanly, and the agents do worse with them.
- identity.md. Legal name, what they sell in one sentence, three competitors, and what makes them different. Keep it under 400 words. Long identity files make agents vague.
- offer.md. Every product, every price, what is included, what is explicitly out of scope. This is the file agents misread most often, so it gets a table, not prose. Scope creep usually starts with an agent inventing a deliverable.
- people.md. Names, roles, decision rights. Who can approve a $500 spend without asking, who has to see anything public-facing before it ships.
- voice.md. Five sentences the client has actually written, plus five they hate. Real samples, not adjectives. If you write "professional but friendly" you have written nothing.
- guardrails.md. The hard nos. Regulated claims, topics that require a human, anything legal has to review.
Plain markdown, stored in Google Drive or Notion. I only move to Airtable once a client count passes ten, and even then Airtable is just the index that generates the files.
The rule I follow: if an agent cannot act differently because of a line, the line does not belong in the folder.
How do you get a client to fill this out without a 40 minute interview?
You send a form. Tally is what I use, and a shared Notion page with required fields works just as well. Sixteen questions, and I tell them it takes 20 minutes, because it does.
The concrete build looks like this:
- Build the form once, then clone it per client. Tally and Airtable both handle this without code.
- Mark the guardrails and offer fields required, so nobody skips the two that matter most.
- Run a Make or n8n automation that takes each form response and writes the five markdown files into a client folder in Drive, then posts the link in Slack.
When a client does not fill it out, and some will not, I record the kickoff call and drop the transcript into Claude with one prompt: generate these five files as drafts from this transcript. Then I send the drafts back with "correct anything that is wrong." Correcting takes them ten minutes. Writing from scratch takes them forty, which is why they never do it.
How do your agents read that context on every task?
No code AI agents have gotten genuinely good at this in the last year, so there are three levels and you can start at level one today.
Level one: Claude Projects or ChatGPT Projects. Drop the five files into the project knowledge and write a 200 word instruction that maps question types to files. Pricing questions read offer.md. Anything public-facing reads voice.md and guardrails.md. Cheap, fast, and fine for one or two clients.
Level two: a retrieval layer. Chunk the files into Supabase, Pinecone, or a basic vector store inside n8n, and have every agent run a search before it answers. This is the standard setup for an AI agent OS for agencies, because it scales past the point where you can fit everything in a prompt.
Level three: an MCP server with read access to the client folder. Every agent, every tool, same source of truth. No duplicated copies drifting out of sync.
The detail that matters more than the plumbing: every agent run starts with the same three files. Identity, offer, guardrails. Voice and people load only when the task is client-facing. That one rule cuts token cost and stops agents from inventing positioning they were never given.
And context goes stale. That is the failure mode nobody warns you about. A client changes their pricing in March and by June four agents are quoting the old number. I put a 15 minute review on the calendar every Monday, and one agent scans recent client messages for contradictions against the folder. When it finds one, it proposes a diff and I approve it. Takes me a minute.
What changes once this is running?
Onboarding a new client goes from two weeks of scattered answers to an afternoon. The client notices something subtler than speed: the agents already seem to know them. Nobody has to repeat the origin story.
What I care about more is that the work stops being reactive. When the context layer is solid, agents finish things instead of starting things, and the client relationship stops depending on me being in the room. That is the difference between owning a business and being the bottleneck in one.
FAQ
Do I need to write any code? No. Tally, Airtable, Make, and Claude Projects cover levels one and two without a single line. The MCP server at level three is the only technical piece, and plenty of people run for a year without it.
What if the client's business changes every month? Keep identity.md and guardrails.md stable, and let offer.md be the volatile one. Point one agent at it and have it propose updates from the client's own emails, so the edits are proposed rather than remembered.
How do I know the context is actually good? Ask the same question three different ways across three agents. If the answers disagree, your context is inconsistent, and no amount of prompting will fix that. Fix the folder first.
The setup is not complicated, but the order matters, and most people get the order wrong by starting with agents before they have anywhere for the context to live. I wrote the full playbook for this. How to Build Your Own AI Agent Operating System walks you through the exact architecture I use, step by step. You can get it at a.mastermindshq.business/ai-os-book. And if you hit a wall on your own client folder, hit reply and tell me what broke. I read everything.
