The Agent Trio: How I Run Three AI Agents on One Shared Memory
Three agent planes, one shared memory, zero self-modification: the shape that keeps a multi-agent system debuggable as it grows.
Run three specialised agent planes, build (Claude Code), schedule and delivery (Hermes), execution and reach (OpenClaw), over one shared typed memory, and let nothing modify itself without human sign-off. The split keeps failures debuggable: where a job breaks tells you which plane to inspect.
The Agent Trio: Three Agents, One Memory
Most people who “run AI agents” run one mega-agent that schedules, executes, reaches the network, remembers, and notifies, and then can’t tell you why it broke. Or they run several agents that each keep their own context, so nothing one learns is available to the others.
I run three agent planes. They share one memory, and none of them can change itself without my sign-off. That shape, specialise the planes, share the memory, gate the changes, is the whole design. Here’s why each part earns its place.
The shape, in one line each
- Claude Code, the build / config plane. Where I design, code, configure, and decide, with a human in the loop.
- Hermes, the delivery / schedule plane. It owns the clock (cron) and the phone channel (Telegram). It’s how the system reaches me.
- OpenClaw, the execution / reach plane. Privileged shell, browser, secrets, and the network paths the others can’t take. It’s how the system reaches the world.
Underneath all three: one shared memory, a single store every plane reads and writes. Not three siloed contexts. One.
Claude Code = BUILD / CONFIG , design, code, approve (you in the loop)
Hermes = DELIVERY / SCHEDULE, cron + Telegram out
OpenClaw = EXECUTION / REACH , privileged exec, browser, secrets, network
└────────── one shared memory spans all three ──────────┘
The split was not a taste decision. It followed walls. Hermes was the only plane with a Telegram channel and a native scheduler. OpenClaw was the only one with a network path to my feed reader, a browser, and my secrets. Claude Code was the only one with the full repo and me in the loop. Every job I wanted to run needed at least two of those, so I stopped asking one agent to do everything and wrote the capabilities down as a table. The planes are just that table made real.
Decision 1: one memory, three readers
The planes are specialised, but they are not strangers. Every one of them reads and writes the same store (memory_search over a shared database). A decision made while building in Claude Code is visible to Hermes when it schedules, and to OpenClaw when it executes.
This is the part most multi-agent setups get wrong. They give each agent its own memory, and then spend all their time passing context between agents by hand. Shared memory means the context is the substrate, no hand-off, no drift, no “which agent knew what.”
The cost: you have to be disciplined about what goes in. One shared memory with junk in it poisons all three planes at once. So the store is typed, segments, summaries, decisions, not a dumping ground.
Decision 2: nothing modifies itself
The second rule is stricter than it sounds: automation may propose changes, but it never applies them.
A daily job can analyse my sessions and say “here’s a new skill you should add” or “this agent definition has a gap.” It writes that as a proposal. It does not touch a single agent or skill definition. Nothing changes until I read the proposal and say yes, at which point Claude Code (the build plane) makes the actual change and commits it.
I hold this line because I have watched the default. OpenClaw’s skill workshop ships set to apply its own proposals. In August five of them were applied with nobody asked, and a weekly unattended review of the skill collection failed six weeks in a row without anyone noticing. Nothing exploded, and that is the problem. A system that quietly rewrites its own definitions leaves you with a setup you can no longer predict, and you find out months later when something breaks and no one can say why it changed. I switched it to propose-only with approval pending. Automation can spot a gap and write it down. A human decides. If I cannot tell you why an agent changed, I cannot trust what it does next.
The boundary that makes it debuggable
The real payoff of the split shows up the moment something breaks. Take a daily digest job:
OpenClaw → reaches the source, fetches + reads the data (EXECUTION / REACH)
↓ hands the finished payload to
Hermes → fires on schedule, delivers to Telegram (DELIVERY / SCHEDULE)
If the digest is late, it’s a Hermes problem, schedule or delivery. If the digest is empty, it’s an OpenClaw problem, fetch or reach. If the digest is wrong, it’s a build problem, I fix the logic in Claude Code. The boundary tells you where to look before you’ve opened a single log.
That is the entire argument for specialising the planes: separation of concerns you can debug six months later.
My first RSS digest was a Hermes cron job that called OpenClaw over MCP to fetch the feeds, read them and write the digest, all in one agent loop. The fetch kept timing out and no digest arrived for days. A digest that never arrives gives you no error to read, only an empty phone. The boundary would have told me in one look: an empty digest is a fetch problem, fetching belongs to the reach plane, and I had put an agent loop between the scheduler and the reach. On 12 June I replaced it with a plain script that fetches and ranks on a schedule, then one model call to write the summary, then Hermes delivers it. No agent loop in the middle.
When to reach for which
| You need to… | Plane |
|---|---|
| Design, code, configure, or decide | Claude Code |
| Run something on a schedule | Hermes |
| Send anything to my phone | Hermes (only) |
| Run a privileged command, browser task, or read a secret | OpenClaw |
| Reach a source behind a network the scheduler can’t | OpenClaw |
| Recall anything any plane has seen | shared memory (any plane) |
Building or deciding → Claude Code. Scheduling or announcing → Hermes. Doing or reaching → OpenClaw. Remembering → all three, one store.
The takeaway
The lesson isn’t “use these three tools.” It’s the shape:
- Specialise the planes so each does one job and a failure tells you where to look.
- Share one memory so context is the substrate, not a hand-off.
- Gate every change so nothing rewrites itself behind your back.
A mega-agent gives you power and no legibility. Isolated agents give you legibility and no shared context. Three specialised planes on one gated memory give you both, and that’s the only reason the system still makes sense to me months after I built it.
Two things. First, I let the boundaries blur. In August I gave Hermes the vault mounted read-write and a local shell, so two planes could now write my files, and capability stopped telling me who owned what. I had to write ownership down separately: Hermes owns scheduling writes, the other plane owns the rest. I would draw ownership first and grant access second. Second, I built the planes before I built one place to see them. A job that wrote its result only to a log file on one machine was invisible. Now every scheduled run records to one database and shows up in one dashboard, and I would have required that before the third job, not after dozens.
Part of how I run a personal multi-agent system. The two execution planes get a deep-dive of their own; the shared-memory layer gets another.