how it works
Multi-agent systems: how they actually work, and when they're worth it
“Agents” are everywhere this year. But one agent running one task and a multi-agent system coordinating several of them in parallel are very different things. Here’s how the second one actually works.
What is an AI agent?
Start with the base unit.
An AI agent is a language model that doesn’t just answer — it acts: it calls tools, makes decisions based on their output, and iterates until it reaches a goal.
The core pattern is a loop:
Input → LLM → Action (tool, search, write) → Observation → LLM → ...
A single Claude agent with a browser, an API and file access can already do a lot. But it hits hard limits:
- Limited context: one LLM has one context window. Make the task long enough and it starts losing information.
- Speed: it works sequentially — one action at a time.
- Specialisation: one prompt can’t be a sharp researcher, a precise copywriter and a data analyst at the same time.
That’s where multi-agent systems come in.
What is a multi-agent system?
A multi-agent system is an architecture where several distinct agents collaborate on a complex task. Each one has a role, its own tools, and often its own context.
The most useful analogy is a team:
- An orchestrator (the manager) takes the overall goal and breaks it into sub-tasks.
- Specialised agents (the workers) run those sub-tasks on their own.
- Results come back, and the orchestrator decides the next step.
A concrete example — a system that produces a market report:
Orchestrator
├── Research agent → scrape + summarise recent news
├── Data agent → dataset analysis, charts
├── Writer agent → report draft
└── Review agent → fact-check, style, length
They run in parallel wherever they can. The orchestrator coordinates, evaluates each output, and either approves it or sends it back for revision.
The main architectural patterns
1. Orchestrator + workers (the most common)
A central agent splits the task, delegates to specialists, and aggregates the results.
Use it for: complex tasks with parallel phases — content pipelines, data analysis, multi-source research.
2. Sequential pipeline
Agents hand output down a chain. Agent A’s output is agent B’s input.
Research → Synthesis → Drafting → Review → Publishing
Use it for: editorial processes, workflows with strict dependencies.
3. Peer-to-peer network (experimental)
Agents talk to each other directly, with no central orchestrator. Any agent can ask another for help.
Use it for: debugging scenarios, cross-validation systems.
4. Human-in-the-loop
At every critical point, a person approves before the system continues. Non-negotiable for irreversible actions: publishing, sending email, payments.
How I do it: every system I run has a Telegram checkpoint before any public action.
The 2026 toolset
Claude as the orchestrator
The tool_use API makes Claude a good fit for the orchestrator role. Claude can:
- Define sub-tasks as structured JSON
- Call external APIs as tools
- Read their output and decide the next step
- Handle loops and retries on its own
Projects and longer-lived memory let agents keep context across sessions — a big change from 12 months ago.
n8n as the infrastructure
n8n has become the de facto orchestration layer for self-hosted multi-agent systems. With the AI Agent nodes in v2.0 you can:
- Build an agent loop with a single node
- Chain several agents as sub-workflows
- Handle errors on dedicated branches
- Pass context between agents through session variables
A typical multi-agent n8n workflow:
Webhook
└── Orchestrator agent (Claude)
├── Sub-workflow: Web research
├── Sub-workflow: Data analysis
└── Sub-workflow: Drafting
└── Telegram: approve?
└── Publish
LangGraph and AutoGen
If you build agents in Python, LangGraph (from LangChain) and AutoGen (Microsoft) are the most mature frameworks:
- LangGraph: models the flow as a directed graph. Strong for complex loops and conditional branching.
- AutoGen: built around multi-agent conversation, with native support for groups of agents talking to each other.
Both work with Claude through the Anthropic SDK.
Costs and practical trade-offs
Multi-agent systems multiply LLM costs. If each agent makes 5 API calls per task and you have 4 agents, that’s 20 calls per run.
How to keep it in check:
-
Use smaller models for simple tasks. The web-research agent doesn’t need Opus — Sonnet is plenty for classification and summarisation. Haiku is cheaper still, but only worth it at high volume on a task you have already tested on real cases.
-
Cache shared work. If two agents need the same document, cache the read once instead of fetching it twice.
-
Cap iterations. Every agent loop needs an explicit
max_iterations— otherwise you risk an infinite loop with a growing bill. -
Log everything, structured. In n8n, add a logging node after each agent to record input, output, duration and tokens used.
The real gain: parallelism
The most underrated advantage of multi-agent systems isn’t specialisation — it’s parallelism.
A single agent researching 10 sources takes 10 sequential steps. Ten parallel workers do the same job in one step of wall-clock time.
For anything that means synthesising many sources — market research, competitor monitoring, legal review — the time saved is large.
When NOT to use a multi-agent system
Not every problem needs one. Complexity you don’t need gives you a fragile, expensive system.
Stay with a single agent if:
- The task has fewer than 5–6 steps
- There’s no real parallelism to exploit
- The LLM budget is tight
- You’re still prototyping
Move to multi-agent when:
- One agent’s context isn’t enough
- You have genuinely parallel tasks
- You need specialisation and error isolation
- The system is going to production at high volume
How I approach it at Mimir Lab
Every system I build starts as a single agent. It becomes multi-agent only when one of these triggers fires:
- The context window becomes the bottleneck
- Run time is unacceptable and the work can be parallelised
- Output quality measurably improves with specialisation
The Mimir Command Center, for example, runs three agents: Researcher (scraping and synthesis), Writer (LinkedIn draft), Reviewer (style and fact-checking). It used to be one agent — splitting it improved output quality by roughly 40%.
The extra complexity pays off, but only when the numbers say so.
Want to see the Command Center’s n8n workflow? It’s in the Workflows section. Upcoming posts go into the implementation, with code and templates.