Multi-Agent Orchestration
with LangChain & MCP
Learn how to build specialist AI agents, give them tools (skills), and wire them together using an orchestrator that routes tasks intelligently — all in Python, step by step.
What Are We Building?
In this guide you will build a multi-agent system — three specialist AI agents, each with their own skills, coordinated by a single master orchestrator. By the end you will understand how real-world AI products like AutoGPT, Devin, and enterprise AI assistants work under the hood.
Calls OpenWeatherMap API to report live temperature, humidity, and conditions for any city.
Authenticates with Gmail via OAuth and reads your unread inbox, summarising emails naturally.
Performs multi-step calculations — add, subtract, multiply, divide — using dedicated tool functions.
The master agent. Reads your query, decides which specialist(s) to call, and unifies results.
After this guide you will understand: what a Skill is, what an Agent is, how the ReAct reasoning loop works, how to wrap agents as tools for orchestration, and what MCP means in practice.
Core Concepts — Three Things to Understand First
1. What is a Skill?
A Skill is a plain Python function decorated with @tool from LangChain.
That decorator registers the function so an LLM Agent can discover and call it automatically.
Normal function → only YOU can call it in code
@tool function → the LLM Agent can call it too
def add(a, b): ← only callable by your code
return a + b
@tool
def add(a: float, b: float) -> float: ← agent can call this!
"""Add two numbers."""
return a + b
The docstring is critical — the LLM reads it to decide when to use the tool. Clear docstrings = smart tool selection. Vague docstrings = confused agent.
2. What is an Agent?
An Agent is an LLM brain connected to a set of tools, running in a reasoning loop. The loop is called ReAct (Reason + Act):
LLM reads your query and decides: "which tool do I need?"
LLM calls the chosen tool with the right arguments.
Tool returns a result. LLM reads the output.
LLM decides: "do I need another tool, or can I answer now?" Loop continues until done.
3. What is the Orchestrator / MCP?
The orchestrator is itself an agent — but instead of calling weather or maths tools, its tools are the other agents. MCP (Model Context Protocol) is the pattern/standard for exposing these agent-as-tool endpoints so any LLM client can discover and call them.
User query
↓
Orchestrator Agent
↓ (reads query, decides which specialists needed)
├── weather_agent_tool("What's the weather in Tokyo?")
│ ↓ calls Weather Agent → get_weather skill → API → result
├── arithmetic_agent_tool("What is 42 × 7?")
│ ↓ calls Arithmetic Agent → multiply skill → result
└── Combines both results → single unified response to user
The orchestrator treats each specialist agent as just another @tool. This is the same pattern all the way up — you could even have an orchestrator of orchestrators. This is how large AI systems scale.
Project Structure
Before writing any code, create this folder layout. Understanding why it's structured this way is as important as the code itself.
├── skills/ # Reusable @tool functions (the building blocks)
│ ├── __init__.py
│ ├── arithmetic_skill.py # add, subtract, multiply, divide
│ ├── weather_skill.py # get_weather via OpenWeatherMap API
│ └── gmail_skill.py # check_new_emails via Gmail OAuth
├── agents/ # One agent per domain, each uses its skills
│ ├── __init__.py
│ ├── arithmetic_agent.py
│ ├── weather_agent.py
│ └── gmail_agent.py
├── orchestrator/ # Master router — uses agents as tools
│ ├── __init__.py
│ └── mcp_orchestrator.py
├── .env # API keys (never commit this!)
├── credentials.json # Gmail OAuth credentials from Google Cloud
├── token.json # Auto-generated after first Gmail login
├── main.py # Entry point — runs all test queries
└── requirements.txt
Create the structure
Terminal — run from your project rootmkdir multi_agent_project && cd multi_agent_project
mkdir agents skills orchestrator
touch agents/__init__.py skills/__init__.py orchestrator/__init__.py
touch main.py requirements.txt .env
requirements.txt
requirements.txtlangchain
langchain-community
langchain-openai
langgraph
openai
requests
google-auth
google-auth-oauthlib
google-auth-httplib2
google-api-python-client
python-dotenv
Install
pip install -r requirements.txt
.env file
.envOPENAI_API_KEY=your_openai_key_here
OPENWEATHER_API_KEY=your_openweathermap_key_here
Add .env and token.json to your .gitignore. These files contain private API keys and OAuth tokens.
Use Python 3.11. LangChain's Pydantic V1 dependency breaks on Python 3.14+. Create your venv with: ~/.pyenv/versions/3.11.9/bin/python -m venv .venv311
Building Skills — The @tool Decorator
Skills are the atomic building blocks. Each skill is a focused Python function that does one job well. We decorate them with @tool so LangChain agents can discover and call them.
Skill 1 — Arithmetic
No API needed — perfect for learning the @tool pattern. Notice each operation is a separate tool so the agent can choose the right one.
from langchain_core.tools import tool
# Each operation is its own @tool
# Why? The agent needs to choose WHICH operation to use.
# One big function would confuse the agent.
@tool
def add(a: float, b: float) -> float:
"""Add two numbers together. Use this for addition or sum calculations."""
return a + b
@tool
def subtract(a: float, b: float) -> float:
"""Subtract b from a. Use this for subtraction or difference calculations."""
return a - b
@tool
def multiply(a: float, b: float) -> float:
"""Multiply two numbers. Use this for multiplication or product calculations."""
return a * b
@tool
def divide(a: float, b: float) -> float:
"""Divide a by b. Use this for division calculations."""
if b == 0:
return "Error: cannot divide by zero"
return a / b
if __name__ == "__main__":
print(add.invoke({"a": 10, "b": 5})) # 15.0
print(multiply.invoke({"a": 4, "b": 6})) # 24.0
Skill 2 — Weather
Calls the free OpenWeatherMap REST API and returns a human-readable weather string.
skills/weather_skill.pyimport requests
import os
from dotenv import load_dotenv
from langchain_core.tools import tool
load_dotenv()
@tool
def get_weather(city: str) -> str:
"""Get the current weather for a given city.
Returns temperature, humidity, and weather conditions."""
api_key = os.getenv("OPENWEATHER_API_KEY")
url = f"http://api.openweathermap.org/data/2.5/weather?q={city}&appid={api_key}&units=metric"
response = requests.get(url)
if response.status_code != 200:
return f"Could not fetch weather for {city}."
data = response.json()
temp = data["main"]["temp"]
description = data["weather"][0]["description"]
humidity = data["main"]["humidity"]
city_name = data["name"]
return f"Weather in {city_name}: {description}, Temp: {temp}°C, Humidity: {humidity}%"
Skill 3 — Gmail
Uses Google's OAuth 2.0 flow. First run opens a browser for login; after that token.json is reused silently.
import os
from dotenv import load_dotenv
from langchain_core.tools import tool
from google.auth.transport.requests import Request
from google.oauth2.credentials import Credentials
from google_auth_oauthlib.flow import InstalledAppFlow
from googleapiclient.discovery import build
load_dotenv()
SCOPES = ["https://www.googleapis.com/auth/gmail.readonly"]
def get_gmail_service():
"""OAuth helper — handles first-run browser login and token refresh."""
creds = None
if os.path.exists("token.json"):
creds = Credentials.from_authorized_user_file("token.json", SCOPES)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file("credentials.json", SCOPES)
creds = flow.run_local_server(port=0)
with open("token.json", "w") as token:
token.write(creds.to_json())
return build("gmail", "v1", credentials=creds)
@tool
def check_new_emails(max_results: int = 5) -> str:
"""Check Gmail inbox for the most recent unread emails.
Returns sender, subject, and a short snippet for each email."""
service = get_gmail_service()
results = service.users().messages().list(
userId="me", labelIds=["INBOX", "UNREAD"], maxResults=max_results
).execute()
messages = results.get("messages", [])
if not messages:
return "No new unread emails found."
summaries = []
for msg in messages:
detail = service.users().messages().get(
userId="me", id=msg["id"], format="metadata",
metadataHeaders=["From", "Subject"]
).execute()
headers = detail["payload"]["headers"]
snippet = detail.get("snippet", "")
sender = next((h["value"] for h in headers if h["name"] == "From"), "Unknown")
subject = next((h["value"] for h in headers if h["name"] == "Subject"), "No Subject")
summaries.append(f"From: {sender}\nSubject: {subject}\nPreview: {snippet[:100]}")
return "\n\n---\n\n".join(summaries)
Before running gmail_skill.py: (1) Create a project at console.cloud.google.com, (2) Enable Gmail API, (3) Create OAuth 2.0 Desktop credentials → download as credentials.json, (4) Add your Gmail address as a Test User under APIs & Services → OAuth consent screen → Audience.
Building the Three Agents
Now we wire skills into agents. Every agent follows the same 4-step pattern: tools list → LLM → create_react_agent → public run function. Learn this pattern once and you can build any agent.
Import skills → tools = [skill1, skill2] → llm = ChatOpenAI(...) → agent = create_react_agent(model, tools, prompt) → def run_agent(query). Every agent in this project follows this exact structure.
Agent 1 — Arithmetic Agent
agents/arithmetic_agent.pyimport sys, os, warnings
warnings.filterwarnings("ignore")
# Add project root to Python path so 'skills' module is findable
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from skills.arithmetic_skill import add, subtract, multiply, divide
load_dotenv()
# Step 1: tools — the skills this agent can use
tools = [add, subtract, multiply, divide]
# Step 2: LLM brain
llm = ChatOpenAI(model="gpt-4o-mini")
# Step 3: create_react_agent wires LLM + tools + reasoning loop
agent = create_react_agent(
model=llm,
tools=tools,
prompt="You are a helpful maths assistant. Use your tools to perform arithmetic calculations."
)
# Step 4: public function — other files call this
def run_arithmetic_agent(query: str) -> str:
result = agent.invoke({"messages": [{"role": "user", "content": query}]})
return result["messages"][-1].content
if __name__ == "__main__":
print(run_arithmetic_agent("What is 42 multiplied by 7, then subtract 100?"))
# → 42 × 7 = 294, 294 − 100 = 194
Agent 2 — Weather Agent
Same pattern — only the imported skill and prompt change.
agents/weather_agent.pyimport sys, os, warnings
warnings.filterwarnings("ignore")
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from skills.weather_skill import get_weather
load_dotenv()
tools = [get_weather]
llm = ChatOpenAI(model="gpt-4o-mini")
agent = create_react_agent(
model=llm, tools=tools,
prompt="""You are a helpful weather assistant.
Use get_weather to fetch real data. Always mention temperature and conditions."""
)
def run_weather_agent(query: str) -> str:
result = agent.invoke({"messages": [{"role": "user", "content": query}]})
return result["messages"][-1].content
if __name__ == "__main__":
print(run_weather_agent("What is the weather like in Tokyo?"))
Agent 3 — Gmail Agent
agents/gmail_agent.pyimport sys, os, warnings
warnings.filterwarnings("ignore")
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from skills.gmail_skill import check_new_emails
load_dotenv()
tools = [check_new_emails]
llm = ChatOpenAI(model="gpt-4o-mini")
agent = create_react_agent(
model=llm, tools=tools,
prompt="""You are a helpful email assistant.
Use check_new_emails to fetch emails. Summarise clearly — who sent it and what it's about."""
)
def run_gmail_agent(query: str) -> str:
result = agent.invoke({"messages": [{"role": "user", "content": query}]})
return result["messages"][-1].content
if __name__ == "__main__":
print(run_gmail_agent("Do I have any new emails? Give me a summary."))
Run python agents/arithmetic_agent.py from multi_agent_project/, not from inside agents/. Python resolves imports relative to where you run from. The sys.path.insert lines handle this automatically too.
The Orchestrator — Agents as Tools
This is the most powerful concept in this guide. The orchestrator is an agent whose tools are other agents.
Each specialist agent is wrapped in a @tool decorator so the master LLM can route tasks to the right specialist automatically.
┌─────────────────────────────────┐
│ ORCHESTRATOR AGENT │
│ LLM reads query and decides │
│ which agent-tools to call │
└────────┬──────────┬──────────────┘
│ │
┌──────────────┘ └──────────────┐
▼ ▼
┌───────────────────┐ ┌───────────────────┐
│ weather_agent │ │ arithmetic_agent │
│ _tool(@tool) │ │ _tool(@tool) │
└────────┬──────────┘ └────────┬──────────┘
│ │
▼ ▼
┌───────────────────┐ ┌───────────────────┐
│ run_weather_ │ │ run_arithmetic_ │
│ agent(query) │ │ agent(query) │
└────────┬──────────┘ └────────┬──────────┘
│ │
▼ ▼
┌───────────────────┐ ┌───────────────────┐
│ get_weather │ │ add/subtract/ │
│ @tool skill │ │ multiply/divide │
└───────────────────┘ └───────────────────┘
import sys, os, warnings
warnings.filterwarnings("ignore")
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool
from langgraph.prebuilt import create_react_agent
# Import each specialist agent's run function
from agents.weather_agent import run_weather_agent
from agents.arithmetic_agent import run_arithmetic_agent
from agents.gmail_agent import run_gmail_agent
load_dotenv()
# ── KEY CONCEPT: wrap each agent as a @tool ───────────────────────
# The orchestrator's LLM reads these docstrings to decide
# which specialist to call — same as how agents choose skills!
@tool
def weather_agent_tool(query: str) -> str:
"""Use this for ANY question about weather, temperature,
or climate conditions in any city or location."""
return run_weather_agent(query)
@tool
def arithmetic_agent_tool(query: str) -> str:
"""Use this for ANY mathematical calculations including
addition, subtraction, multiplication, division, or
multi-step arithmetic problems."""
return run_arithmetic_agent(query)
@tool
def gmail_agent_tool(query: str) -> str:
"""Use this to check Gmail inbox, read new emails,
or summarise recent email messages."""
return run_gmail_agent(query)
# ── Build the orchestrator agent ─────────────────────────────────
llm = ChatOpenAI(model="gpt-4o-mini")
tools = [weather_agent_tool, arithmetic_agent_tool, gmail_agent_tool]
orchestrator = create_react_agent(
model=llm,
tools=tools,
prompt="""You are a master orchestrator with 3 specialist agents:
1. weather_agent_tool — for weather queries
2. arithmetic_agent_tool — for maths calculations
3. gmail_agent_tool — for email checking
Route each query to the right agent(s). If a query needs
multiple agents, call all of them and combine the results."""
)
def run_orchestrator(query: str) -> str:
result = orchestrator.invoke({
"messages": [{"role": "user", "content": query}]
})
return result["messages"][-1].content
main.py — Wiring Everything Together
The entry point runs four test queries — single-agent, single-agent, single-agent, and finally a multi-agent query that calls two specialists in one shot.
main.pyimport warnings
warnings.filterwarnings("ignore")
from orchestrator.mcp_orchestrator import run_orchestrator
print("=" * 50)
print(" Multi-Agent Orchestrator Ready!")
print("=" * 50)
# Test 1 — single agent: weather
print("\n🌤️ Test 1: Weather query")
print(run_orchestrator("What is the weather in Paris?"))
# Test 2 — single agent: maths
print("\n🔢 Test 2: Maths query")
print(run_orchestrator("What is 156 divided by 12, then multiply by 5?"))
# Test 3 — single agent: email
print("\n📧 Test 3: Email query")
print(run_orchestrator("Check my emails and give me a brief summary"))
# Test 4 — MULTI-AGENT: orchestrator calls weather + arithmetic
print("\n🚀 Test 4: Multi-agent query")
print(run_orchestrator("What's the weather in London AND what is 99 multiplied by 99?"))
Run it
cd multi_agent_project
python main.py
Expected Output
Terminal output==================================================
Multi-Agent Orchestrator Ready!
==================================================
🌤️ Test 1: Weather query
The current weather in Paris is partly cloudy with a temperature of 10.87°C and humidity of 66%.
🔢 Test 2: Maths query
The result of 156 divided by 12 is 13, and 13 multiplied by 5 is 65.
📧 Test 3: Email query
Here is a summary of your recent emails:
1. AI Engineering — Day 13: AI Model Abstraction Layer guide
2. ServiceNow — Community event on March 24
3. The New Stack — Kubernetes webinar invitation
🚀 Test 4: Multi-agent query (TWO agents called!)
The current weather in London is 10.28°C with few clouds.
Additionally, 99 multiplied by 99 equals 9801.
Test 4 is the magic moment — one query triggers two separate agents running sequentially. The orchestrator called weather_agent_tool AND arithmetic_agent_tool, got both results, and combined them into one clean response. That is multi-agent orchestration.
The MCP Pattern — What It Really Means
MCP stands for Model Context Protocol. It's a standard developed by Anthropic for how AI models discover and call external tools and agents. Understanding it conceptually is as important as the code.
Three Layers of the Pattern
Skills Layer — @tool functions
Raw Python functions decorated with @tool. Each does one atomic job. This is where real work happens — API calls, calculations, file reads.
Agent Layer — LLM + skills + ReAct loop
A specialist LLM that owns a set of skills and reasons about which to call. Each agent is an expert in its domain.
Orchestrator Layer — agents as tools
A master LLM that treats each agent as a tool. Routes queries to specialists and assembles unified responses. This is the MCP layer.
Why This Architecture Scales
| Principle | What it means in practice |
|---|---|
| Separation of concerns | Each agent only knows its own domain. Adding a new agent doesn't touch existing ones. |
| Composability | Orchestrator can call 1, 2, or all 3 agents per query. Combinations are unlimited. |
| Replaceability | Swap gpt-4o-mini for Claude or Llama in any single agent without touching others. |
| Testability | Each skill and agent can be tested in isolation with if __name__ == "__main__". |
| Extensibility | Add a Calendar Agent or Search Agent by repeating the same 4-step pattern. |
In production, the orchestrator runs as an HTTP server (using Anthropic's mcp Python SDK or FastAPI). Agents are exposed as endpoints. Any MCP-compatible client — Claude Desktop, VS Code extensions, custom apps — can then discover and call your agents automatically over the network.
What You Built
A decorated Python function the LLM can call. Docstring = when to use it. Type hints = what to pass in.
LLM + tools + reasoning loop. Think → Act → Observe → Repeat until answer is ready.
LangGraph's modern function that wires model, tools, and prompt into a runnable agent.
An agent whose tools are other agents. Routes tasks to specialists and combines results.
Secure Google login flow. First run opens browser → token.json saved → reused silently.
Skill → Agent → Orchestrator. Three layers. Each wraps the one below as a callable tool.
sys.path.insert(0, parent_dir) makes Python find your modules regardless of where you run from.
Always isolate dependencies. Python 3.11 is the sweet spot for LangChain stability today.
Most people who use AI assistants have no idea what's underneath. You've built the architecture that powers every multi-agent AI product — skills, specialist agents, orchestration, and the MCP routing pattern. This knowledge doesn't go out of date when frameworks change, because you understand the fundamentals.
Where to Go Next
Immediate Upgrades to This Project
- Add a web search agent using the free DuckDuckGo API or SerpAPI
- Add a calendar agent using the Google Calendar API (same OAuth pattern as Gmail)
- Add memory so the orchestrator remembers context across queries using LangGraph's checkpointing
- Build a Flask or FastAPI HTTP server around the orchestrator to expose it as a real MCP endpoint
- Add a simple chat UI in Streamlit or plain HTML so non-developers can use it
Natural Next Steps in Learning
- LangGraph deep dive — state machines, conditional edges, parallel agent execution
- RAG (Retrieval Augmented Generation) — give your agents access to your own documents
- Anthropic's MCP SDK — the official Python library for production MCP servers
- LlamaIndex — connecting agents to large document collections with semantic search
- ChromaDB or Qdrant — vector databases for agent long-term memory at scale
- Multi-agent debate patterns — two agents that argue and synthesise better answers
Common Questions at This Stage
Can I use a local LLM instead of OpenAI?
Yes. Replace ChatOpenAI(model="gpt-4o-mini") with ChatOllama(model="llama3.1:8b") from langchain_community.chat_models. Only the LLM line changes — skills and agents stay identical.
How do I add a new agent?
1. Create skills/new_skill.py with @tool functions. 2. Create agents/new_agent.py with the 4-step pattern. 3. Import and wrap it as a @tool in the orchestrator. Done.
What's the difference between LangChain and LangGraph?
LangChain provides the building blocks (@tool, ChatOpenAI, prompts). LangGraph is a layer on top that adds stateful agent graphs, loops, and the create_react_agent we used.