0%
⚡ SustainSys AI Academy · Module 2

Multi-Agent Orchestration
with LangChain & MCP

Learn how to build specialist AI agents, give them tools (skills), and wire them together using an orchestrator that routes tasks intelligently — all in Python, step by step.

Python 3.11 LangChain LangGraph OpenAI GPT-4o-mini Gmail API OpenWeatherMap MCP Pattern OAuth 2.0

What Are We Building?

In this guide you will build a multi-agent system — three specialist AI agents, each with their own skills, coordinated by a single master orchestrator. By the end you will understand how real-world AI products like AutoGPT, Devin, and enterprise AI assistants work under the hood.

🌤️
Weather Agent

Calls OpenWeatherMap API to report live temperature, humidity, and conditions for any city.

📧
Gmail Agent

Authenticates with Gmail via OAuth and reads your unread inbox, summarising emails naturally.

🔢
Arithmetic Agent

Performs multi-step calculations — add, subtract, multiply, divide — using dedicated tool functions.

🎛️
Orchestrator

The master agent. Reads your query, decides which specialist(s) to call, and unifies results.

🎯 Learning Goals

After this guide you will understand: what a Skill is, what an Agent is, how the ReAct reasoning loop works, how to wrap agents as tools for orchestration, and what MCP means in practice.

Core Concepts — Three Things to Understand First

1. What is a Skill?

A Skill is a plain Python function decorated with @tool from LangChain. That decorator registers the function so an LLM Agent can discover and call it automatically.

Normal function   →   only YOU can call it in code
@tool function    →   the LLM Agent can call it too

def add(a, b):          ← only callable by your code
    return a + b

@tool
def add(a: float, b: float) -> float:    ← agent can call this!
    """Add two numbers."""
    return a + b

The docstring is critical — the LLM reads it to decide when to use the tool. Clear docstrings = smart tool selection. Vague docstrings = confused agent.

2. What is an Agent?

An Agent is an LLM brain connected to a set of tools, running in a reasoning loop. The loop is called ReAct (Reason + Act):

01
Think

LLM reads your query and decides: "which tool do I need?"

02
Act

LLM calls the chosen tool with the right arguments.

03
Observe

Tool returns a result. LLM reads the output.

04
Repeat or Answer

LLM decides: "do I need another tool, or can I answer now?" Loop continues until done.

3. What is the Orchestrator / MCP?

The orchestrator is itself an agent — but instead of calling weather or maths tools, its tools are the other agents. MCP (Model Context Protocol) is the pattern/standard for exposing these agent-as-tool endpoints so any LLM client can discover and call them.

User query
    ↓
Orchestrator Agent
    ↓ (reads query, decides which specialists needed)
    ├── weather_agent_tool("What's the weather in Tokyo?")
    │       ↓ calls Weather Agent → get_weather skill → API → result
    ├── arithmetic_agent_tool("What is 42 × 7?")
    │       ↓ calls Arithmetic Agent → multiply skill → result
    └── Combines both results → single unified response to user
💡 Key Insight

The orchestrator treats each specialist agent as just another @tool. This is the same pattern all the way up — you could even have an orchestrator of orchestrators. This is how large AI systems scale.

Project Structure

Before writing any code, create this folder layout. Understanding why it's structured this way is as important as the code itself.

multi_agent_project/
├── skills/ # Reusable @tool functions (the building blocks)
│ ├── __init__.py
│ ├── arithmetic_skill.py # add, subtract, multiply, divide
│ ├── weather_skill.py # get_weather via OpenWeatherMap API
│ └── gmail_skill.py # check_new_emails via Gmail OAuth
├── agents/ # One agent per domain, each uses its skills
│ ├── __init__.py
│ ├── arithmetic_agent.py
│ ├── weather_agent.py
│ └── gmail_agent.py
├── orchestrator/ # Master router — uses agents as tools
│ ├── __init__.py
│ └── mcp_orchestrator.py
├── .env # API keys (never commit this!)
├── credentials.json # Gmail OAuth credentials from Google Cloud
├── token.json # Auto-generated after first Gmail login
├── main.py # Entry point — runs all test queries
└── requirements.txt

Create the structure

Terminal — run from your project root
mkdir multi_agent_project && cd multi_agent_project
mkdir agents skills orchestrator
touch agents/__init__.py skills/__init__.py orchestrator/__init__.py
touch main.py requirements.txt .env

requirements.txt

requirements.txt
langchain
langchain-community
langchain-openai
langgraph
openai
requests
google-auth
google-auth-oauthlib
google-auth-httplib2
google-api-python-client
python-dotenv
Install
pip install -r requirements.txt

.env file

.env
OPENAI_API_KEY=your_openai_key_here
OPENWEATHER_API_KEY=your_openweathermap_key_here
⚠️ Never Commit .env

Add .env and token.json to your .gitignore. These files contain private API keys and OAuth tokens.

🐍 Python Version

Use Python 3.11. LangChain's Pydantic V1 dependency breaks on Python 3.14+. Create your venv with: ~/.pyenv/versions/3.11.9/bin/python -m venv .venv311

Building Skills — The @tool Decorator

Skills are the atomic building blocks. Each skill is a focused Python function that does one job well. We decorate them with @tool so LangChain agents can discover and call them.

Skill 1 — Arithmetic

No API needed — perfect for learning the @tool pattern. Notice each operation is a separate tool so the agent can choose the right one.

skills/arithmetic_skill.py
from langchain_core.tools import tool

# Each operation is its own @tool
# Why? The agent needs to choose WHICH operation to use.
# One big function would confuse the agent.

@tool
def add(a: float, b: float) -> float:
    """Add two numbers together. Use this for addition or sum calculations."""
    return a + b

@tool
def subtract(a: float, b: float) -> float:
    """Subtract b from a. Use this for subtraction or difference calculations."""
    return a - b

@tool
def multiply(a: float, b: float) -> float:
    """Multiply two numbers. Use this for multiplication or product calculations."""
    return a * b

@tool
def divide(a: float, b: float) -> float:
    """Divide a by b. Use this for division calculations."""
    if b == 0:
        return "Error: cannot divide by zero"
    return a / b

if __name__ == "__main__":
    print(add.invoke({"a": 10, "b": 5}))       # 15.0
    print(multiply.invoke({"a": 4, "b": 6}))   # 24.0

Skill 2 — Weather

Calls the free OpenWeatherMap REST API and returns a human-readable weather string.

skills/weather_skill.py
import requests
import os
from dotenv import load_dotenv
from langchain_core.tools import tool

load_dotenv()

@tool
def get_weather(city: str) -> str:
    """Get the current weather for a given city.
    Returns temperature, humidity, and weather conditions."""

    api_key = os.getenv("OPENWEATHER_API_KEY")
    url = f"http://api.openweathermap.org/data/2.5/weather?q={city}&appid={api_key}&units=metric"
    response = requests.get(url)

    if response.status_code != 200:
        return f"Could not fetch weather for {city}."

    data = response.json()
    temp        = data["main"]["temp"]
    description = data["weather"][0]["description"]
    humidity    = data["main"]["humidity"]
    city_name   = data["name"]

    return f"Weather in {city_name}: {description}, Temp: {temp}°C, Humidity: {humidity}%"

Skill 3 — Gmail

Uses Google's OAuth 2.0 flow. First run opens a browser for login; after that token.json is reused silently.

skills/gmail_skill.py
import os
from dotenv import load_dotenv
from langchain_core.tools import tool
from google.auth.transport.requests import Request
from google.oauth2.credentials import Credentials
from google_auth_oauthlib.flow import InstalledAppFlow
from googleapiclient.discovery import build

load_dotenv()
SCOPES = ["https://www.googleapis.com/auth/gmail.readonly"]

def get_gmail_service():
    """OAuth helper — handles first-run browser login and token refresh."""
    creds = None
    if os.path.exists("token.json"):
        creds = Credentials.from_authorized_user_file("token.json", SCOPES)
    if not creds or not creds.valid:
        if creds and creds.expired and creds.refresh_token:
            creds.refresh(Request())
        else:
            flow = InstalledAppFlow.from_client_secrets_file("credentials.json", SCOPES)
            creds = flow.run_local_server(port=0)
        with open("token.json", "w") as token:
            token.write(creds.to_json())
    return build("gmail", "v1", credentials=creds)

@tool
def check_new_emails(max_results: int = 5) -> str:
    """Check Gmail inbox for the most recent unread emails.
    Returns sender, subject, and a short snippet for each email."""
    service = get_gmail_service()
    results = service.users().messages().list(
        userId="me", labelIds=["INBOX", "UNREAD"], maxResults=max_results
    ).execute()
    messages = results.get("messages", [])
    if not messages:
        return "No new unread emails found."
    summaries = []
    for msg in messages:
        detail = service.users().messages().get(
            userId="me", id=msg["id"], format="metadata",
            metadataHeaders=["From", "Subject"]
        ).execute()
        headers = detail["payload"]["headers"]
        snippet = detail.get("snippet", "")
        sender  = next((h["value"] for h in headers if h["name"] == "From"), "Unknown")
        subject = next((h["value"] for h in headers if h["name"] == "Subject"), "No Subject")
        summaries.append(f"From: {sender}\nSubject: {subject}\nPreview: {snippet[:100]}")
    return "\n\n---\n\n".join(summaries)
📋 Gmail Setup — Google Cloud Console

Before running gmail_skill.py: (1) Create a project at console.cloud.google.com, (2) Enable Gmail API, (3) Create OAuth 2.0 Desktop credentials → download as credentials.json, (4) Add your Gmail address as a Test User under APIs & Services → OAuth consent screen → Audience.

Building the Three Agents

Now we wire skills into agents. Every agent follows the same 4-step pattern: tools list → LLM → create_react_agent → public run function. Learn this pattern once and you can build any agent.

💡 The Universal Agent Pattern

Import skills → tools = [skill1, skill2] → llm = ChatOpenAI(...) → agent = create_react_agent(model, tools, prompt) → def run_agent(query). Every agent in this project follows this exact structure.

Agent 1 — Arithmetic Agent

agents/arithmetic_agent.py
import sys, os, warnings
warnings.filterwarnings("ignore")
# Add project root to Python path so 'skills' module is findable
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))

from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from skills.arithmetic_skill import add, subtract, multiply, divide

load_dotenv()

# Step 1: tools — the skills this agent can use
tools = [add, subtract, multiply, divide]

# Step 2: LLM brain
llm = ChatOpenAI(model="gpt-4o-mini")

# Step 3: create_react_agent wires LLM + tools + reasoning loop
agent = create_react_agent(
    model=llm,
    tools=tools,
    prompt="You are a helpful maths assistant. Use your tools to perform arithmetic calculations."
)

# Step 4: public function — other files call this
def run_arithmetic_agent(query: str) -> str:
    result = agent.invoke({"messages": [{"role": "user", "content": query}]})
    return result["messages"][-1].content

if __name__ == "__main__":
    print(run_arithmetic_agent("What is 42 multiplied by 7, then subtract 100?"))
    # → 42 × 7 = 294, 294 − 100 = 194

Agent 2 — Weather Agent

Same pattern — only the imported skill and prompt change.

agents/weather_agent.py
import sys, os, warnings
warnings.filterwarnings("ignore")
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))

from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from skills.weather_skill import get_weather

load_dotenv()

tools = [get_weather]
llm = ChatOpenAI(model="gpt-4o-mini")
agent = create_react_agent(
    model=llm, tools=tools,
    prompt="""You are a helpful weather assistant.
    Use get_weather to fetch real data. Always mention temperature and conditions."""
)

def run_weather_agent(query: str) -> str:
    result = agent.invoke({"messages": [{"role": "user", "content": query}]})
    return result["messages"][-1].content

if __name__ == "__main__":
    print(run_weather_agent("What is the weather like in Tokyo?"))

Agent 3 — Gmail Agent

agents/gmail_agent.py
import sys, os, warnings
warnings.filterwarnings("ignore")
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))

from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from skills.gmail_skill import check_new_emails

load_dotenv()

tools = [check_new_emails]
llm = ChatOpenAI(model="gpt-4o-mini")
agent = create_react_agent(
    model=llm, tools=tools,
    prompt="""You are a helpful email assistant.
    Use check_new_emails to fetch emails. Summarise clearly — who sent it and what it's about."""
)

def run_gmail_agent(query: str) -> str:
    result = agent.invoke({"messages": [{"role": "user", "content": query}]})
    return result["messages"][-1].content

if __name__ == "__main__":
    print(run_gmail_agent("Do I have any new emails? Give me a summary."))
🔑 Always Run from Project Root

Run python agents/arithmetic_agent.py from multi_agent_project/, not from inside agents/. Python resolves imports relative to where you run from. The sys.path.insert lines handle this automatically too.

The Orchestrator — Agents as Tools

This is the most powerful concept in this guide. The orchestrator is an agent whose tools are other agents. Each specialist agent is wrapped in a @tool decorator so the master LLM can route tasks to the right specialist automatically.

                    ┌─────────────────────────────────┐
                    │       ORCHESTRATOR AGENT         │
                    │   LLM reads query and decides    │
                    │   which agent-tools to call      │
                    └────────┬──────────┬──────────────┘
                             │          │
              ┌──────────────┘          └──────────────┐
              ▼                                        ▼
  ┌───────────────────┐                    ┌───────────────────┐
  │  weather_agent    │                    │ arithmetic_agent  │
  │  _tool(@tool)     │                    │  _tool(@tool)     │
  └────────┬──────────┘                    └────────┬──────────┘
           │                                        │
           ▼                                        ▼
  ┌───────────────────┐                    ┌───────────────────┐
  │  run_weather_     │                    │ run_arithmetic_   │
  │  agent(query)     │                    │ agent(query)      │
  └────────┬──────────┘                    └────────┬──────────┘
           │                                        │
           ▼                                        ▼
  ┌───────────────────┐                    ┌───────────────────┐
  │  get_weather      │                    │ add/subtract/     │
  │  @tool skill      │                    │ multiply/divide   │
  └───────────────────┘                    └───────────────────┘
orchestrator/mcp_orchestrator.py
import sys, os, warnings
warnings.filterwarnings("ignore")
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))

from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool
from langgraph.prebuilt import create_react_agent

# Import each specialist agent's run function
from agents.weather_agent import run_weather_agent
from agents.arithmetic_agent import run_arithmetic_agent
from agents.gmail_agent import run_gmail_agent

load_dotenv()

# ── KEY CONCEPT: wrap each agent as a @tool ───────────────────────
# The orchestrator's LLM reads these docstrings to decide
# which specialist to call — same as how agents choose skills!

@tool
def weather_agent_tool(query: str) -> str:
    """Use this for ANY question about weather, temperature,
    or climate conditions in any city or location."""
    return run_weather_agent(query)

@tool
def arithmetic_agent_tool(query: str) -> str:
    """Use this for ANY mathematical calculations including
    addition, subtraction, multiplication, division, or
    multi-step arithmetic problems."""
    return run_arithmetic_agent(query)

@tool
def gmail_agent_tool(query: str) -> str:
    """Use this to check Gmail inbox, read new emails,
    or summarise recent email messages."""
    return run_gmail_agent(query)

# ── Build the orchestrator agent ─────────────────────────────────
llm = ChatOpenAI(model="gpt-4o-mini")
tools = [weather_agent_tool, arithmetic_agent_tool, gmail_agent_tool]

orchestrator = create_react_agent(
    model=llm,
    tools=tools,
    prompt="""You are a master orchestrator with 3 specialist agents:
    1. weather_agent_tool    — for weather queries
    2. arithmetic_agent_tool — for maths calculations
    3. gmail_agent_tool      — for email checking
    Route each query to the right agent(s). If a query needs
    multiple agents, call all of them and combine the results."""
)

def run_orchestrator(query: str) -> str:
    result = orchestrator.invoke({
        "messages": [{"role": "user", "content": query}]
    })
    return result["messages"][-1].content

main.py — Wiring Everything Together

The entry point runs four test queries — single-agent, single-agent, single-agent, and finally a multi-agent query that calls two specialists in one shot.

main.py
import warnings
warnings.filterwarnings("ignore")

from orchestrator.mcp_orchestrator import run_orchestrator

print("=" * 50)
print("  Multi-Agent Orchestrator Ready!")
print("=" * 50)

# Test 1 — single agent: weather
print("\n🌤️  Test 1: Weather query")
print(run_orchestrator("What is the weather in Paris?"))

# Test 2 — single agent: maths
print("\n🔢  Test 2: Maths query")
print(run_orchestrator("What is 156 divided by 12, then multiply by 5?"))

# Test 3 — single agent: email
print("\n📧  Test 3: Email query")
print(run_orchestrator("Check my emails and give me a brief summary"))

# Test 4 — MULTI-AGENT: orchestrator calls weather + arithmetic
print("\n🚀  Test 4: Multi-agent query")
print(run_orchestrator("What's the weather in London AND what is 99 multiplied by 99?"))
Run it
cd multi_agent_project
python main.py

Expected Output

Terminal output
==================================================
  Multi-Agent Orchestrator Ready!
==================================================

🌤️  Test 1: Weather query
The current weather in Paris is partly cloudy with a temperature of 10.87°C and humidity of 66%.

🔢  Test 2: Maths query
The result of 156 divided by 12 is 13, and 13 multiplied by 5 is 65.

📧  Test 3: Email query
Here is a summary of your recent emails:
1. AI Engineering — Day 13: AI Model Abstraction Layer guide
2. ServiceNow — Community event on March 24
3. The New Stack — Kubernetes webinar invitation

🚀  Test 4: Multi-agent query (TWO agents called!)
The current weather in London is 10.28°C with few clouds.
Additionally, 99 multiplied by 99 equals 9801.
🎯 Watch Test 4 Carefully

Test 4 is the magic moment — one query triggers two separate agents running sequentially. The orchestrator called weather_agent_tool AND arithmetic_agent_tool, got both results, and combined them into one clean response. That is multi-agent orchestration.

The MCP Pattern — What It Really Means

MCP stands for Model Context Protocol. It's a standard developed by Anthropic for how AI models discover and call external tools and agents. Understanding it conceptually is as important as the code.

Three Layers of the Pattern

1

Skills Layer — @tool functions

Raw Python functions decorated with @tool. Each does one atomic job. This is where real work happens — API calls, calculations, file reads.

2

Agent Layer — LLM + skills + ReAct loop

A specialist LLM that owns a set of skills and reasons about which to call. Each agent is an expert in its domain.

3

Orchestrator Layer — agents as tools

A master LLM that treats each agent as a tool. Routes queries to specialists and assembles unified responses. This is the MCP layer.

Why This Architecture Scales

Principle What it means in practice
Separation of concerns Each agent only knows its own domain. Adding a new agent doesn't touch existing ones.
Composability Orchestrator can call 1, 2, or all 3 agents per query. Combinations are unlimited.
Replaceability Swap gpt-4o-mini for Claude or Llama in any single agent without touching others.
Testability Each skill and agent can be tested in isolation with if __name__ == "__main__".
Extensibility Add a Calendar Agent or Search Agent by repeating the same 4-step pattern.
🔌 Production MCP

In production, the orchestrator runs as an HTTP server (using Anthropic's mcp Python SDK or FastAPI). Agents are exposed as endpoints. Any MCP-compatible client — Claude Desktop, VS Code extensions, custom apps — can then discover and call your agents automatically over the network.

What You Built

✓ @tool Skill

A decorated Python function the LLM can call. Docstring = when to use it. Type hints = what to pass in.

✓ ReAct Agent

LLM + tools + reasoning loop. Think → Act → Observe → Repeat until answer is ready.

✓ create_react_agent

LangGraph's modern function that wires model, tools, and prompt into a runnable agent.

✓ Orchestrator

An agent whose tools are other agents. Routes tasks to specialists and combines results.

✓ OAuth 2.0

Secure Google login flow. First run opens browser → token.json saved → reused silently.

✓ MCP Pattern

Skill → Agent → Orchestrator. Three layers. Each wraps the one below as a callable tool.

✓ sys.path trick

sys.path.insert(0, parent_dir) makes Python find your modules regardless of where you run from.

✓ venv + pyenv

Always isolate dependencies. Python 3.11 is the sweet spot for LangChain stability today.

🎓 You Are Now a Builder

Most people who use AI assistants have no idea what's underneath. You've built the architecture that powers every multi-agent AI product — skills, specialist agents, orchestration, and the MCP routing pattern. This knowledge doesn't go out of date when frameworks change, because you understand the fundamentals.