0%
SustainSys AI Academy · Student Guide

Building a Local AI Agent
From Scratch

A complete, hands-on guide to understanding AI agents, MCP, orchestration, tools, and memory — by building a working research assistant on your own machine. No paid APIs. No black boxes.

Python 3.10+ Ollama · llama3.1:8b MCP Protocol ReAct Pattern Open-Meteo API Wikipedia API 9 Stages

What is an AI Agent?

Most people's first experience with AI is a chatbot — you ask a question, it answers. That's a single LLM call. An AI agent is something fundamentally different: it's an LLM that can act, not just respond.

The critical difference is the loop. A plain LLM call is one-shot — in, out, done. An agent loops: it thinks, acts, observes what happened, thinks again, acts again — until the task is complete.

Plain LLM Call
  • One question, one answer
  • No tools or real-world access
  • Can't check facts or do maths
  • Hallucinates when uncertain
  • Forgets everything instantly
AI Agent
  • Loops until task is done
  • Uses tools: weather, search, files
  • Verifies with real-world calls
  • Admits uncertainty, looks it up
  • Maintains conversation memory

A Live Example

Ask a plain LLM: "What's the weather in Tokyo and should I pack a coat?" — it will fabricate an answer or admit it can't access real-time data.

Ask an agent the same question:

1
Thinks

"To answer this I need real weather data. I'll use the get_weather tool."

2
Acts

Calls get_weather("Tokyo") via MCP → gets back live temperature, wind, conditions.

3
Observes

Sees: "8°C, partly cloudy, wind 12 mph"

4
Answers

"It's 8°C and cloudy in Tokyo right now — yes, definitely pack a coat."

💡 The One-Line Definition

An AI agent = LLM + loop + tools. The LLM decides what to do. The loop keeps running until it's done. The tools give it real-world capabilities. That's it.

What Makes a Good Agent?

  • A clear system prompt defining role, tools, and output format
  • A reliable decision loop that terminates gracefully
  • Well-described tools so the LLM knows when to use them
  • Proper memory so it doesn't forget context mid-task
  • Solid error handling so one tool failure doesn't crash everything

What is MCP?

MCP stands for Model Context Protocol — an open standard created by Anthropic that defines how AI applications discover and use tools hosted on external servers.

Think of it like USB. Before USB, every device needed its own special cable. USB standardised the connection so any device works with any port. MCP does the same for AI tools and agents.

The Problem MCP Solves

Imagine building 5 different AI applications — each needing weather data, calendar access, and search. Without MCP, you copy tools into every project. With MCP, tools live on one server and every application connects to it.

Without MCP
  • Copy tools into every project
  • Fix a bug? Fix it 5 times
  • Every agent has its own tool API
  • No standardised discovery
  • Tight coupling everywhere
With MCP
  • Tools live on one server
  • Fix once, all agents benefit
  • Standard protocol for all
  • Agents discover tools dynamically
  • Plug-and-play architecture

The Three MCP Endpoints

Every MCP server exposes three standard endpoints — this is the protocol:

The MCP Protocol
# 1. Health check — "Are you alive?"
GET  /health
# → {"status": "ok", "server": "My Tool Server", "version": "2.0"}

# 2. Tool discovery — "What can you do?"
GET  /tools
# → {"tools": [{"name": "calculate", "description": "..."}], "count": 7}

# 3. Tool execution — "Do this thing"
POST /tools/call
# Body:   {"tool": "get_weather", "input": "London"}
# → {"result": "Weather in London: 14°C, Partly cloudy", "success": true}

Client / Server Architecture


  ┌──────────────────────────────────┐
  │         MCP SERVER               │
  │      mcp_server.py               │
  │      Flask on :5001              │
  │                                  │
  │  Tools hosted here:              │
  │  • calculate()                   │
  │  • get_weather()                 │
  │  • search_wikipedia()            │
  │  • save_note() / read_notes()    │
  │  • get_current_time()            │
  │  • summarise_text()              │
  └───────────────┬──────────────────┘
                  │  HTTP / JSON
                  │  POST /tools/call
  ┌───────────────▼──────────────────┐
  │         MCP CLIENT               │
  │      mcp_client.py               │
  │                                  │
  │  connect()       health check    │
  │  discover_tools()  list tools    │
  │  call_tool()     run a tool      │
  │                                  │
  │  Used by agent.py to act         │
  └──────────────────────────────────┘
      
🌐 Real-World MCP Today

Claude Desktop, Cursor IDE, and many other AI tools already support MCP natively. Build an MCP server and any of those tools can connect to it immediately — your tools become a plugin for the entire AI ecosystem.

MCP Server — The Core Pattern

src/mcp_server.py — simplified
from flask import Flask, request, jsonify

app = Flask(__name__)

# 1. Define tools as plain Python functions
def calculate(expression: str) -> str:
    return str(eval(expression))

# 2. Register in a dictionary
TOOLS = {
    "calculate": {"function": calculate, "description": "Evaluate a math expression"}
}

# 3. Expose via HTTP
@app.route("/tools", methods=["GET"])
def list_tools():
    return jsonify({"tools": [{"name": k, "description": v["description"]} for k,v in TOOLS.items()]})

@app.route("/tools/call", methods=["POST"])
def call_tool():
    d = request.get_json()
    result = TOOLS[d["tool"]]["function"](d["input"])
    return jsonify({"result": result, "success": True})

System Architecture

Here is the complete picture of the system and how every file connects:


  ┌──────────────────────────────────────────────────────────┐
  │                  RESEARCH ASSISTANT                      │
  │                                                          │
  │  ┌──────────┐     ┌───────────────────────────────┐     │
  │  │ main.py  │────►│        agent.py               │     │
  │  │ CLI loop │     │   Orchestrator · ReAct loop   │     │
  │  └──────────┘     └──────┬───────────┬────────────┘     │
  │                          │           │                  │
  │               ┌──────────┘           └────────────┐     │
  │               │                                   │     │
  │          ┌────▼────┐  ┌──────────┐  ┌────────────▼─┐   │
  │          │ llm.py  │  │memory.py │  │mcp_client.py │   │
  │          │ Ollama  │  │short +   │  │HTTP client   │   │
  │          │ :11434  │  │long term │  └──────┬───────┘   │
  │          └─────────┘  └──────────┘         │           │
  │                                            │ HTTP       │
  │                                            ▼           │
  │                              ┌─────────────────────┐   │
  │                              │   mcp_server.py     │   │
  │                              │   Flask  :5001      │   │
  │                              ├─────────────────────┤   │
  │                              │ calculate()         │   │
  │                              │ get_weather()       │   │
  │                              │ search_wikipedia()  │   │
  │                              │ save_note()         │   │
  │                              │ read_notes()        │   │
  │                              │ get_current_time()  │   │
  │                              │ summarise_text()    │   │
  │                              └─────────────────────┘   │
  └──────────────────────────────────────────────────────────┘
      

Data Flow for a Single Query

What happens when you type "What's the weather in Paris?":

1
main.py receives input

Reads user text, passes it to run_agent()

2
agent.py builds context

Assembles system prompt with live tool list from MCP + long-term facts from memory.json

3
llm.py calls Ollama

Sends messages to localhost:11434/api/chat, waits for response

4
agent.py parses response

Regex finds Action: get_weather and Input: Paris

5
mcp_client.py calls server

POSTs {"tool": "get_weather", "input": "Paris"} to localhost:5001/tools/call

6
mcp_server.py runs the tool

Calls real get_weather("Paris") → fetches live data from Open-Meteo API

7
Result flows back

Weather data → mcp_client → agent appends "Observation" to conversation → calls LLM again

8
Final answer

LLM formats answer → Answer: It's 16°C in Paris... → printed to user, saved to memory

Tech Stack

ComponentTechnologyWhy
LanguagePython 3.10+Industry standard for AI/ML. Huge ecosystem.
LLMllama3.1:8b via OllamaRuns 100% locally. Free. No API key. Strong enough for agent tasks.
LLM ServerOllamaExposes local models via OpenAI-compatible HTTP API on port 11434.
MCP ServerFlaskMinimal HTTP server. Easy to understand. One file.
HTTP ClientrequestsThe standard Python library for HTTP calls.
Configpython-dotenvLoads .env files. Keeps secrets out of code. Industry standard.
Weather APIOpen-MeteoCompletely free. No signup. Returns real live data.
SearchWikipedia APIFree factual knowledge. No key needed.
MemoryRAM + JSON fileSimple, transparent, no database needed for learning.
LoggingPython loggingBuilt-in. Timestamps, levels, file + console output.
💡 Local-First Approach

Everything runs on your machine. No cloud costs, no rate limits, no data leaving your computer. You can see and control every single piece — the best way to truly learn the patterns.

The 9 Build Stages

Each stage adds exactly one new concept. Never skip — each builds directly on the last.

01

Environment Setup

Install Python, create a virtual environment, set up project folder structure, install dependencies. Learn why virtual environments exist.

02

First LLM Call

Send a message to llama3.1:8b via Ollama and print the response. Understand prompts, system prompts, and the message format all LLMs use.

03

Build a Simple Agent

Implement the ReAct loop. Give the LLM a role, tools, and a strict output format. See how the agent thinks before acting.

04

Add Tools

Create a calculator, note-saver, and live weather fetcher. Learn how agents choose which tool to use and how results feed back into the loop.

05

Add Memory

Implement short-term (RAM) and long-term (JSON) memory. Understand context windows, conversation history, and memory management.

06

Introduce MCP

Build a real MCP server with Flask. Learn the client/server pattern, HTTP endpoints, and tool discovery. Run two processes talking to each other.

07

Wire MCP to the Agent

Replace local tool calls with MCP client calls. One line of code change transforms the architecture. Tool discovery becomes dynamic.

08

Clean Up the Code

Add config.py for centralised settings, a proper logging system, and layered error handling. What separates working code from production code.

09

Final Project

Add Wikipedia search and text summarisation. Polish main.py with a command interface. Combine everything into one coherent, usable product.

The ReAct Pattern

ReAct stands for Reason + Act. It's the most widely-used pattern for AI agents. ChatGPT plugins, Claude tools, and LangChain agents all use it under the hood. The core idea: force the LLM to show its reasoning before every action — making behaviour transparent and debuggable.

The Output Format

The ReAct format we enforce via system prompt
# When the agent needs a tool:
Thought: I need current weather data for Tokyo. I'll use get_weather.
Action:  get_weather
Input:   Tokyo

# After seeing the result, if another tool is needed:
Thought: Got the weather. Now I'll save it to notes as requested.
Action:  save_note
Input:   Tokyo weather: 8°C, partly cloudy

# When the agent has a final answer:
Thought: I have all information needed to answer fully.
Answer:  The weather in Tokyo is 8°C and partly cloudy. Pack a coat.

Parsing the Response

Without a strict format, the LLM responds any way it wants — prose, JSON, bullet points. We need to parse it with code, so we enforce a machine-readable structure and extract the action with regex:

src/agent.py — parse_response()
import re

def parse_response(response: str) -> dict:
    # Look for "Action: tool_name"
    action_match = re.search(r"Action:\s*(\w+)", response)
    input_match  = re.search(r"Input:\s*(.+?)(?:\n|$)", response)

    if action_match:
        return {
            "type":  "action",
            "tool":  action_match.group(1).strip(),
            "input": input_match.group(1).strip() if input_match else ""
        }

    answer_match = re.search(r"Answer:\s*(.+)", response, re.DOTALL)
    if answer_match:
        return {"type": "answer", "content": answer_match.group(1).strip()}

    return {"type": "answer", "content": response.strip()}

The Full Agent Loop

src/agent.py — run_agent() simplified
def run_agent(user_query: str) -> str:
    context = user_query
    step = 0

    while step < MAX_STEPS:     # safety limit
        step += 1

        # Ask the LLM: what should I do next?
        response = chat(context, system_prompt=build_system_prompt())
        parsed   = parse_response(response)

        if parsed["type"] == "answer":
            return parsed["content"]   # ← done!

        elif parsed["type"] == "action":
            # Call tool via MCP, feed result back in
            result  = mcp.call_tool(parsed["tool"], parsed["input"])
            context = f"Observation: {result}\n\nContinue."
            # ↑ loop again with the new context

    return "Reached maximum steps."

Tools & Tool Calling

A tool is a Python function the agent can choose to call. Tools give capabilities beyond what the LLM knows — real-time data, computation, file access, external APIs. The agent doesn't call tools randomly; it reads their descriptions and uses language understanding to match the right tool to the task.

Why Tool Descriptions Matter

Vague Description
  • "does weather stuff"
  • "calculator tool"
  • "note thing"
Clear Description
  • "Get current weather for any city. Input: city name e.g. 'London'"
  • "Evaluate any Python math expression. e.g. '2**10' or 'math.sqrt(144)'"
  • "Save text to notes file for retrieval. Input: text to save."

The Weather Tool — Deep Dive

The best example of a realistic tool — calls an external API, parses JSON, handles errors, returns clean formatted text:

src/mcp_server.py — get_weather()
def get_weather(city: str) -> str:
    try:
        # Step 1: City name → coordinates (free geocoding API)
        geo = requests.get(
            "https://geocoding-api.open-meteo.com/v1/search",
            params={"name": city, "count": 1},
            timeout=10
        ).json()["results"][0]
        lat, lon = geo["latitude"], geo["longitude"]

        # Step 2: Live weather data from Open-Meteo (free, no key)
        data = requests.get(
            "https://api.open-meteo.com/v1/forecast",
            params={
                "latitude": lat, "longitude": lon,
                "current": ["temperature_2m", "weather_code"],
                "timezone": "auto"
            }, timeout=10
        ).json()

        current = data["current"]
        return f"Weather in {city}: {current['temperature_2m']}°C"

    except Exception as e:
        return f"Error: {str(e)}"
⚠️ Critical Tool Rule

Every tool function must accept exactly one string argument and return a string. The agent loop always passes one string input and expects a string back. If a tool needs no input, accept it anyway and ignore it: def read_notes(input: str = "") -> str:

Memory Systems

Without memory, every message starts fresh. Say "My name is James" — next message ask "What is my name?" and the agent has no idea. We solve this with two distinct memory types.

Short-Term Memory — Conversation History

Lives in RAM. Holds every message in the current session. Disappears when the program closes.

src/memory.py — ShortTermMemory
class ShortTermMemory:
    def __init__(self, max_messages: int = 20):
        self.messages = []
        self.max_messages = max_messages

    def add(self, role: str, content: str):
        self.messages.append({"role": role, "content": content})
        # Trim oldest messages when we hit the limit
        if len(self.messages) > self.max_messages:
            self.messages = self.messages[-self.max_messages:]

    def get_history(self) -> list:
        return [{"role": m["role"], "content": m["content"]} for m in self.messages]

Long-Term Memory — Persistent Facts

Saved to a JSON file on disk. Survives restarts. Injected into the system prompt every session so the agent "knows" it from the start.

src/memory.py — LongTermMemory
class LongTermMemory:
    def __init__(self, filepath: str = "memory.json"):
        self.facts = self._load()  # load from disk on startup

    def remember(self, key: str, value: str):
        self.facts[key] = {"value": value}
        self._save()  # write to disk immediately

    def get_all_facts(self) -> str:
        # Injected into system prompt every session
        return "\n".join(f"- {k}: {v['value']}" for k,v in self.facts.items())

What is a Context Window?

Every LLM has a context window — the maximum text it can see at once. For llama3.1:8b this is ~128,000 tokens. Everything — system prompt, conversation history, tool results — must fit within this limit. When conversations grow long, we trim oldest messages with messages[-max_messages:].

Orchestration

Orchestration is the coordination layer — the code that manages the overall flow. It's not a single feature; it's the glue that holds the whole system together. In this project, agent.py is the orchestrator.

It coordinates all five subsystems in the right order:

  • Receives user input from main.py
  • Assembles system prompt with dynamic tool list + user facts
  • Calls the LLM via llm.py and parses the response
  • Calls tools via mcp_client.py when needed
  • Updates conversation history in memory.py
  • Decides when the task is complete and returns the answer

The System Prompt — Heart of the Orchestrator

Rebuilt fresh for every query — dynamically including the live tool list and current facts:

agent.py — build_system_prompt()
def build_system_prompt() -> str:
    tool_descriptions = mcp.get_tool_descriptions()  # live from MCP server
    long_term_facts   = long_term.get_all_facts()       # from memory.json

    return f"""You are a helpful research assistant.

WHAT YOU KNOW ABOUT THE USER:
{long_term_facts}

TOOLS AVAILABLE:
{tool_descriptions}

Use this EXACT format when using a tool:
Thought: [reason]
Action:  [tool_name]
Input:   [input]

When done:
Thought: [reason]
Answer:  [complete answer]
"""
🔑 The Key Insight

Because tool descriptions are fetched live from the MCP server, if you add a new tool to mcp_server.py and restart it — the agent discovers it automatically. agent.py never needs to change. This is the power of the orchestration + MCP combination.

Key Terms Glossary

Every term you'll encounter in this project and the broader AI engineering field — explained with examples from the code you'll build.

Agent

LLM + loop + tools. Can take actions and iterate until a task is done. In this project: agent.py.

Orchestration

The coordination logic that manages agent flow — calling LLMs, tools, and memory in the right order.

Tool Calling

When an agent triggers a function to get real-world data or take action. e.g. calling get_weather("Tokyo").

Function Calling

Same as tool calling — the term used by OpenAI. "Tool calling" is more generic across providers.

MCP

Model Context Protocol — a standard for hosting tools on a server so any AI agent can discover and use them.

ReAct Pattern

Reason + Act. The agent shows its Thought before every Action. Makes agents transparent and debuggable.

System Prompt

Hidden instructions sent to the LLM before every conversation. Defines role, tools, and output format.

Context Window

Maximum text an LLM can process at once. llama3.1:8b = ~128k tokens. Everything must fit inside it.

Short-Term Memory

The conversation history in RAM for the current session. Lost when the program closes.

Long-Term Memory

Facts saved to disk (memory.json) that persist between sessions.

Local Model

An LLM running on your machine via Ollama. Free, private, no internet required.

Hosted Model

An LLM accessed via API (OpenAI, Anthropic). More powerful but costs money per request.

API Connector

Code connecting to an external service. In this project: Open-Meteo for weather, Wikipedia for search.

Structured Output

Forcing the LLM to respond in a specific format (Thought/Action/Answer) so code can parse it reliably.

Prompt Design

Crafting system prompts that reliably guide LLM behaviour. The most important skill in AI engineering.

Workflow vs Agent

A workflow follows a fixed script. An agent makes decisions. This project builds a true decision-making agent.

Step-by-Step Build Guide

Prerequisites

  • macOS or Linux (Windows users: use WSL2)
  • Python 3.10 or higher — check with python3 --version
  • Ollama installed with at least one model pulled
  • ~6 GB free disk space for the LLM model

Stage 1 — Environment Setup

Terminal
# Check Python version (need 3.10+)
python3 --version

# Create project folder and navigate into it
mkdir research-assistant && cd research-assistant

# Create and activate virtual environment
python3 -m venv .venv
source .venv/bin/activate   # you'll see (.venv) in your prompt

# Install all dependencies
pip install requests python-dotenv flask wikipedia-api

# Create folder structure
mkdir src logs
touch src/__init__.py .env .gitignore requirements.txt main.py
research-assistant/
├── .venv/ ← virtual environment (auto-created)
├── src/ ← all Python modules live here
│ └── __init__.py ← makes src a Python package
├── logs/ ← log files written here at runtime
├── .env ← config and secrets (never commit to git)
├── .gitignore ← tells git what to ignore
├── requirements.txt ← pinned dependencies
└── main.py ← entry point — this is what you run

Stage 2 — First LLM Call

src/llm.py
import requests

OLLAMA_URL = "http://localhost:11434/api/chat"
MODEL      = "llama3.1:8b"

def chat(user_message: str, system_prompt: str = None) -> str:
    messages = []
    if system_prompt:
        messages.append({"role": "system", "content": system_prompt})
    messages.append({"role": "user", "content": user_message})

    response = requests.post(
        OLLAMA_URL,
        json={"model": MODEL, "messages": messages, "stream": False},
        timeout=60
    )
    return response.json()["message"]["content"]

if __name__ == "__main__":
    print(chat("What is the capital of France?"))
    # → "The capital of France is Paris."

Running the Complete Final Project

The final project requires two terminal windows running simultaneously — this is the MCP client/server pattern in action:

Terminal 1 — MCP Tool Server
cd ~/research-assistant
source .venv/bin/activate
python3 src/mcp_server.py

# Expected output:
# 🔌 MCP Tool Server v2.0
# Tools available : 7
# Health : http://localhost:5001/health
# Tools  : http://localhost:5001/tools
Terminal 2 — Research Assistant Agent
cd ~/research-assistant
source .venv/bin/activate
python3 main.py

# Expected output:
# 🤖 Research Assistant Agent
# ✅ Connected to MCP server: Research Assistant MCP Server
# 🔧 Discovered 7 tools from MCP server
# Ready! Type /help for commands or just ask me anything.

Test Queries — Try These in Order

Covers every system in the project
# 1. Tests tool calling (calculator)
What is 2 to the power of 16?

# 2. Tests external API (Open-Meteo)
What's the weather in Tokyo?

# 3. Tests Wikipedia search tool
Search Wikipedia for the James Webb Space Telescope

# 4. Tests multi-tool chaining
What's the weather in Paris and save a note about it

# 5. Tests short-term memory
My name is [your name]
What is my name?

# 6. Tests long-term memory (survives restarts)
/remember city=London
quit → restart → What is the weather where I live?

# 7. Tests command interface
/tools
/memory
/facts

The .env Configuration File

.env
# LLM
OLLAMA_MODEL=llama3.1:8b
OLLAMA_URL=http://localhost:11434/api/chat

# MCP Server
MCP_SERVER_HOST=localhost
MCP_SERVER_PORT=5001

# Agent
AGENT_MAX_STEPS=5
AGENT_MEMORY_LIMIT=20

# Logging — DEBUG | INFO | WARNING | ERROR
LOG_LEVEL=INFO