0%
SustainSys AI Academy · Full Curriculum

AI Architect
Mastery Curriculum

From Python fundamentals to production AI systems. A structured, hands-on learning path covering LLMs, agents, RAG, enterprise architecture, multi-cloud, and GenAIOps — built for the AI Architect role.

Python 3.10+ LangChain LlamaIndex Ollama ChromaDB / FAISS Kubernetes 7 Modules 16–22 weeks
ModuleTopicDuration
1Foundations of Modern AI Systems2–3 weeks
2Multi-Agent AI Systems3–4 weeks
3Memory Systems in AI Agents2–3 weeks
4Retrieval Augmented Generation (RAG)3–4 weeks
5Enterprise AI Architecture2–3 weeks
6Cross-Cloud AI Architecture2–3 weeks
7GenAIOps2–3 weeks
Module 01

Foundations of Modern AI Systems

Est. 2–3 weeks · 7 sub-sections · 2 hands-on labs

This module builds your mental model of how modern AI systems work, from the mathematics of tokens and embeddings through to architecture decisions about where and how to run models. By the end, you will be able to call LLMs, generate embeddings, and reason about model selection trade-offs.

1.1 What Modern AI Architecture Looks Like

A modern AI system is not a single model sitting on a server. It is an orchestrated pipeline of components: user interfaces, API gateways, prompt management layers, one or more language models, retrieval systems, memory stores, tool connectors, and observability infrastructure. Think of it like a microservices architecture, but where the core compute is a probabilistic language model rather than deterministic business logic.

┌─────────────┐   ┌──────────────┐   ┌─────────────┐
│  User / UI  │───│  API Gateway  │───│ Orchestrator│
└─────────────┘   └──────────────┘   └─────┬───────┘
                                           │
                     ┌─────────────────┼───────────────┐
                     │                 │               │
               ┌─────┴─────┐  ┌─────┴─────┐  ┌────┴──────┐
               │  LLM(s)    │  │  Retrieval │  │  Tools /  │
               │ (Cloud/    │  │  (Vector   │  │  APIs     │
               │  Local)    │  │  DB + RAG) │  │          │
               └───────────┘  └───────────┘  └───────────┘

The Orchestrator is the brain. It receives a user request, decides which tools or models to use, retrieves relevant context, constructs a prompt, calls the LLM, and returns a response. Frameworks like LangChain, LlamaIndex, and Semantic Kernel provide this orchestration layer.

What Large Language Models (LLMs) Are

A Large Language Model is a neural network trained on vast amounts of text. Its fundamental operation is next-token prediction: given a sequence of tokens, it predicts the probability distribution over what comes next. This simple objective, scaled to billions of parameters and trillions of training tokens, produces emergent capabilities like reasoning, coding, and summarisation.

⚠ Key Insight: LLMs Are Probabilistic

An LLM does not "know" facts the way a database does. It has learned statistical patterns from training data. This is why LLMs can hallucinate — they generate text that is statistically plausible but factually wrong. RAG systems (Module 4) exist precisely to ground LLM outputs in verified data.

The major LLM families you will work with as an AI Architect include frontier models (GPT-4o, Claude, Gemini) that offer maximum capability, and small language models (Phi-3, Llama 3, Mistral) that can run locally or at the edge for lower cost and latency.

Tokens, Embeddings, and Prompts

Tokens

Tokens are the atomic units of text that LLMs process. A token is roughly 3–4 characters of English text, or about 0.75 words. The sentence "AI architecture is fascinating" is approximately 5 tokens. Tokenisation matters because LLM pricing, context windows, and performance are all measured in tokens.

Different models use different tokenisers. OpenAI uses tiktoken (BPE-based), while Llama models use SentencePiece. The key point: you cannot assume a one-to-one mapping between words and tokens.

Embeddings

An embedding is a dense vector representation of text in a high-dimensional space (typically 384 to 4096 dimensions). Texts with similar meanings are placed close together in this vector space. Embeddings are the foundation of semantic search, RAG, and many other AI patterns.

          meaning axis
              ^
              |     * "machine learning"
              |    * "deep learning"
              |
              |          * "neural networks"
              |
              |                        * "cooking recipes"
              |                       * "baking bread"
              +--------------------------------------> topic axis

  Similar concepts cluster together in vector space

We measure similarity between embeddings using cosine similarity, which ranges from -1 (opposite) to 1 (identical). In practice, related texts typically score 0.7–0.95.

Prompts

A prompt is the text input you send to an LLM. Effective prompting is a core skill for an AI Architect. The main patterns are:

  • Zero-shot: Ask the model directly with no examples.
  • Few-shot: Provide 2–5 examples of desired input/output format before your actual question.
  • Chain-of-thought: Ask the model to "think step by step" to improve reasoning accuracy.
  • System prompts: Instructions that set the model's persona, constraints, and output format.

Hands-On: Exploring Tokens and Embeddings

This script calls a local Ollama model to generate embeddings and measure similarity. Make sure you have Ollama installed and the nomic-embed-text model pulled.

Prerequisites: pip install ollama numpy

tokens_and_embeddings.py
# Explore how LLMs tokenise text and generate embeddings

import ollama
import numpy as np

# --- Part 1: Generate Embeddings ---
texts = [
    "Machine learning is a branch of artificial intelligence",
    "Deep learning uses neural networks with many layers",
    "I enjoy baking sourdough bread on weekends",
]

embeddings = []
for text in texts:
    response = ollama.embed(model="nomic-embed-text", input=text)
    emb = response["embeddings"][0]
    embeddings.append(emb)
    print(f"Text: {text[:50]}...")
    print(f"  Vector dimensions: {len(emb)}")
    print(f"  First 5 values: {emb[:5]}")

# --- Part 2: Measure Similarity ---
def cosine_similarity(a, b):
    a, b = np.array(a), np.array(b)
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

for i in range(len(texts)):
    for j in range(i + 1, len(texts)):
        sim = cosine_similarity(embeddings[i], embeddings[j])
        print(f"Similarity: {sim:.4f}")
💡 Expected Output

The two AI-related sentences will have high similarity (around 0.7–0.85), while comparing either to the baking sentence will produce low similarity (around 0.2–0.4). This is the foundation of semantic search — and the core mechanism behind RAG.

LLM Inference vs Training

Training is the process of adjusting a model's billions of parameters by feeding it enormous datasets. Pre-training GPT-4-class models costs tens of millions of dollars and months of compute on thousands of GPUs.

Inference is running a trained model to generate outputs. This is what happens when you call an API or run Ollama locally. Inference is orders of magnitude cheaper than training, but it's still the primary cost driver for production AI systems.

AspectTrainingInference
PurposeLearn patterns from dataGenerate predictions/text
CostMillions of dollars for frontier modelsFractions of a penny per request
ComputeThousands of GPUs for weeksSingle GPU or CPU per request
FrequencyOnce (or periodic fine-tuning)Every user interaction
Your roleChoose when/whether to fine-tuneOptimise latency, cost, throughput

As an AI Architect, you will spend 95% of your time on inference-side architecture: how to route requests, cache responses, manage context windows, and optimise cost.

Frontier Models vs Small Language Models

This is one of the most important architectural decisions you will make. Frontier models (Claude Opus, GPT-4o, Gemini Ultra) have the highest capability but highest cost and latency. Small Language Models (Phi-3, Llama 3 8B, Mistral 7B) are 10–100x cheaper, can run locally, but have lower reasoning capability.

The key insight: most production tasks don't need frontier models. Classification, extraction, summarisation, and simple Q&A can often be handled by SLMs. Reserve frontier models for complex reasoning, multi-step planning, and creative generation.

FactorFrontier ModelsSmall Language Models
Parameters100B – 1T+1B – 13B
ReasoningExcellentGood for focused tasks
Cost / 1M tokens$2 – $60$0.05 – $0.50 (or free locally)
Latency1 – 5 seconds0.1 – 1 second locally
PrivacyData sent to providerRuns entirely on-premise
Use casesComplex analysis, coding, planningClassification, extraction, chat
ℹ Architecture Pattern: Model Routing

A production system often uses BOTH frontier and small models. A lightweight classifier analyses the incoming request and routes simple queries to an SLM while sending complex queries to a frontier model. This can cut costs by 60–80% while maintaining quality.

Edge Inference vs Cloud Inference

Cloud inference means your model runs on a provider's servers (OpenAI, Anthropic, Azure, AWS). Benefits: no GPU management, easy scaling, access to frontier models. Drawbacks: latency, data privacy concerns, ongoing costs, vendor dependency.

Edge inference means running models directly on user devices or local servers. Ollama on your Mac M4 is edge inference. Benefits: zero latency to network, complete data privacy, no per-request costs. Drawbacks: limited model size, device resource constraints.

                 Need frontier-level reasoning?
                      /             \
                   YES               NO
                    |                 |
              Use CLOUD           Data sensitive?
              (GPT-4o,            /          \
               Claude,         YES            NO
               Gemini)          |              |
                          Use EDGE         Use CLOUD SLM
                          (Ollama,         (cheaper endpoint)
                           llama.cpp)

Hybrid AI Architecture

In enterprise settings, you almost always build a hybrid architecture that combines cloud and edge, frontier and small models. The orchestration layer decides at runtime which model to use based on the task requirements, data sensitivity, cost constraints, and latency needs.

┌─────────────────────────────────────────────┐
│            ORCHESTRATION LAYER              │
│  (LangChain / LlamaIndex / Custom)          │
└─────────┬───────────┬───────────┬───────────┘
          │           │           │
    ┌─────┴───┐  ┌────┴─────┐  ┌───┴───────┐
    │ CLOUD    │  │ EDGE     │  │ RETRIEVAL │
    │ (Claude  │  │ (Ollama  │  │ (ChromaDB │
    │  GPT-4o) │  │  llama3) │  │  FAISS)   │
    └──────────┘  └──────────┘  └───────────┘
    Complex         Simple        Knowledge
    reasoning       tasks         grounding

Hands-On: Calling LLMs Locally and Comparing Models

Prerequisites: ollama pull llama3.1:8b && ollama pull phi3:mini

compare_models.py
import ollama
import time

PROMPT = """You are a senior AI architect. Explain in 3 sentences
why RAG is preferred over fine-tuning for enterprise knowledge bases."""

models = ["llama3.1:8b", "phi3:mini"]

for model in models:
    print(f"\nModel: {model}")
    start = time.time()
    response = ollama.chat(
        model=model,
        messages=[{"role": "user", "content": PROMPT}]
    )
    elapsed = time.time() - start
    print(f"Response ({elapsed:.1f}s):")
    print(response["message"]["content"])
    eval_count = response.get("eval_count", 0)
    if elapsed > 0 and eval_count > 0:
        print(f"Speed: {eval_count/elapsed:.1f} tokens/sec")
💡 What to Observe

llama3.1:8b will likely give more nuanced answers. phi3:mini may be faster. Both are running entirely on your Mac M4 with zero API costs. This is the foundation of edge inference.

🧪 Module 1 Exercises
  1. Modify the embeddings script to add 5 more sentences and find which pairs are most/least similar
  2. Write a script that counts tokens for different inputs using tiktoken (pip install tiktoken)
  3. Create a cost calculator: given token counts and model pricing, estimate monthly API costs
  4. Build a model router: classify user queries as 'simple' or 'complex' and route to different models
🚀 Mini-Project: Model Comparison Dashboard

Build a Python script that sends 10 different prompts (classification, summarisation, code generation, reasoning, etc.) to both llama3.1:8b and phi3:mini. Record response time, token count, and manually rate quality 1–5. Output a comparison table. This teaches you the evaluation mindset every AI Architect needs.

Module 02

Multi-Agent AI Systems

Est. 3–4 weeks · Agents, ReAct, Orchestration · 3 hands-on labs

Agents are LLMs that can take actions. Instead of just generating text, an agent can call tools, search the web, write code, query databases, and make decisions about what to do next. Multi-agent systems coordinate multiple specialised agents to solve complex problems.

2.1 What AI Agents Are

An AI agent has three core capabilities: perception (understanding the current state), reasoning (deciding what to do), and action (executing tools or generating outputs). The LLM provides the reasoning; the framework provides the perception and action layers.

  ┌─────────────────────────────────────┐
  │              AI AGENT                │
  │                                      │
  │  ┌───────────┐  Input  ┌─────────┐  │
  │  │ Perceive  │─────────│ Reason  │  │
  │  │ (context, │         │ (LLM    │  │
  │  │  memory)  │         │  core)  │  │
  │  └───────────┘         └────┬────┘  │
  │                             │       │
  │                        ┌────┴────┐  │
  │                        │  Act    │  │
  │                        │ (tools, │  │
  │                        │  APIs)  │  │
  │                        └─────────┘  │
  └─────────────────────────────────────┘

2.2 The ReAct Pattern

ReAct (Reasoning + Acting) is the most important agent pattern. The agent iterates through a loop:

1
Thought

Reason about what to do — "I need to calculate 15% of 2340"

2
Action

Call a tool — calculate("2340 * 0.15")

3
Observation

Process the result — "Result: 351.0"

4
Repeat or Answer

If done, return final answer. Otherwise, loop back to Thought.

react_agent.py
# A simple ReAct agent using LangChain with Ollama
# pip install langchain langchain-community langchain-ollama

from langchain_ollama import ChatOllama
from langchain.agents import AgentExecutor, create_react_agent
from langchain.tools import tool
from langchain import hub

@tool
def calculate(expression: str) -> str:
    """Evaluate a mathematical expression."""
    try:
        return str(eval(expression))
    except Exception as e:
        return f"Error: {e}"

@tool
def get_word_count(text: str) -> str:
    """Count the number of words in a text string."""
    return str(len(text.split()))

llm = ChatOllama(model="llama3.1:8b", temperature=0)
prompt = hub.pull("hwchase17/react")
tools = [calculate, get_word_count]
agent = create_react_agent(llm, tools, prompt)
executor = AgentExecutor(agent=agent, tools=tools, verbose=True)

result = executor.invoke({
    "input": "What is 15% of 2340?"
})
print(f"Final Answer: {result['output']}")

2.3 Multi-Agent Architecture Patterns

  • Hierarchical: A manager agent delegates tasks to worker agents. Best for well-defined workflows.
  • Peer-to-Peer: Agents communicate directly. Best for collaborative problem-solving.
  • Sequential (Pipeline): Agent A's output feeds into Agent B, then Agent C. Best for staged processing.
  • Task Decomposition: A planner breaks a complex task into subtasks, each handled by a specialist agent.
                ┌─────────────┐
                │  MANAGER    │
                │  AGENT      │
                └───┬───┬─────┘
                    │   │
           ┌───────┘   └───────┐
     ┌────┴─────┐        ┌────┴─────┐
     │ Research  │        │ Writing   │
     │ Agent     │        │ Agent     │
     └──────────┘        └──────────┘
     Tools: search        Tools: format,
     retrieve, analyse    write, edit
🧪 Module 2 Exercises
  1. Add a third agent (Editor) to a pipeline that checks the summary for factual consistency against the research
  2. Convert a sequential pipeline to hierarchical: create a Manager agent that decides which specialist to call
  3. Build a tool-using agent that can search files on your local machine (use Python's os module as a tool)
  4. Implement error handling: what happens when an agent produces invalid output? Build a retry mechanism
Module 03

Memory Systems in AI Agents

Est. 2–3 weeks · Short-term, Long-term, Episodic, Audit · 2 hands-on labs

LLMs are stateless by default — they have no memory of previous conversations. Every interaction starts from scratch unless you explicitly provide context. Memory systems solve this by storing and retrieving relevant information from past interactions, documents, and structured knowledge.

3.1 Types of Memory

Memory TypeWhat It StoresLifespanImplementation
Short-term (Buffer)Current conversation turnsSingle sessionIn-memory list of messages
Long-term (Semantic)Facts, preferences, knowledgePersistentVector database (Chroma, FAISS)
EpisodicSpecific past interactionsPersistentTimestamped vector entries
Audit / ComplianceAll interactions for reviewPermanentAppend-only log / database

3.2 Long-Term Memory with Vector Databases

  User says something    ┌────────────┐    Query top-k
  worth remembering  ───>│  Embedding  │───>┌───────────┐
                         │  Model      │    │  Vector   │
  Agent needs context    └────────────┘    │  Database  │
  about user         ───> embed query ─────>│ (Chroma/  │
                                            │  FAISS)   │
                                            └───────────┘

Hands-On: Agent with Persistent Memory

Prerequisites: pip install chromadb

persistent_memory_agent.py
import chromadb, ollama
from datetime import datetime

client = chromadb.PersistentClient(path="./agent_memory")
collection = client.get_or_create_collection(
    name="user_memories",
    metadata={"hnsw:space": "cosine"}
)

def store_memory(text):
    emb = ollama.embed(model="nomic-embed-text", input=text)["embeddings"][0]
    mem_id = f"mem_{datetime.now().strftime('%Y%m%d_%H%M%S_%f')}"
    collection.add(ids=[mem_id], embeddings=[emb], documents=[text])

def recall_memories(query, n=3):
    emb = ollama.embed(model="nomic-embed-text", input=query)["embeddings"][0]
    results = collection.query(query_embeddings=[emb], n_results=n)
    return results["documents"][0] if results["documents"] else []

def chat_with_memory(user_input):
    memories = recall_memories(user_input)
    ctx = "\n".join(f"- {m}" for m in memories) or "No memories yet."
    prompt = f"Memories:\n{ctx}\n\nUser: {user_input}\nRespond helpfully:"
    resp = ollama.chat(model="llama3.1:8b", messages=[{"role":"user", "content":prompt}])
    store_memory(f"User said: {user_input}")
    return resp["message"]["content"]

print(chat_with_memory("I am building a RAG system for a bank in London"))
print(chat_with_memory("What project am I working on?"))
💡 Key Insight

The second query retrieves the memory from the first exchange. Even if you restart the script, ChromaDB persists to disk, so the memories survive. This is how production agents maintain context across sessions.

3.4 Knowledge Graphs for Structured Memory

Vector databases store unstructured semantic memory. Knowledge graphs store structured relationships: "Satya WORKS_AT Sustainsys", "Sustainsys SERVES financial_services". Combining both gives agents the richest possible memory. In production, tools like Neo4j or Amazon Neptune provide the graph database.

🧪 Module 3 Exercises
  1. Extend the persistent memory agent to categorise memories (personal, project, preference) using metadata
  2. Implement a memory decay function: older memories get lower retrieval scores
  3. Build an audit memory layer: log every interaction to a separate append-only collection
  4. Create a memory management UI: list all memories, delete specific ones, search by date range
Module 04

Retrieval Augmented Generation (RAG)

Est. 3–4 weeks · Chunking, Embeddings, Retrieval, RAGAS · 3 hands-on labs

RAG is the most important pattern in enterprise AI. It solves the core problem of LLMs: they only know what was in their training data. RAG lets you ground LLM responses in your own documents — company policies, financial reports, medical records, product manuals — without retraining the model.

4.1 Why RAG Exists

  • No retraining needed: Update documents instantly without expensive fine-tuning.
  • Verifiable sources: Every answer can cite the specific document it came from.
  • Access control: You can restrict which documents different users can access.
  • Cost-effective: Works with any LLM, including free local models.

4.2 The RAG Pipeline

  INGESTION (offline)                 QUERY (realtime)
  ─────────────────                 ──────────────────

  ┌────────┐                        ┌─────────┐
  │ Docs   │  Load                  │  User   │
  │ (PDF,  │────┐                   │  Query  │
  │  CSV)  │    v                   └────┬────┘
  └────────┘  ┌────────┐                v
              │ Chunk  │          ┌─────────┐
              │ Text   │          │ Embed   │
              └───┬────┘          │ Query   │
                  v               └────┬────┘
              ┌────────┐              v
              │ Embed  │   ┌─────────────────┐
              │ Chunks │   │  VECTOR DATABASE │
              └───┬────┘   │  (similarity     │
                  └──────> │   search)        │
                           └────────┬────────┘
                                    v
                             ┌──────────┐
                             │ Reranker │  (optional)
                             └───┬──────┘
                                 v
                           ┌────────┐   ┌────────┐
                           │  LLM   │──>│ Answer │
                           └────────┘   │ + cite │
                                        └────────┘

4.3 Document Ingestion and Chunking

Chunking is the most impactful design decision in RAG. The main strategies are:

  • Fixed-size chunking: Split every N characters/tokens with overlap. Simple but can break mid-sentence.
  • Recursive character splitting: Split by paragraph, then sentence, then character. Preserves structure.
  • Semantic chunking: Use embeddings to find natural topic boundaries. Best quality but slower.
  • Hierarchical chunking: Create parent (large) and child (small) chunks. Search children, retrieve parents for context.

Hands-On: Complete RAG Pipeline

Prerequisites: pip install langchain langchain-community langchain-ollama chromadb

rag_pipeline.py
from langchain_community.document_loaders import TextLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Chroma
from langchain_ollama import OllamaEmbeddings, ChatOllama
from langchain.chains import RetrievalQA
from langchain.prompts import PromptTemplate

# Step 1: Load & Chunk Documents
loader = TextLoader("sample_doc.txt")
splitter = RecursiveCharacterTextSplitter(
    chunk_size=200, chunk_overlap=50,
    separators=["\n\n", "\n", ". ", " "]
)
chunks = splitter.split_documents(loader.load())

# Step 2: Create Vector Store
embeddings = OllamaEmbeddings(model="nomic-embed-text")
vectorstore = Chroma.from_documents(chunks, embeddings,
    persist_directory="./rag_chroma_db")

# Step 3: RAG Chain
retriever = vectorstore.as_retriever(search_kwargs={"k": 3})
llm = ChatOllama(model="llama3.1:8b", temperature=0)

prompt = PromptTemplate(
    template="""Use ONLY the context below to answer.
Context: {context}
Question: {question}
Answer:""",
    input_variables=["context", "question"]
)

qa = RetrievalQA.from_chain_type(
    llm=llm, retriever=retriever,
    chain_type_kwargs={"prompt": prompt},
    return_source_documents=True
)

result = qa.invoke({"query": "What retrieval approach does Sustainsys use?"})
print(f"Answer: {result['result']}")

4.4 RAG Evaluation with RAGAS

RAGAS measures four key metrics:

  • Faithfulness: Does the answer only contain information from the retrieved context?
  • Answer Relevancy: Is the answer actually relevant to the question asked?
  • Context Precision: Are the retrieved documents actually relevant?
  • Context Recall: Did retrieval find all the relevant information?

4.5 Cache Augmented Generation (CAG)

CAG is an optimisation where frequently-asked questions and their RAG-generated answers are cached. This reduces latency from seconds to milliseconds and cuts compute costs dramatically for repetitive queries. Implement using Redis or a simple dictionary with TTL expiry.

🧪 Module 4 Exercises
  1. Extend the RAG pipeline to ingest PDF files using PyPDFLoader
  2. Implement hybrid retrieval: combine Chroma vector search with BM25 keyword search
  3. Add a reranking step using a cross-encoder model to reorder retrieved chunks
  4. Build a RAG evaluation harness: create 10 question-answer pairs and measure retrieval precision
  5. Implement a simple CAG layer using a Python dictionary with timestamp-based expiry
🚀 Mini-Project: Multi-Document RAG for Financial Services

Build a RAG system that ingests 5+ documents (mix of PDF, CSV, and text), uses hybrid retrieval (vector + BM25), includes metadata filtering (document type, date), and returns answers with source citations. Add a simple Flask API endpoint so it can be called from other systems.

Module 05

Enterprise AI Architecture

Est. 2–3 weeks · Auth, Integration, Kubernetes · Architecture patterns

Enterprise AI is not just about making a model work — it's about making it work securely, at scale, with auditability, within existing IT ecosystems.

5.1 Authentication and Identity

Microsoft Entra ID (formerly Azure AD) is the dominant identity provider. Your AI system must authenticate users via OAuth 2.0 / OpenID Connect, map user identity to document access permissions, and log all interactions for audit.

User ──> Entra ID ──> Token ──> API Gateway ──> AI Orchestrator
                                   |                    |
                                   v                    v
                               Validate             Filter docs
                               token +              by user's
                               extract              permissions
                               roles                (RBAC)

5.2 Enterprise Integration Patterns

  • Microsoft Teams / Slack: Chatbot interfaces — users ask questions in chat, AI retrieves from RAG and responds.
  • Salesforce / CRM: AI analyses customer data, generates summaries, predicts churn.
  • ERP Systems (SAP, Oracle): AI assists with procurement, inventory forecasting.
  • Email (Exchange / Gmail): AI drafts responses, categorises incoming mail, extracts action items.

5.3 Containerisation and Kubernetes

  ┌──────────── KUBERNETES CLUSTER ────────────┐
  │                                              │
  │  ┌──────────┐  ┌──────────┐  ┌────────┐    │
  │  │ API      │  │ RAG      │  │ Vector │    │
  │  │ Gateway  │  │ Service  │  │ DB     │    │
  │  │ (3 pods) │  │ (5 pods) │  │(3 pods)│    │
  │  └──────────┘  └──────────┘  └────────┘    │
  │                                              │
  │  ┌──────────┐  ┌──────────┐  ┌────────┐    │
  │  │ LLM      │  │ Agent    │  │ Cache  │    │
  │  │ Proxy    │  │ Workers  │  │ (Redis)│    │
  │  │ (2 pods) │  │ (auto)   │  │(2 pods)│    │
  │  └──────────┘  └──────────┘  └────────┘    │
  └──────────────────────────────────────────────┘
🧪 Module 5 Exercises
  1. Create a Flask API wrapper for your RAG pipeline with /query and /health endpoints
  2. Implement role-based document filtering: admin users see all docs, regular users see only their department's
  3. Write a Kubernetes deployment YAML for the RAG service with 3 replicas and resource limits
  4. Design a complete enterprise AI architecture for a bank including all security and integration points
Module 06

Cross-Cloud AI Architecture

Est. 2–3 weeks · Azure AI, Vertex AI, AWS Bedrock · Multi-cloud patterns

6.1 Cloud AI Platform Comparison

CapabilityAzure AIGoogle Vertex AIAWS Bedrock
LLM AccessOpenAI (exclusive), Llama, MistralGemini, Claude, LlamaClaude, Llama, Titan, Mistral
Vector DBAzure AI SearchVertex AI Vector SearchOpenSearch, pgvector on RDS
RAG SupportAzure AI Studio with groundingVertex AI RAG EngineKnowledge Bases for Bedrock
Agent FrameworkSemantic Kernel / AutoGenVertex AI Agent BuilderBedrock Agents
IdentityEntra ID (native)Cloud IAM + Workforce IdentityIAM + Cognito
StrengthsEnterprise Microsoft ecosystemML/data science toolingBroadest model selection

6.2 Avoiding Vendor Lock-In

ℹ Architecture Pattern: Cloud-Agnostic AI Layer

Build your AI logic in LangChain/LlamaIndex. Use environment variables to switch between providers. Store vector data in pgvector (PostgreSQL extension) which runs identically on all clouds. This lets you move between clouds or run multi-cloud without rewriting application code.

6.3 Multi-Cloud Decision Framework

  • Existing ecosystem: If the client runs Microsoft 365, Azure is the path of least resistance.
  • Model requirements: If you need GPT-4o, you need Azure. If you need Claude with native tooling, AWS Bedrock.
  • Data residency: Check which cloud has regions in the required jurisdictions.
  • Cost: Compare compute costs AND data transfer fees, which can be substantial in multi-cloud setups.
🧪 Module 6 Exercises
  1. Create a LangChain script that switches between Ollama, OpenAI, and Anthropic using only environment variables
  2. Design a multi-cloud architecture for healthcare: patient data on-premise, LLM on Azure, analytics on GCP
  3. Compare pricing: monthly cost of 100,000 RAG queries/day on Azure AI vs AWS Bedrock vs local Ollama
  4. Build a provider abstraction class wrapping embedding generation for Ollama, OpenAI, and Google
Module 07

GenAIOps

Est. 2–3 weeks · CI/CD for AI, IaC, Monitoring, Deployment · 2 hands-on labs

GenAIOps is MLOps evolved for the generative AI era. It covers the full lifecycle of AI systems: development, testing, deployment, monitoring, and continuous improvement.

7.1 The GenAI Lifecycle

  ┌──────────┐     ┌──────────┐     ┌──────────┐
  │  DEVELOP  │────>│   TEST    │────>│  DEPLOY   │
  │ Prompts   │     │ Evals     │     │ CI/CD     │
  │ RAG pipe  │     │ RAGAS     │     │ IaC       │
  │ Agents    │     │ A/B test  │     │ Canary    │
  └──────────┘     └──────────┘     └─────┬─────┘
       ^                                   │
       │          ┌──────────┐             │
       └──────────│  MONITOR  │<────────────┘
                  │ Latency   │
                  │ Quality   │
                  │ Cost      │
                  │ Drift     │
                  └──────────┘

7.2 CI/CD for AI Systems

  • Code tests: Standard unit/integration tests for application logic.
  • Prompt regression tests: Run a fixed set of prompts and compare outputs to baselines.
  • RAG evaluation: Run RAGAS metrics. Fail the build if faithfulness drops below thresholds.
  • Cost estimation: Calculate projected costs. Alert if a prompt change increases token usage.

7.3 Infrastructure as Code

terraform/main.tf — simplified example
resource "azurerm_cognitive_account" "openai" {
  name                = "sustainsys-openai"
  location            = "uksouth"
  resource_group_name = azurerm_resource_group.ai.name
  kind                = "OpenAI"
  sku_name            = "S0"
}

resource "azurerm_postgresql_flexible_server" "pgvector" {
  name                = "sustainsys-pgvector"
  location            = "uksouth"
  resource_group_name = azurerm_resource_group.ai.name
  sku_name            = "GP_Standard_D2s_v3"
  storage_mb          = 65536
}

7.4 Monitoring and Observability

MetricWhat It MeasuresTargetAlert Threshold
Latency (P50/P95)Response time< 2s P50> 5s P95
Token usage/queryCost efficiency< 2000 tokens avg> 5000 tokens
Retrieval precisionRAG quality> 0.85< 0.70
Hallucination rateAnswer accuracy< 5%> 15%
Error rateSystem reliability< 0.1%> 1%
User satisfactionThumbs up/down ratio> 80% positive< 60%

Tools like LangSmith, Weights & Biases, and Azure AI Studio provide dashboards and tracing specifically designed for GenAI systems.

7.5 Model Deployment Strategies

  • Canary deployment: Route 5% of traffic to the new version. Monitor for quality regressions.
  • Blue/Green deployment: Run old and new versions simultaneously. Instant rollback if needed.
  • A/B testing: Run two prompt variants with different user groups. Measure which performs better.
🧪 Module 7 Exercises
  1. Extend the logger to generate daily summary reports: avg latency, total tokens, error count
  2. Build a prompt regression test suite: 10 questions with expected answers, run nightly, flag quality drops
  3. Create a Dockerfile + docker-compose.yml that runs your RAG system + Redis cache + monitoring logger
  4. Design a complete GenAIOps pipeline diagram: from git push to production deployment with all quality gates
Capstone Project

AI-Powered Financial Advisory Platform

Integrates all 7 modules into a production-grade system

Build an AI system for a financial services client that:

  • Ingests regulatory documents, company policies, and market reports (Module 4: RAG)
  • Uses a multi-agent architecture: Router Agent, Compliance Agent, Research Agent, Response Agent (Module 2: Agents)
  • Maintains conversation history and user preferences across sessions (Module 3: Memory)
  • Runs locally via Ollama for sensitive queries, routes complex analysis to cloud LLMs (Module 1: Hybrid)
  • Exposes a REST API with JWT authentication (Module 5: Enterprise)
  • Uses pgvector for cloud-agnostic vector storage (Module 6: Cross-Cloud)
  • Includes monitoring dashboard, evaluation suite, and Dockerised deployment (Module 7: GenAIOps)

Deliverables

  • Python application with all components integrated
  • Architecture diagram showing all components and data flows
  • Docker Compose file for local deployment
  • Evaluation report: RAGAS scores, latency benchmarks, cost projections
  • README documenting design decisions and trade-offs
💡 Learning Path Recommendation

Work through the modules sequentially. Each module's hands-on exercises build on the previous modules. Budget 2–3 weeks per module. Run all code examples locally on your Mac M4 using Ollama. When you reach the capstone, you'll have all the components ready to integrate. Total estimated time: 16–22 weeks.

This curriculum was designed to take you from Python fundamentals to AI Architect capability. Every concept connects directly to the skills demanded in the job description. The hands-on exercises use the exact tools and frameworks — LangChain, LlamaIndex, Ollama, ChromaDB, FAISS — that you will use in production. Build each project, extend each exercise, and by the end you will be designing AI systems with confidence.