Introduction to AI Agents

AI Agents represent a fundamental shift in how we build intelligent systems. Rather than creating static models that map inputs to outputs, agents are autonomous systems that can perceive their environment, reason about it, take actions, and learn from the results. They combine language models, planning, memory, and tool use to solve complex, multi-step problems that require interaction with the real world.

The emergence of capable language models like GPT-4, Claude, and others has made building agents practical and powerful. An AI agent can be thought of as a loop: Observe โ†’ Think โ†’ Act โ†’ Observe. At each step, the agent uses reasoning frameworks (like ReAct), calls tools (APIs, databases, search engines), maintains memory of what it has learned, and coordinates with other agents to accomplish complex tasks.

What makes agents different from simple LLM applications is their agency โ€” the ability to make decisions autonomously, call tools as needed, recover from errors, and interact with dynamic environments. A chatbot responds to one message at a time. An agent can decompose a task into subtasks, fetch information, integrate results, and iterate toward a goal.

What You'll Learn

Agent Fundamentals

Understand the agent loop, perception, reasoning, planning, and action. Learn why agents are fundamentally different from prompt engineering.

Reasoning Frameworks

Explore ReAct (Reasoning + Acting), Chain-of-Thought, Tree-of-Thought, and Plan-and-Execute approaches for structured agent behavior.

Tool Use & Grounding

Learn how agents use tools (APIs, search, databases) to ground language in reality. Implement custom tools and handle failures gracefully.

Memory & Learning

Master conversation memory, episodic memory, semantic memory, and vector-based retrieval to help agents learn and leverage past interactions.

Multi-Agent Systems

Design systems where multiple agents collaborate (or compete) to solve problems. Understand emergent behaviors and coordination mechanisms.

Practical Implementation

Build real agents using LangChain, AutoGPT, CrewAI, and other frameworks. Deploy, evaluate, and iterate on agent behavior.

Prerequisites

Understanding of LLMs (transformers, prompt engineering), Python programming, basic NLP concepts, and APIs. Familiarity with LangChain or similar frameworks is helpful but not required.

Why AI Agents Matter

Agents represent the next frontier of AI capability. While large language models are powerful, they are fundamentally reactionary โ€” they generate responses based on input. Agents are proactive and goal-directed. They persist over time, maintain state, take actions, and adapt based on results.

Agents vs. Traditional LLM Applications

Dimension Traditional LLM AI Agent
Control Flow Single pass: Prompt โ†’ Response Loop: Observe โ†’ Think โ†’ Act โ†’ Observe
Tool Use No tools; only generates text Calls APIs, databases, search, calculators
Autonomy Responds to user input Pursues goals with minimal supervision
Memory Context window only Persistent memory, retrieval, learning
Error Handling One chance; errors end the interaction Detects failures, retries, adapts strategy
Scalability Handles one task per prompt Breaks complex tasks into subtasks
Collaboration Single model only Multiple agents coordinate & specialize

Real-World Agent Applications

Research Agents

Autonomously search scientific literature, summarize findings, generate hypotheses, and propose experiments. Multi-day research cycles with human oversight.

Code Generation Agents

Generate, test, debug, and refactor code. Agents that write code, run tests, analyze failures, and iterate until tests pass.

Customer Service Agents

Resolve customer issues, access CRM systems, process returns, escalate to humans when needed. Available 24/7 with context awareness.

Business Intelligence Agents

Query databases, generate reports, extract insights, create visualizations, explain trends to executives in plain language.

Robotics & Embodied AI

Control robots, perceive environments, plan actions, adjust based on feedback. The agent loop becomes the control loop.

Scientific Discovery

Run simulations, analyze results, propose new experiments, iterate toward breakthrough discoveries with minimal human intervention.

Key Insight: Agents are not just a technical advancement โ€” they fundamentally change how we think about AI systems. Instead of asking 'What is this text about?' we can ask 'What should this system do to achieve this goal?' This shift unlocks orders of magnitude more capability.

The Agent Capability Spectrum

Different applications require different levels of agent sophistication:

Simple Tool Use (90% of current use cases)
Single tool call with reasoning
Multi-Step Agents (70% adoption in enterprise)
Multiple tools, basic planning
Autonomous Agents (emerging frontier)
Full autonomy, learning, adaptation
Multi-Agent Systems (research & advanced)
Coordination, specialization, emergence

Core Concepts and Terminology

Before building agents, we need to establish a shared vocabulary. These concepts form the foundation of all agent design.

1. The Agent Loop: Observe โ†’ Think โ†’ Act โ†’ Observe

Every agent operates on a fundamental cycle:

Observe
Perceive environment
Read sensors, fetch data, get context
Think
Reason about situation
Use reasoning framework, plan actions
Act
Execute actions
Call tools, modify environment
Observe
Perceive results
Check success, handle errors

This loop continues until the agent reaches a terminal state (goal achieved, max steps reached, or unrecoverable error).

State Management

Between iterations, the agent maintains state: the conversation history, memory of past actions, progress toward goals, and learned information. This state is crucial for coherent agent behavior across multiple steps.

2. Perception and Grounding

Perception is how agents understand the world. Unlike human perception (visual, auditory), AI agents perceive through:

  • Language input: User queries, descriptions of the environment
  • Structured data: Database queries, API responses, sensor readings
  • Multimodal: Text + images, text + audio, text + video
  • Implicit signals: Rewards, success/failure feedback, error messages

Grounding means connecting abstract concepts to concrete reality. An agent that only generates text is not grounded โ€” it can hallucinate facts. An agent that calls APIs is grounded โ€” it can verify facts against reality.

Why Grounding Matters

A language model asked 'What's the current weather?' will make up an answer. An agent with a weather API tool will fetch real data. This distinction between knowledge and access is fundamental to building trustworthy AI systems.

3. Goals and Task Decomposition

Agents are goal-directed. Given a goal, they must:

  • Break it into subgoals or subtasks
  • Reason about dependencies between tasks
  • Execute subtasks in the right order
  • Integrate results to achieve the overall goal

This is fundamentally different from single-prompt response. A user might say "Compare the 2026 revenue of Competitor A and Competitor B." The agent must decompose this into subtasks:

  1. Find Competitor A's 2026 revenue (from their financial reports or APIs)
  2. Find Competitor B's 2026 revenue (same source type)
  3. Compare the two numbers and explain the difference
  4. Draw conclusions about market position

4. Tool Use and Grounding Functions

Tools are how agents interact with the world. Common tools include:

Knowledge Tools

Search engines, Wikipedia, knowledge bases, vector databases for semantic search.

Computational Tools

Calculators, code execution, scientific computing, symbolic math.

Data Tools

Database queries, API calls, data transformation, analytics.

Action Tools

Sending emails, posting to social media, triggering workflows, file I/O.

5. Memory: The Agent's Long-Term Storage

While LLM context windows are limited (4K to 200K tokens), agents need to remember:

  • Episodic memory: What happened in past interactions (conversation history)
  • Semantic memory: Facts and knowledge learned (vector embeddings)
  • Procedural memory: How to accomplish tasks (learned patterns, heuristics)
  • State memory: Current progress toward goals

Agents use techniques like retrieval-augmented generation (RAG) to efficiently access relevant memories without overloading the context window.

6. Reasoning and Decision-Making

How agents decide what to do next is critical. Different reasoning frameworks include:

  • Chain-of-Thought (CoT): Generate intermediate reasoning steps
  • ReAct: Interleave reasoning (Thought) and acting (Action)
  • Plan-and-Execute: Create a full plan first, then execute it
  • Tree-of-Thought: Explore multiple reasoning paths and backtrack
  • Hierarchical Planning: Abstract high-level goals into lower-level tactics

7. Error Handling and Recovery

Agents must be robust. When something goes wrong:

  • Detection: Recognize that an action failed (error message, validation failure)
  • Analysis: Understand why it failed
  • Recovery: Try a different approach or ask for help
  • Learning: Update mental model to avoid the same error next time

A naive agent that stops after the first failure is useless. A sophisticated agent adapts, retries with different parameters, or escalates to a human operator.

The Agent Loop in Detail

Let's zoom in on each phase of the agent loop and understand what happens.

Phase 1: Observation

The agent perceives its environment. This might be:

  • A user query: "What is the market cap of Tesla?"
  • Sensor data: A robot perceives an obstacle in its path
  • API response: The agent receives feedback from a previous action
  • Internal state: The agent checks its memory and current progress

The observation is encoded as context โ€” usually text โ€” that the agent processes.

Phase 2: Thinking

Given the observation, the agent thinks about what to do. The agent's "thinking" involves:

  • Reasoning: Apply reasoning frameworks (ReAct, CoT) to understand the situation
  • Planning: Generate one or more candidate next actions
  • Evaluation: Estimate which action is most likely to make progress toward the goal
  • Decision: Select the best action

This is where the LLM's reasoning comes in. Modern LLMs are remarkably good at this step when given proper prompting.

Internal vs External Reasoning

The agent can use tokens for 'internal' reasoning (thoughts that don't appear to the user) or 'external' reasoning (visible to observers). Some agent frameworks maximize reasoning steps; others minimize them for latency.

Phase 3: Action

The agent executes the selected action. This might be:

  • Calling a tool: "Search the web for Tesla's market cap"
  • Memory operation: "Store this fact for later retrieval"
  • Communication: "Send this email"
  • Environment modification: "Move the robot forward 1 meter"
  • Special actions: "Ask the user for clarification" (escalation)

The action is typically structured (not free-form text). Modern agent frameworks use function calling โ€” the LLM outputs a structured representation of what it wants to do.

Phase 4: Observation of Results

After the action, the agent observes the results:

  • Success: The tool returned useful data
  • Partial success: The tool returned something, but it's incomplete or noisy
  • Failure: The tool returned an error

The agent feeds this result back into the reasoning loop. If the goal is achieved, it stops. Otherwise, it loops back to thinking.

Termination Conditions

Agents must have clear stopping conditions:

  • Goal achieved: The agent has successfully completed the task
  • Max iterations reached: The agent has taken N steps and should stop
  • Unrecoverable error: The agent has encountered a fatal failure
  • User interruption: The user has stopped the agent
  • Escalation: The agent has decided the task requires human intervention

Design Principle: Always set a maximum number of agent iterations. An agent in an infinite loop is worse than a failed agent. Most tasks should complete in 5-15 steps; if it takes more, the problem decomposition is likely wrong.

Animated Loop Diagram

GOAL Observe Think Act Process Result

Continuous loop until goal achieved or termination condition reached

Reasoning Frameworks: ReAct, Chain-of-Thought, and Beyond

Different situations require different reasoning approaches. Here are the major frameworks and when to use each.

1. Chain-of-Thought (CoT): Making Reasoning Visible

Chain-of-Thought is the simplest reasoning framework. The agent generates intermediate reasoning steps before arriving at a conclusion. Instead of jumping directly to an answer, it shows its work.

Python โ€” Chain-of-Thought Example
# Prompt You are a helpful assistant. Think through this step by step. Question: Alice has 3 apples. She buys 5 more. Her friend gives her 2. How many apples does Alice have now? # Expected output Thought: Let me work through this step by step. - Alice starts with 3 apples - She buys 5 more: 3 + 5 = 8 apples - Her friend gives her 2: 8 + 2 = 10 apples - So Alice has 10 apples total. Answer: Alice has 10 apples.

CoT is helpful because:

  • It improves accuracy on multi-step reasoning tasks (often 5-10% improvement)
  • It makes the agent's thinking transparent and debuggable
  • It helps the LLM avoid rushing to conclusions

However, CoT has limitations:

  • It increases token usage (longer responses)
  • It doesn't improve performance on all tasks (only complex reasoning)
  • It still can't fix hallucinations from lack of knowledge

2. ReAct: Reasoning + Acting (The Agent Framework)

ReAct interleaves reasoning with action. Instead of thinking of everything then acting, the agent thinks, acts, observes results, then thinks again based on what it learned.

Python โ€” ReAct Loop Example
User: "Who won the 2024 Nobel Prize in Physiology or Medicine?" Agent Loop: [Thought 1] I need to find information about the 2024 Nobel Prize in Physiology or Medicine. Let me search for this information. [Action 1] Search: "2024 Nobel Prize Physiology Medicine winner" [Observation 1] The search returned: "The 2024 Nobel Prize in Physiology or Medicine was awarded to Victor Ambros and Gary Ruvkun for the discovery of microRNA and its role in post-transcriptional gene regulation." [Thought 2] Great! I found the answer. The 2024 Nobel Prize was awarded to two scientists: Victor Ambros and Gary Ruvkun. Let me provide this information to the user. [Action 2] Respond to user [Final Response] The 2024 Nobel Prize in Physiology or Medicine was awarded to Victor Ambros and Gary Ruvkun for their discovery of microRNA and its role in post-transcriptional gene regulation.

ReAct is powerful because:

  • Agents can verify their reasoning against reality
  • If reasoning was wrong, they can correct it based on observation
  • Tool calls are motivated by explicit reasoning
  • It's interpretable โ€” you can see why the agent took each action

3. Plan-and-Execute: Two-Phase Reasoning

Some tasks benefit from planning before acting. In Plan-and-Execute, the agent first creates a full plan, then executes it step by step.

Python โ€” Plan-and-Execute Example
User: "Plan a 3-day trip to Tokyo for someone interested in technology and food." [Planning Phase] I'll break this into a plan: Plan: Day 1: - Visit teamLab Borderless (digital art/tech experience) - Lunch at a tech-themed cafรฉ - Visit Akihabara electronics district - Dinner at Michelin-starred restaurant Day 2: - Visit Sony Building - Lunch at conveyor belt sushi - Visit Mori Art Museum (tech architecture) - Dinner at traditional izakaya Day 3: - Visit Robot Restaurant (closing soon, see it!) - Lunch at food market - Visit Roppongi tech district - Dinner recommendations and packing [Execution Phase] Now I'll flesh out each item: Day 1: - teamLab Borderless: Located in Odaiba. 2-3 hours needed. Entry fee: ยฅ3,400. Hours: 10 AM - 7 PM. (Getting more details about current exhibits)...

Use Plan-and-Execute when:

  • The task is complex with multiple steps
  • There are dependencies between steps (some tasks must happen before others)
  • You want to ensure comprehensive coverage of the task
  • You want the user to see and approve the plan before execution

4. Tree-of-Thought: Exploring Multiple Paths

For problems where the solution path is uncertain, Tree-of-Thought explores multiple reasoning branches and selects the most promising ones.

Python โ€” Tree-of-Thought Pseudocode
def tree_of_thought(problem, max_depth=3): """Explore multiple reasoning paths like a game tree.""" root = Node(problem) nodes_to_explore = [root] for depth in range(max_depth): next_nodes = [] for node in nodes_to_explore: # Generate multiple possible next steps candidates = generate_candidates(node.state, k=3) for candidate in candidates: child = Node(candidate) # Evaluate if this path looks promising score = evaluate_promising(child) next_nodes.append((child, score)) # Keep only the top K most promising paths (prune the tree) nodes_to_explore = [n for n, s in sorted(next_nodes, key=lambda x: x[1], reverse=True)][:5] # Return the best final state found return best_solution(nodes_to_explore)

Tree-of-Thought is useful for:

  • Mathematical problem-solving
  • Complex reasoning tasks where multiple solution paths exist
  • Scenarios where the LLM might initially choose a suboptimal path

However, it's more computationally expensive (multiple LLM calls) and better for problems where verification is possible.

Comparison: Which Framework to Use?

Framework When to Use Token Cost Accuracy Gain
Chain-of-Thought Simple reasoning, need transparency ~+20% +5-10%
ReAct Need grounding in tools/reality ~+40% +15-30%
Plan-and-Execute Complex multi-step tasks ~+50% +10-25%
Tree-of-Thought Uncertain solution paths, math ~+200-300% +20-50%

Practical Recommendation: Start with ReAct for most agent applications. It provides excellent results without excessive token cost. Use Tree-of-Thought only for high-stakes problems where accuracy matters more than latency.

Tool Use and Grounding: Connecting Agents to Reality

A language model without tools is just a text predictor. Tools are what make agents real. Tools connect the agent to databases, APIs, sensors, and the outside world.

Why Tools Matter

  • Grounding: Agents can verify knowledge against reality
  • Actions: Agents can modify the world, not just talk about it
  • Freshness: Agents can access real-time information (stock prices, weather)
  • Accuracy: Agents can use specialized tools (calculators, code execution)
  • Delegation: Agents can offload computation to appropriate systems

Types of Tools

API Tools

Call REST APIs, GraphQL endpoints, RPC services. Examples: weather API, stock data, payment processors.

Database Tools

Query databases, retrieve structured data. Examples: SQL queries, document stores, search indices.

Search Tools

Search engines, vector similarity search, RAG. Examples: web search, internal knowledge base.

Compute Tools

Code execution, calculations, simulations. Examples: Python interpreter, math engines, ML inference.

File System Tools

Read/write files, directory operations. Examples: document processing, logs.

Custom Tools

Domain-specific operations. Examples: domain registration, sensor control, payment integration.

Implementing a Simple Tool: Weather Lookup

Python โ€” Tool Definition in LangChain
from langchain.agents import tool import requests @tool def get_weather(location: str) -> str: """Get current weather for a location. Args: location: City name, e.g. "London" or "London, UK" Returns: Current weather conditions and temperature """ try: # Call a weather API (example: Open-Meteo which is free) url = f"https://geocoding-api.open-meteo.com/v1/search?name={location}&count=1" geo_response = requests.get(url).json() if not geo_response.get('results'): return f"Location '{location}' not found" result = geo_response['results'][0] lat, lon = result['latitude'], result['longitude'] # Get weather data weather_url = f"https://api.open-meteo.com/v1/forecast?latitude={lat}&longitude={lon}&current=temperature_2m,weather_code" weather_data = requests.get(weather_url).json() current = weather_data['current'] temp = current['temperature_2m'] return f"Current weather in {location}: {temp}ยฐC, conditions: {current.get('weather_code', 'Clear')}" except Exception as e: return f"Error getting weather: {str(e)}" # The @tool decorator makes this available to the agent # The agent can call it like: get_weather(location="London")

Best Practices for Tool Design

  • Clear Names: Tool names should clearly indicate what they do
  • Clear Docstrings: The agent uses the docstring to decide whether to call the tool
  • Typed Arguments: Include type hints so the agent knows what to pass
  • Error Handling: Tools should never crash; return error messages instead
  • Reasonable Output: Tool output should be in a form the LLM can understand
  • Feedback: Include information in the output that helps the agent decide next steps

Tool Composition

Complex tools can be built from simpler tools. For example, a 'research' tool might internally use web search, then summarize results. This layering makes agents more maintainable.

Common Tool Patterns

Python โ€” Search + Summarize Pattern
@tool def research_topic(topic: str) -> str: """Research a topic and return key findings. Args: topic: The topic to research Returns: Summary of key findings """ # Step 1: Search for information search_results = search_web(topic) if not search_results: return f"No results found for '{topic}'" # Step 2: Summarize results summaries = [r['summary'] for r in search_results[:3]] combined = "\n".join(summaries) # Step 3: Let the LLM synthesize (via calling another tool) synthesis = summarize_text(combined) return synthesis

Tool Chaining: Complex Operations from Simple Tools

Python โ€” Tool Chaining Example
# Scenario: User asks "What was Tesla's Q3 2024 revenue compared to Q3 2023?" # The agent decomposes this into steps: Step 1: Call search_financial_data Result: "Tesla Q3 2024 revenue: $24.18 billion (found in investor relations)" Step 2: Call search_financial_data Result: "Tesla Q3 2023 revenue: $21.45 billion (found in investor relations)" Step 3: Call calculate_comparison Input: (24.18, 21.45) Result: "Q3 2024 revenue was $2.73B more than Q3 2023, a 12.7% year-over-year increase" Step 4: Format and respond to user Result: "Tesla's Q3 2024 revenue of $24.18B represented a 12.7% increase over Q3 2023's $21.45B..."

Error Handling: Making Tools Robust

Python โ€” Robust Tool with Fallbacks
@tool def fetch_document(doc_id: str) -> str: """Fetch a document by ID with graceful error handling. Args: doc_id: Document identifier Returns: Document content or helpful error message """ # Try primary source first try: return fetch_from_database(doc_id) except DatabaseTimeout: # If primary source is slow, try cache try: return fetch_from_cache(doc_id) except CacheError: # If everything fails, explain to the agent return f"Document {doc_id} is temporarily unavailable. Try again in a few moments." except DocumentNotFound: # Clear error message return f"Document {doc_id} does not exist in the system." except PermissionError: # Escalation info return f"You don't have permission to access document {doc_id}. Request access from your manager."

Key Principle: Good tool error messages contain actionable information. Instead of 'Error', say 'Database timeout - retrying would likely succeed' or 'Permission denied - contact [email protected]'.

Planning and Task Decomposition

The difference between a capable agent and a floundering one often comes down to planning. Complex tasks require breaking them into subtasks, understanding dependencies, and executing in the right order.

Task Decomposition: The Core Skill

Given a goal, the agent must ask:

  • What are the subgoals needed to achieve this goal?
  • What information do I need to gather first?
  • Are there any prerequisites or dependencies?
  • What could go wrong at each step?
  • How will I know if I've succeeded?
Python โ€” Task Decomposition Example
User Goal: "Create a marketing strategy for our new AI product targeting SMBs" Decomposition: Level 1 (High-level tasks): 1. Research the SMB market and competitive landscape 2. Define our target audience segments 3. Analyze competitor positioning 4. Develop messaging and value propositions 5. Plan marketing channels and tactics 6. Create a budget and resource allocation plan 7. Define success metrics and KPIs Level 2 (Expanding task 1): 1.1. Define "SMB" for our product (size, industry, geography) 1.2. Research SMB AI adoption rates and needs 1.3. Research SMB tech budgets and purchasing processes 1.4. Analyze market size and growth trends 1.5. Identify pain points SMBs face that our product solves Level 2 (Expanding task 3): 3.1. List 5-10 competitors in the SMB AI space 3.2. For each competitor: pricing, features, target market, messaging 3.3. Identify competitive gaps (features they lack, markets they ignore) 3.4. Identify our unique advantages 3.5. Determine how to position against each competitor ...and so on

Hierarchical Planning

Agents work best with hierarchical plans โ€” breaking problems from abstract to concrete:

Level 1
High-level goal
Level 2
Subgoals
Level 3
Concrete tasks
Level 4
Actions

Handling Dependencies

Not all subtasks can be done in parallel. Some depend on results from others:

Python โ€” Dependency-Aware Planning
# Tasks and their dependencies tasks = { "research_market": { "depends_on": [], # No dependencies "effort": "high" }, "define_target_audience": { "depends_on": ["research_market"], # Must do market research first "effort": "medium" }, "analyze_competitors": { "depends_on": ["define_target_audience"], # Need to know audience first "effort": "high" }, "develop_messaging": { "depends_on": ["define_target_audience", "analyze_competitors"], "effort": "medium" }, "plan_channels": { "depends_on": ["develop_messaging"], "effort": "medium" }, "allocate_budget": { "depends_on": ["plan_channels"], # Need to know channels first "effort": "low" } } # Execution order respects dependencies: # 1. research_market (no deps) # 2. define_target_audience (after 1) # 3. analyze_competitors (after 2) # 4. develop_messaging (after 3) # 5. plan_channels (after 4) # 6. allocate_budget (after 5)

Adaptive Planning: Replanning When Things Change

Agents should monitor if their plan is still valid and replan if needed:

  • Assumption violation: A fact the plan depends on turned out to be wrong
  • New information: Data discovered mid-plan suggests a better approach
  • Tool failure: A tool the plan relied on is unavailable
  • Resource constraint: Time or budget is running out

Replanning Threshold

Agents should replan if the cost of replanning is less than the cost of continuing with a suboptimal plan. This is a heuristic decision, not always optimal.

Memory Management: Building Agent Context

Memory is the agent's long-term knowledge. Without memory, each agent interaction is isolated. With good memory, agents build understanding over time.

Types of Agent Memory

Episodic Memory

What happened during this interaction. Conversation history, past actions, outcomes.

Semantic Memory

Facts and knowledge. Learned during training or from documents. Indexed for retrieval.

Procedural Memory

How to do things. Learned skills, patterns, heuristics. Implicit in agent behavior.

State Memory

Current progress. What goals are being pursued, what's been completed, what's next.

Episodic Memory: Conversation History

The simplest form of memory is keeping the full conversation history:

Python โ€” Basic Conversation Memory
from langchain.memory import ConversationBufferMemory from langchain.agents import AgentExecutor, create_openai_tools_agent from langchain_openai import ChatOpenAI llm = ChatOpenAI(model="gpt-4") # Create memory that stores all messages memory = ConversationBufferMemory( memory_key="chat_history", return_messages=True ) # Agents can access this memory # Problem: Memory grows infinitely, eventually exceeds context window

Problem: Full conversation history grows too large. Eventually exceeds the LLM's context window.

Episodic Memory: Summarization and Sliding Window

Instead of storing everything, summarize old conversations or use a sliding window:

Python โ€” Summarized Memory
from langchain.memory import ConversationSummaryMemory from langchain_openai import ChatOpenAI llm = ChatOpenAI(model="gpt-4") # Memory that summarizes old conversations memory = ConversationSummaryMemory( llm=llm, buffer="Last 10 messages are stored verbatim.\ Older messages are summarized to key points." ) # Example: # After 50 messages, the memory might look like: # [Summary of messages 1-40 in 2-3 sentences] # [Messages 41-50 verbatim] # This keeps recent context while being compact

Semantic Memory: Vector-Based Retrieval

For large knowledge bases, agents use vector embeddings and semantic search:

Python โ€” Vector-Based Memory (RAG)
from langchain.embeddings import OpenAIEmbeddings from langchain.vectorstores import Chroma from langchain.text_splitter import CharacterTextSplitter # Create vector embeddings of documents documents = load_documents("company_knowledge_base/") # PDFs, docs, etc. text_splitter = CharacterTextSplitter(chunk_size=1000) texts = text_splitter.split_documents(documents) embeddings = OpenAIEmbeddings() vector_store = Chroma.from_documents(texts, embeddings) # Agent can now retrieve relevant documents def search_memory(query: str) -> str: """Search the knowledge base.""" results = vector_store.similarity_search(query, k=3) return "\n".join([r.page_content for r in results]) # Agent uses this for grounding knowledge: # User: "What's our refund policy?" # Agent calls: search_memory("refund policy") # Gets: [Relevant policy documents] # Responds with accurate information grounded in reality

State Memory: Progress Tracking

Agents need to track what they're doing and what they've accomplished:

Python โ€” State Management
class AgentState: """Track agent progress toward a goal.""" def __init__(self, goal: str): self.goal = goal self.completed_steps = [] self.pending_steps = [] self.learned_facts = {} self.failures = [] self.max_iterations = 15 self.iteration_count = 0 def add_completed_step(self, step: str, result: str): """Record a completed step.""" self.completed_steps.append({ "step": step, "result": result, "timestamp": datetime.now() }) def mark_failure(self, step: str, error: str, retry_count: int): """Track failures for debugging.""" self.failures.append({ "step": step, "error": error, "retry_count": retry_count }) def should_continue(self) -> bool: """Check if agent should keep going.""" if self.iteration_count >= self.max_iterations: return False if too_many_failures(self.failures): return False return True

Conversation Memory + Vector Memory Integration

Python โ€” Hybrid Memory System
class HybridMemory: """Combine conversation history with semantic search.""" def __init__(self, llm, vector_store): self.conversation_memory = ConversationSummaryMemory(llm=llm) self.semantic_memory = vector_store # Vector database def remember(self, input_text: str) -> str: """Get relevant context for the agent. Returns both: 1. Recent conversation history (episodic) 2. Relevant documents from knowledge base (semantic) """ # Get recent conversation conversation = self.conversation_memory.load_memory_variables()["history"] # Get semantically similar documents relevant_docs = self.semantic_memory.similarity_search(input_text, k=3) # Combine them context = f""" Current Conversation Context: {conversation} Relevant Knowledge Base Documents: {relevant_docs} """ return context # Usage in agent: # Agent receives: # 1. The current user query # 2. Recent conversation history # 3. Semantically relevant documents # 4. This combination provides grounded, contextual understanding

Memory Design Principle: Different types of information need different storage. Recent conversations stay in full detail. Old conversations get summarized. Semantic facts go in vector databases. Current progress goes in state objects.

Multi-Agent Systems: Specialization and Coordination

Some problems are too complex for a single agent. Multi-agent systems use multiple specialized agents that collaborate (or compete) to solve problems.

Why Multi-Agent Systems?

  • Specialization: Different agents can specialize in different domains (research, writing, analysis)
  • Scalability: Agents can work in parallel on different subproblems
  • Robustness: If one agent fails, others can continue
  • Diversity: Multiple agents with different reasoning styles may find better solutions
  • Realistic: Mirrors how humans solve complex problems (teamwork)

Types of Multi-Agent Architectures

Pipeline Architecture

Agents process in sequence: Agent A's output becomes Agent B's input. Like an assembly line.

Hierarchical Architecture

Manager agent decomposes tasks and delegates to worker agents. Manager coordinates results.

Democracy/Voting Architecture

Multiple agents propose solutions. Best is selected by consensus or voting.

Pool Architecture

All agents work on the same problem in parallel. Results are merged or ranked.

Example: Research Paper Writing Pipeline

Python โ€” Multi-Agent Writing Pipeline
# Agent 1: Researcher Agent researcher = Agent( role="Research Specialist", tools=["search_papers", "fetch_pdf", "extract_key_findings"], goal="Find relevant papers and extract key findings" ) # Agent 2: Analysis Agent analyst = Agent( role="Data Analyst", tools=["analyze_trends", "identify_gaps", "compare_methodologies"], goal="Analyze research trends and identify knowledge gaps" ) # Agent 3: Writing Agent writer = Agent( role="Technical Writer", tools=["outline_paper", "draft_sections", "citations"], goal="Write a comprehensive research paper" ) # Coordinator: Managers the flow coordinator = Agent( role="Project Manager", agents=[researcher, analyst, writer], goal="Ensure all agents work together effectively" ) # Pipeline: # 1. Researcher finds papers on topic # 2. Analyst extracts insights and identifies gaps # 3. Writer uses insights to draft the paper # 4. Coordinator validates progress and ensures quality

Consensus-Based Decision Making

When multiple agents disagree, consensus mechanisms help:

Python โ€” Consensus Agent System
def consensus_decision(agents: List[Agent], question: str) -> str: """Get consensus from multiple agents. Useful when you want multiple perspectives on a decision. """ # Step 1: Get each agent's answer answers = [] for agent in agents: response = agent.answer(question) answers.append(response) # Step 2: Check for agreement if all(a == answers[0] for a in answers): return answers[0] # Full consensus # Step 3: If disagreement, debate debate_history = [] for round_num in range(3): # 3 rounds of debate for i, agent in enumerate(agents): # Each agent sees others' views and can change mind other_views = [answers[j] for j in range(len(agents)) if j != i] updated = agent.reconsider(question, other_views) answers[i] = updated debate_history.append((agent.role, updated)) # Step 4: Final vote votes = Counter(answers) return votes.most_common(1)[0][0] # Return most common answer # Usage: # agents = [research_agent, coding_agent, design_agent] # decision = consensus_decision(agents, "Should we use React or Vue?") # Even if one agent has biased preferences, consensus gives balanced decision

CrewAI: A Framework for Multi-Agent Teams

CrewAI is a Python framework that simplifies building multi-agent teams:

Python โ€” CrewAI Example
from crewai import Agent, Task, Crew, Process # Define specialized agents researcher = Agent( role="Market Research Analyst", goal="Find and analyze market opportunities", backstory="You are an expert at analyzing markets and identifying trends", tools=[search_tool, analyze_tool], verbose=True ) writer = Agent( role="Content Writer", goal="Create compelling market analysis reports", backstory="You are an excellent writer skilled at explaining complex data", tools=[write_tool, format_tool], verbose=True ) # Define tasks for each agent research_task = Task( description="Research the AI agent market. Find size, growth, key players, trends.", agent=researcher ) writing_task = Task( description="Write a compelling market analysis report based on the research", agent=writer ) # Create crew (team) with a process crew = Crew( agents=[researcher, writer], tasks=[research_task, writing_task], process=Process.sequential, # Tasks done in order verbose=True ) # Execute result = crew.kickoff() print(result)

Emergent Behavior in Multi-Agent Systems

Interestingly, multi-agent systems can exhibit emergent behavior โ€” collective properties that arise from agent interactions but aren't explicitly programmed.

Examples:

  • A team of agents independently working on subtasks naturally learns to avoid duplicate work
  • Agents sharing findings naturally create higher-quality synthesis than any single agent
  • Weak individual agents in a diverse team can outperform strong individual agents

The Diversity Hypothesis

A team of diverse agents (different backgrounds, reasoning styles, knowledge) often outperforms a team of identical copies of the best agent. This mirrors human team dynamics and suggests depth of expertise matters less than diversity of perspective.

Agent Frameworks: Tools for Building Agents

Building agents from scratch is complex. Frameworks abstract away boilerplate and best practices.

Major Agent Frameworks

LangChain
Most popular, best docs
AutoGPT
Experimental, research-focused
CrewAI
Multi-agent, growing adoption
Anthropic SDK
Native Claude integration

LangChain: The Standard

LangChain is the most widely used agent framework. It provides:

  • Predefined agent types (ReAct, OpenAI Functions, etc.)
  • A large library of tools and integrations
  • Memory management abstractions
  • Chain-of-thought execution
Python โ€” LangChain Agent Quickstart
from langchain_openai import ChatOpenAI from langchain.agents import AgentExecutor, create_openai_tools_agent from langchain.tools import tool from langchain import hub # Define tools @tool def calculator(expression: str) -> str: """Evaluate a mathematical expression.""" return str(eval(expression)) @tool def web_search(query: str) -> str: """Search the web.""" # Actual implementation would call a search API return f"Results for '{query}'..." # Get the prompt template prompt = hub.pull("hwchase17/openai-tools-agent") # Create the agent llm = ChatOpenAI(model="gpt-4") tools = [calculator, web_search] agent = create_openai_tools_agent(llm, tools, prompt) executor = AgentExecutor.from_agent_and_tools( agent=agent, tools=tools, verbose=True, max_iterations=10 ) # Run the agent result = executor.invoke({"input": "What is 42 * 3? Also, search for recent AI news."}) print(result["output"])

Comparison: LangChain vs CrewAI

Feature LangChain CrewAI
Maturity Mature, 2+ years Newer, rapidly evolving
Single vs Multi-Agent Primarily single-agent Multi-agent focus
Learning Curve Steep, lots to learn Gentler, simpler API
Flexibility Highly flexible, low-level Opinionated, high-level
Community Large, many examples Growing fast
Best For Custom solutions, production Teams, prototyping

Building Your First Agent: Implementation Guide

Let's build a complete, working agent from scratch. This agent researches companies and generates investment summaries.

Python โ€” Complete Working Agent
#!/usr/bin/env python3 """Build a simple investment research agent.""" from langchain_openai import ChatOpenAI from langchain.agents import AgentExecutor, create_openai_tools_agent from langchain.tools import tool from langchain import hub # Setup llm = ChatOpenAI(model="gpt-4", temperature=0) # Define tools the agent can use @tool def search_financial_data(company_name: str) -> str: """Search for a company's financial data. Args: company_name: Name of the company (e.g., 'Apple Inc', 'Tesla') Returns: Financial summary including revenue, profit, growth rates """ # In real implementation, query financial database or API data = { "Apple": "Revenue: $383B (2023), Net Income: $100B, Growth: 5%", "Tesla": "Revenue: $82B (2023), Net Income: $13B, Growth: 20%", "Microsoft": "Revenue: $211B (2023), Net Income: $72B, Growth: 15%" } return data.get(company_name, "Company not found in database") @tool def search_recent_news(company_name: str) -> str: """Search for recent news about a company. Args: company_name: Company name Returns: Summary of recent news """ # In real implementation, query news API return f"Recent news about {company_name}: Strong Q4 results, new product launches, market expansion" @tool def get_industry_data(industry: str) -> str: """Get industry trends and benchmarks. Args: industry: Industry name (e.g., 'technology', 'healthcare') Returns: Industry metrics and trends """ return f"Industry {industry}: Average growth 12%, profit margin 15%, key trends: AI adoption" @tool def calculate_ratios(revenue: float, net_income: float) -> str: """Calculate financial ratios. Args: revenue: Company revenue net_income: Company net income Returns: Calculated financial ratios """ profit_margin = (net_income / revenue) * 100 return f"Profit Margin: {profit_margin:.1f}%" # Get prompt template prompt = hub.pull("hwchase17/openai-tools-agent") # Create agent tools = [search_financial_data, search_recent_news, get_industry_data, calculate_ratios] agent = create_openai_tools_agent(llm, tools, prompt) # Create executor with safety limits executor = AgentExecutor.from_agent_and_tools( agent=agent, tools=tools, verbose=True, max_iterations=10, early_stopping_method="force" # Stop if loops too much ) # Run the agent if __name__ == "__main__": result = executor.invoke({ "input": "Compare Apple and Tesla. Which is a better investment right now?" }) print("\n=== FINAL RESPONSE ===") print(result["output"])

Key Implementation Considerations

  • Error Handling: Tools should never crash; return error messages instead
  • Rate Limiting: Don't overwhelm APIs; add delays if needed
  • Timeout Protection: Set timeouts on tool calls to prevent hangs
  • Logging: Log all agent decisions for debugging and auditing
  • Cost Control: Monitor API costs; set spending limits if needed
  • Testing: Test tools independently before adding to agent

Code Examples: Multiple Frameworks

Example 1: ReAct Agent with LangChain

Python โ€” ReAct Pattern
# ReAct = Reasoning + Acting # Agent alternates between: # 1. Thought: Reason about what to do # 2. Action: Call a tool # 3. Observation: See the result # 4. Back to Thought (until goal reached) from langchain.agents import AgentExecutor, create_react_agent from langchain_openai import ChatOpenAI from langchain.tools import tool @tool def get_current_date() -> str: """Get today's date.""" from datetime import datetime return datetime.now().strftime("%Y-%m-%d") @tool def calculate(expression: str) -> str: """Calculate math expressions.""" return str(eval(expression)) llm = ChatOpenAI(model="gpt-4") tools = [get_current_date, calculate] # ReAct automatically structures thoughts and actions agent = create_react_agent(llm, tools, prompt_template) executor = AgentExecutor.from_agent_and_tools(agent=agent, tools=tools, verbose=True) result = executor.invoke({"input": "What date is it? Add 7 days."})

Example 2: Tool-Calling Agent (Native Function Calls)

Python โ€” Function Calling Agent
# Modern LLMs support native function calling # This is simpler than ReAct for many use cases from langchain_openai import ChatOpenAI import json llm = ChatOpenAI(model="gpt-4") # Define tools as JSON schemas tools = [ { "type": "function", "function": { "name": "get_weather", "description": "Get weather for a location", "parameters": { "type": "object", "properties": { "location": {"type": "string", "description": "City name"} }, "required": ["location"] } } } ] # Call LLM with tools response = llm.invoke([ {"role": "user", "content": "What's the weather in London?"} ], tools=tools) # LLM returns structured function calls # Agent executes them and provides results

Example 3: Multi-Agent System with CrewAI

Python โ€” CrewAI Multi-Agent Example
from crewai import Agent, Task, Crew, Process from crewai_tools import SerperDevTool, ScrapeWebsiteTool # Create agents with different specializations researcher = Agent( role="Senior Researcher", goal="Conduct thorough research and summarize findings", backstory="You are an expert researcher with deep knowledge of technology trends", tools=[SerperDevTool()], verbose=True ) analyst = Agent( role="Business Analyst", goal="Analyze information and identify key insights", backstory="You are a seasoned analyst who identifies business opportunities", tools=[ScrapeWebsiteTool()], verbose=True ) # Create tasks research = Task( description="Research AI agent adoption in enterprise", agent=researcher, expected_output="Detailed research report on AI agents in enterprises" ) analysis = Task( description="Analyze the research and identify market opportunities", agent=analyst, expected_output="Analysis of market opportunities and trends" ) # Create crew crew = Crew( agents=[researcher, analyst], tasks=[research, analysis], process=Process.sequential, verbose=True ) # Execute result = crew.kickoff() print(result)

Reasoning Frameworks Comparison

Different frameworks suit different problems. Here's how to choose:

Framework ReAct Plan-and-Execute Tree-of-Thought
How It Works Interleave reasoning + acting. Think โ†’ Act โ†’ Observe โ†’ Think Create full plan first, then execute step by step Explore multiple reasoning paths, keep best ones
Best For Tasks with tools & feedback loops Complex multi-step tasks with clear structure Math, logic, planning where wrong path is costly
Token Cost Moderate (~1.5x base) Higher (~2x base) Very High (~5-10x base)
Accuracy Improvement +15-30% on grounded tasks +10-25% on complex tasks +30-50% on math/logic
Latency Can be quick if few steps Slower (planning + execution phases) Much slower (explores many paths)
Debuggability Very good (see each step) Good (plan is visible) Hard (many parallel paths)
Failure Recovery Excellent (adapts to failures) Good (can replan if needed) Built-in (tries other paths)

Decision Tree: Which Framework?

Do you need real-time tool interaction? ▼

YES: Use ReAct. Your agent will benefit from observing tool results and adapting.

NO: Continue to next question.

Is the task primarily mathematical or logic-heavy? ▼

YES: Use Tree-of-Thought if cost/time allows, otherwise ReAct.

NO: Continue to next question.

Does the task have clear sequential steps? ▼

YES: Use Plan-and-Execute. Create a plan first.

NO: Use ReAct (most flexible).

Memory Patterns in Production

How do production agents manage memory at scale?

Pattern 1: Summary + Sliding Window

Keep recent messages verbatim, summarize older ones:

Python โ€” Memory: Summary + Window
class SmartMemory: def __init__(self, keep_recent=10, max_summary_length=500): self.messages = [] self.keep_recent = keep_recent self.summary = "" self.max_summary_length = max_summary_length def add_message(self, role, content): self.messages.append({"role": role, "content": content}) # If we have too many messages, summarize old ones if len(self.messages) > self.keep_recent * 2: self._compress_memory() def _compress_memory(self): # Keep last N messages, summarize the rest recent = self.messages[-self.keep_recent:] old = self.messages[:-self.keep_recent] if old: # Summarize old messages (in real implementation, use LLM) old_text = "\n".join([m["content"] for m in old]) self.summary = summarize_text(old_text) self.messages = recent def get_context(self): """Get memory in a compact form.""" if self.summary: context = f"[Previous Summary]\n{self.summary}\n\n" else: context = "" context += "[Recent Messages]\n" for msg in self.messages: context += f"{msg['role']}: {msg['content']}\n" return context

Pattern 2: Hierarchical Memory

Different types of information at different levels:

Python โ€” Hierarchical Memory
class HierarchicalMemory: def __init__(self): # Level 1: Very recent (full detail) self.current_session = [] # Level 2: Recent history (compressed) self.recent_history = [] # Level 3: Semantic/vector store self.semantic_memory = VectorStore() # Level 4: Long-term facts (database) self.long_term_facts = Database() def retrieve(self, query): """Retrieve relevant information for a query.""" results = [] # 1. Check current session (most relevant) current_relevant = search(self.current_session, query) results.extend(current_relevant) # 2. Check recent history if needed if len(results) < 3: history_relevant = search(self.recent_history, query) results.extend(history_relevant) # 3. Do semantic search in vector store semantic_relevant = self.semantic_memory.search(query, k=3) results.extend(semantic_relevant) return results[:5] # Return top 5 results

Pattern 3: Fact-Based Memory

Extract and store discrete facts from conversations:

Python โ€” Fact Extraction Memory
class FactMemory: def __init__(self): self.facts = [] # List of {fact, confidence, timestamp, source} def extract_facts(self, text): """Extract facts from text using LLM.""" prompt = f""" Extract key facts from this text. For each fact, rate confidence (0-1). Format: {{fact, confidence}} Text: {text} """ response = llm.invoke(prompt) # Parse response and store facts for fact_line in response.split("\n"): fact, confidence = parse_fact(fact_line) if confidence > 0.7: self.facts.append({ "fact": fact, "confidence": confidence, "timestamp": datetime.now(), "source": "conversation" }) def query_facts(self, question): """Find relevant facts for a question.""" # Score each fact for relevance scored = [] for fact in self.facts: relevance = compute_similarity(question, fact["fact"]) scored.append((fact, relevance)) # Return high-confidence, relevant facts return [f for f, s in sorted(scored, key=lambda x: x[1], reverse=True) if f["confidence"] > 0.7][:5]

Evaluating Agent Performance

How do you know if your agent is good? Evaluation is critical but non-obvious.

Evaluation Metrics

Task Completion

Did the agent achieve its goal? Binary yes/no or success rate %.

Tool Usage Efficiency

How many tool calls were needed? Fewer is better (lower cost, latency).

Reasoning Quality

Are the agent's reasoning steps logical and justified?

Error Recovery

When the agent fails, does it recover gracefully?

Grounding

Does the agent use tools appropriately? Avoids hallucinations?

Alignment

Does agent behavior match intended values and constraints?

Building an Evaluation Framework

Python โ€” Simple Agent Evaluation
from dataclasses import dataclass from typing import List import numpy as np @dataclass class EvaluationMetric: name: str value: float # 0-1 score weight: float # importance weight class AgentEvaluator: def __init__(self): self.test_cases = [] self.results = [] def add_test_case(self, input_text: str, expected_output: str): """Add a test case.""" self.test_cases.append({ "input": input_text, "expected": expected_output }) def evaluate_agent(self, agent): """Run evaluation on agent.""" metrics = [] for test_case in self.test_cases: # Run agent result = agent.run(test_case["input"]) # Metric 1: Task Completion completed = result["success"] # Boolean metrics.append(EvaluationMetric( name="task_completion", value=float(completed), weight=0.4 )) # Metric 2: Token Efficiency token_ratio = result["tokens_used"] / result["expected_tokens"] efficiency = 1.0 / min(token_ratio, 3.0) # Cap penalty at 3x metrics.append(EvaluationMetric( name="efficiency", value=efficiency, weight=0.2 )) # Metric 3: Reasoning Quality (evaluated by another LLM) reasoning_score = evaluate_reasoning(result["reasoning"], test_case["expected"]) metrics.append(EvaluationMetric( name="reasoning_quality", value=reasoning_score, weight=0.2 )) # Metric 4: Tool Usage Appropriateness appropriate = evaluate_tool_usage(result["tool_calls"], test_case["input"]) metrics.append(EvaluationMetric( name="tool_appropriateness", value=appropriate, weight=0.2 )) # Calculate weighted average total_weight = sum(m.weight for m in metrics) weighted_score = sum(m.value * m.weight for m in metrics) / total_weight return { "overall_score": weighted_score, "metrics": metrics, "pass_rate": sum(1 for m in metrics if m.name == "task_completion") / len(metrics) }

Common Evaluation Pitfalls

  • Only measuring final output: Good reasoning matters even if output is slightly off
  • Ignoring token cost: An agent that solves everything in 10x tokens is impractical
  • Not testing edge cases: Test error conditions, ambiguous inputs, adversarial examples
  • Overfitting to benchmarks: Improve on the actual task, not the test set
  • Ignoring human judgment: Sometimes LLM evaluation misses important aspects

Pro Tip: Build a comprehensive test suite with diverse examples. Include normal cases, edge cases, and adversarial cases. Evaluate not just whether the agent gets the 'right' answer, but whether it reasons soundly.

Best Practices for Building Agents

Design Principles

  • Explicit Goals: Agents should have clear, measurable goals. Vague goals lead to vague behavior.
  • Bounded Reasoning: Set max iterations. An infinite loop is worse than failure.
  • Graceful Degradation: When things go wrong, agents should fail safely and informatively.
  • Observable Behavior: Log and monitor agent decisions. Black boxes are dangerous.
  • Human Oversight: Critical actions should require human approval or logging.

Prompt Engineering for Agents

Python โ€” Good Agent Prompt
You are a helpful research assistant. Your goal is to answer questions accurately using available tools. Follow these guidelines: 1. Think carefully about what tools to use 2. Use tools only when necessary 3. If a tool returns an error, try a different approach 4. Never make up information - always ground claims in tool results 5. Show your reasoning step by step 6. If you're unsure, ask clarifying questions rather than guessing Available tools: [web_search, database_query, calculator] Maximum steps: 10 (to prevent infinite loops)

Error Handling and Robustness

Defensive Tool Design

Every tool should handle errors gracefully. Instead of failing, return a helpful message explaining what went wrong and suggesting next steps. Example: Instead of "Error: HTTP 500", return "The data service is temporarily unavailable. Try again in a moment, or I can use cached data from 2 hours ago."

Cost Management

  • Monitor API Costs: Track LLM API calls; they're expensive at scale
  • Use Caching: Cache tool results when possible (e.g., weather for the same location)
  • Optimize Token Usage: Shorter prompts, relevant context only
  • Batch Operations: Process multiple requests together when possible
  • Cheaper Models: Use cheaper models for simple tasks, save expensive models for complex reasoning

Testing and Iteration

  • Start with a simple agent, add complexity gradually
  • Test each tool independently before integrating
  • Build test suites covering normal and edge cases
  • Use logging extensively during development
  • Get feedback from actual users early

The Prototype-to-Production Journey

A simple agent that works is better than a complex agent that doesn't. Start simple, test thoroughly, add features based on real user feedback. Many successful agents are surprisingly simple.

Interview Questions on AI Agents

If you're interviewing for agent-focused roles, you might encounter these questions:

1. What's the difference between an LLM and an AI agent? ▼
An LLM is a tool for generating text; an agent is a system that uses tools (including LLMs) to achieve goals. LLMs are reactive (given input, produce output); agents are proactive (pursue goals over time). Agents have memory, use reasoning, call external tools, and iterate toward success. A chatbot using an LLM is not an agent. A system that decomposes a goal into subtasks, calls APIs, learns from results, and adapts is an agent.
2. Explain ReAct. Why is it effective? ▼
ReAct = Reasoning + Acting. The agent interleaves reasoning (thinking steps) with acting (tool calls). Instead of planning everything then executing, it thinks, acts, observes results, and thinks again. It's effective because: (1) Agents can verify reasoning against reality, catching hallucinations. (2) Observations inform future reasoning. (3) It's interpretable - you see why the agent took each action. (4) It handles uncertain environments better. The cycle is: Thought โ†’ Action โ†’ Observation โ†’ Thought...
3. How would you design an agent to research companies for investment? ▼
I'd structure it as: (1) Define clear goal - 'Analyze company X and recommend buy/hold/sell'. (2) Define tools - financial data API, news search, analyst reports, calculation. (3) Design reasoning - Plan-and-Execute or ReAct. (4) Build memory - store findings, avoid re-fetching. (5) Create evaluation criteria - completeness, accuracy, cost/efficiency. (6) Add safety - human review before making decisions. (7) Test extensively with past companies where the answer is known.
4. What are the main challenges in building production agents? ▼
Key challenges: (1) Reliability - agents must fail gracefully and be debuggable. (2) Cost - LLM API calls are expensive; must optimize token usage. (3) Latency - agent loops are slower than single LLM calls. (4) Tool availability - not all tools work reliably; need fallbacks. (5) Hallucinations - even with grounding, agents can make up information. (6) Scope creep - agents can go off-task if goals aren't clear. (7) Evaluation - hard to measure agent quality; no single metric captures it all.
5. How would you make an agent more robust against errors? ▼
Multi-level approach: (1) Tool level - each tool has error handling and returns helpful messages. (2) Agent level - agent detects failures and tries alternative approaches. (3) Fallbacks - if primary tool fails, try cached data or simpler tool. (4) Escalation - if agent can't solve it, clearly escalate to human. (5) Monitoring - log all decisions for debugging. (6) Timeouts - prevent infinite loops. (7) Validation - check outputs before using them. (8) User feedback - collect user corrections to improve agent.
6. Compare ReAct, Plan-and-Execute, and Tree-of-Thought approaches. ▼
ReAct: Interleave reasoning and acting. Good for tool-based tasks. Moderate cost. Excellent recovery from failures. Best for most practical applications. Plan-and-Execute: Create full plan first, then execute. Good for complex multi-step tasks. Higher cost. Requires upfront planning overhead. Tree-of-Thought: Explore multiple reasoning paths. Highest cost. Best for math/logic. Can backtrack. Use when accuracy >> cost. My recommendation: Start with ReAct; it's the sweet spot for most agents.
7. What are the main types of memory in agents? How would you implement them? ▼
Four types: (1) Episodic - conversation history. Keep recent messages verbatim, summarize old ones. (2) Semantic - facts and knowledge. Use vector databases for similarity search. (3) Procedural - skills/patterns. Implicit in agent behavior and prompting. (4) State - current progress. Explicit tracking of goals and completed steps. Implementation: Hybrid approach - conversation buffer for recent, vector store for semantic, state object for progress tracking.
8. Design a multi-agent system for content creation. ▼
Structure: (1) Researcher agent - finds relevant information, sources, trends. (2) Writer agent - drafts content using research. (3) Editor agent - reviews for clarity and accuracy. (4) SEO agent - optimizes for search. Coordination: Manager agent decomposes 'Create blog post about AI agents' into subtasks for each agent. Handles dependencies (research before writing, writing before editing). Aggregates results. Alternative: Hierarchical - manager delegates, workers execute, manager synthesizes.
9. How do you evaluate whether an agent is working well? ▼
Multi-dimensional evaluation: (1) Task success rate - does it achieve its goal? (2) Token efficiency - how expensive per task? (3) Reasoning quality - are steps logical? (4) Tool appropriateness - does it use tools correctly? (5) Error recovery - graceful failures? (6) Alignment - does it follow guidelines? Build test suites with diverse cases. Use another LLM to evaluate reasoning. Monitor costs. Collect user feedback. A 'good' agent isn't perfect; it's reliable, understandable, and delivers value.
10. What's your opinion on autonomous agents vs. human-in-the-loop? ▼
Current state: Fully autonomous agents are not yet reliable for critical decisions. Human-in-the-loop is the practical approach - agent proposes, human reviews/approves. Future: As agents improve, more autonomy is possible. My take: The right model depends on stakes. Autonomous is fine for low-stakes (recommendations). For high-stakes (financial, medical, legal), always include human review. Transparency matters - users should understand agent reasoning.

Frequently Asked Questions

Can an LLM be an agent? ▼
Not by itself. An LLM is a text generator. It becomes part of an agent when combined with tools, memory, planning, and a control loop. The agent uses the LLM as a reasoning component, but the agent is the larger system.
How many iterations should an agent take? ▼
Usually 3-10. If an agent needs more than 15-20 steps to solve a task, the problem decomposition is likely wrong. Set max_iterations as a safety valve. Most well-designed agents should complete tasks in 5 steps or fewer.
What happens if an agent calls the same tool multiple times? ▼
This is often a sign the agent is stuck in a loop. (1) Add loop detection - track tool calls and bail if a tool is called repeatedly. (2) Improve tool error messages - help the agent understand why the previous call failed. (3) Check tool outputs - are they helpful? (4) Adjust the prompt - make the goal clearer. (5) Simplify task - break it into smaller pieces.
How do I make agents faster? ▼
Latency optimization: (1) Use cheaper/faster models for simple tasks. (2) Reduce reasoning steps - not all tasks need Chain-of-Thought. (3) Cache tool results - reuse recent data. (4) Parallel tool calls - call multiple tools simultaneously. (5) Shorter prompts - relevant context only. (6) Tool optimization - make tools faster. In practice, agents are 2-5x slower than single LLM calls but provide much higher quality.
Can agents learn from past interactions? ▼
Yes, but not automatically. (1) Store past interactions in memory/database. (2) Extract lessons/facts from past interactions. (3) Use those lessons to improve future behavior. This is different from traditional ML learning. It's more like human learning - reflecting on past experiences. Some frameworks like LangChain have memory mechanisms that support this.
How do I prevent agent hallucinations? ▼
Key strategies: (1) Grounding - require tool use for factual claims. (2) Verification - have agent check its own outputs. (3) Clear instructions - 'Never make up information; always use tools'. (4) Tool design - tools return structured, verifiable data. (5) Testing - hallucination detection should be part of evaluation. (6) Monitoring - log and review all outputs. You can't eliminate hallucinations; you can reduce and detect them.
What's the difference between agents and workflows? ▼
Workflows are fixed sequences of steps. Agents are flexible - they decide what to do based on the situation. Workflow: Step 1 โ†’ Step 2 โ†’ Step 3. Agent: Observe โ†’ Decide what to do โ†’ Do it โ†’ Observe โ†’ Decide next step. Agents are more powerful but more complex. Use workflows for deterministic tasks; agents for reasoning-intensive tasks.
How do I deploy agents to production? ▼
Key considerations: (1) Containerize - use Docker. (2) API wrapper - expose agent as HTTP endpoint. (3) Monitoring - log all decisions, monitor costs/latency. (4) Rate limiting - protect against abuse. (5) Error handling - graceful failures with user-friendly messages. (6) Updates - how do you push new prompts/tools? (7) Scaling - agents are compute-intensive; plan for horizontal scaling. (8) User feedback loop - collect feedback to improve the agent.
How do I test agents? ▼
Multi-level testing: (1) Unit tests - test each tool independently. (2) Integration tests - test agent with real tools. (3) Benchmark tests - does agent solve example problems? (4) Edge case tests - how does agent handle unusual inputs? (5) Regression tests - does agent still work after changes? (6) A/B tests - compare agent versions. (7) User acceptance tests - real users try the agent. Start with unit tests; build up test coverage iteratively.
What are the ethical concerns with agents? ▼
Important concerns: (1) Autonomy - agents making decisions with real consequences need careful design. (2) Transparency - users should understand what agent can/can't do. (3) Bias - agents inherit biases from training data. (4) Privacy - agents accessing sensitive data. (5) Security - agents calling external APIs. (6) Liability - who's responsible if agent fails? Build in human oversight, make reasoning transparent, handle errors gracefully, and respect user data.

Practical Exercises

Exercise 1: Build a Simple Tool

Create a Python function that acts as a tool. It should have a clear docstring, take typed arguments, and return structured output. Example: a tool that fetches the current time in different timezones.

Python โ€” Starter Code
@tool def get_timezone_time(timezone: str) -> str: """Get current time in a timezone. Args: timezone: Timezone name (e.g., 'US/Eastern', 'Europe/London') Returns: Current time in that timezone """ from datetime import datetime import pytz tz = pytz.timezone(timezone) current_time = datetime.now(tz).strftime("%Y-%m-%d %H:%M:%S") return f"Current time in {timezone}: {current_time}"

Exercise 2: Write Agent Prompts

Write two versions of a system prompt - one for a bad agent (vague goals, no constraints) and one for a good agent (clear goals, error handling, tool usage guidelines). Compare them.

Python โ€” Starter Code
# BAD: Vague agent prompt You are a helpful assistant. Be nice and try your best. # GOOD: Clear agent prompt You are a financial research assistant. Your goal is to analyze companies and provide investment recommendations. Use the available tools (financial_data, news_search) to ground your analysis in reality. Never make up numbers. If unsure, ask for clarification. Maximum 10 steps. Always show your reasoning.

Exercise 3: Design a Multi-Step Task

Think of a complex task that requires multiple steps and tool use (e.g., 'Plan a trip to Japan'). Break it into subtasks. Identify dependencies. Sketch which tools you'd need.

Python โ€” Starter Code
Task: "Plan a week-long trip to Japan for a family with teenagers" Subtasks: 1. Research visa requirements (tool: search_visa_info) 2. Find flights (tool: flight_search) 3. Book accommodations (tool: hotel_search) 4. Plan activities (tool: attraction_search) 5. Check weather & pack advice (tool: weather_forecast) 6. Create budget (tool: currency_converter, cost_calculator) Dependencies: - Know dates before searching flights/hotels - Know rough itinerary before planning activities - All of the above before calculating budget Tools needed: search, flight API, hotel API, currency converter

Exercise 4: Error Handling Scenarios

For an agent that answers customer service questions, write error messages for these scenarios: (1) Database is down, (2) Tool times out, (3) Answer is unclear. Make the messages actionable, not just 'Error'.

Python โ€” Starter Code
Scenarios and good error messages: 1. Database down BAD: "Error: Database connection failed" GOOD: "I'm having trouble accessing customer data right now. This should be resolved in a few minutes. I can provide general help while we wait." 2. Tool timeout BAD: "Timeout" GOOD: "The search is taking longer than expected. Would you like me to: (A) Wait a bit longer, (B) Provide cached results from yesterday, (C) Search a smaller date range?" 3. Answer unclear BAD: "I don't know" GOOD: "I found some information, but it's not completely clear. Let me ask: Are you asking about return policy for physical products or digital services?"

Summary: Key Takeaways

The Big Picture

  • Agents are systems, not models: Built by combining LLMs, tools, memory, and reasoning frameworks.
  • The agent loop is fundamental: Observe โ†’ Think โ†’ Act โ†’ Observe. Repeat until goal achieved.
  • Grounding is critical: Agents without tools are just text generators. Tools connect agents to reality.
  • Memory matters: Episodic, semantic, procedural, and state memory together enable learning over time.
  • Design trumps complexity: A well-designed simple agent beats a complex one that doesn't work.

When to Use Agents

  • Tasks requiring multiple steps and external tools
  • Problems needing reasoning and adaptation
  • Systems that must interact with dynamic environments
  • Applications where transparency (showing reasoning) is important
  • Domains where errors need graceful recovery

When NOT to Use Agents

  • Simple classification or generation tasks (use LLM directly)
  • When latency is critical and you can't afford LLM loops
  • Fully deterministic workflows (use traditional software)
  • When you need guaranteed correctness (agents aren't 100% reliable)

Getting Started

  1. Pick a simple task (e.g., answer a frequently asked question)
  2. Build a basic agent with 1-2 tools using LangChain
  3. Test thoroughly with diverse inputs
  4. Add complexity gradually (more tools, memory, reasoning)
  5. Collect user feedback and iterate

Key Resources

Further Learning

Online Courses

  • DeepLearning.AI: "Agentic Design Patterns with Claude" - Excellent practical course
  • Anthropic: "Building with Claude" - Free resources on agent building

Books

Community Resources

  • GitHub Discussions - LangChain and CrewAI communities are active and helpful
  • Discord servers - Many AI communities discussing agents and frameworks
  • Academic conferences - ICLR, NeurIPS, ACL publish cutting-edge agent research

Final Thought: The field of AI agents is rapidly evolving. What you learned here is foundational, but the best practices will change. Stay current by reading research papers, following key researchers on social media, and building real projects. Theory is important; practice is essential.