Agent Architecture and Design
The Components of an AI Agent
An AI agent is not a single thing -- it is a system made of multiple components working together. Understanding how these components fit together is essential before you can design agents that work reliably in production.
The core components of any AI agent are: the LLM (the reasoning engine), tools (the action capabilities), memory (context and history), and the orchestration layer (the code that manages the loop between them).
Component 1: The LLM (Reasoning Engine)
The LLM is the brain of the agent. It reads the current state (goal, conversation history, tool results), reasons about what to do next, and decides which tool to call and with what inputs.
The quality of the LLM directly affects agent performance:
- More capable models (GPT-4o, Claude 3.5 Sonnet) make better decisions and handle complex multi-step reasoning
- Faster, cheaper models (GPT-4o-mini) work well for simpler tasks but may struggle with complex reasoning chains
- For production agents: start with the most capable model, optimise down only after validating quality
Prompting the LLM in an agent context
The system prompt for an agent is more complex than a simple task prompt. It must:
- Define the agent's role and goal
- List the available tools with descriptions of when to use each
- Define the output format for tool calls
- Set constraints (what the agent should not do)
- Define when to stop and produce a final answer
Component 2: Tools
Tools are functions the LLM can call to interact with the world. Each tool has:
- A name: What the LLM calls it
- A description: When and why to use it (this is what the LLM reads to decide which tool to use)
- Input parameters: What the tool expects
- Output: What the tool returns
Common agent tools:
// Example tool definitions for a research agent
const tools = [
{
name: "web_search",
description: "Search the internet for current information. Use when you need facts, prices, news, or data not in your training.",
parameters: {
query: { type: "string", description: "The search query" }
}
},
{
name: "read_webpage",
description: "Read the full content of a webpage. Use after a search to get detailed information from a specific URL.",
parameters: {
url: { type: "string", description: "The URL to read" }
}
},
{
name: "write_file",
description: "Write content to a file. Use when you need to save your findings or create a document.",
parameters: {
filename: { type: "string" },
content: { type: "string" }
}
},
{
name: "finish",
description: "Use this when you have completed the task. Provide your final answer.",
parameters: {
answer: { type: "string", description: "The complete final answer or result" }
}
}
];
Critical design rule: Tool descriptions are the most important part of tool design. The LLM decides which tool to call based entirely on reading the description. Write clear, specific descriptions that explain not just what the tool does but when to use it.
Component 3: Memory
Memory determines what the agent knows about its history and current state. There are three types:
Working Memory (In-Context)
The conversation history -- the accumulated Thought/Action/Observation steps -- that fits within the LLM's context window. This is the simplest form of memory and is always present.
Limitation: Context windows have finite size. For very long tasks, earlier context gets dropped.
External Memory (Database)
Storing information outside the context window in a database or file system:
- Summary of completed steps
- Retrieved documents or data
- Results from previous agent runs
The agent can query external memory using a tool, bringing relevant information back into context when needed.
Semantic Memory (Vector Database)
Storing information as embeddings in a vector database (covered in detail in Module 6). This enables the agent to search its memory semantically: "Find what I know about X" rather than looking up by exact key.
Component 4: The Orchestration Layer
The orchestration layer is the code that manages the agent loop. It:
- Sends the initial goal and context to the LLM
- Receives the LLM's response (a tool call or a final answer)
- If it is a tool call: executes the tool and appends the result to the conversation
- If it is a final answer: returns it and stops
- Enforces limits (maximum iterations, timeout)
// Simplified agent loop in Node.js
async function runAgent(goal, tools, maxIterations = 10) {
const messages = [
{ role: 'system', content: buildSystemPrompt(tools) },
{ role: 'user', content: goal }
];
for (let i = 0; i < maxIterations; i++) {
const response = await callLLM(messages, tools);
if (response.finish_reason === 'stop') {
// Model produced a final answer
return response.content;
}
if (response.tool_calls) {
// Execute the requested tool call
for (const call of response.tool_calls) {
const result = await executeTool(call.name, call.arguments);
messages.push({ role: 'tool', content: JSON.stringify(result), tool_call_id: call.id });
}
}
}
return 'Max iterations reached. Task may be incomplete.';
}
Designing for Reliability
Set iteration limits
Always set a maximum number of iterations (typically 10 to 20). Without a limit, a stuck agent will loop forever at significant cost.
Design tool error handling
Every tool should return structured responses indicating success or failure:
// Good tool response structure
{ success: true, data: { ... } }
{ success: false, error: "Rate limit exceeded. Try again in 60 seconds." }
When a tool fails, the agent can decide whether to retry, use a different tool, or report the failure.
Validate tool inputs
Before executing a tool, validate that the LLM provided all required inputs in the correct format. Silently failing on bad inputs is worse than returning a clear error.
Log every step
Log every Thought/Action/Observation cycle with timestamps, tool names, inputs, and outputs. This is essential for debugging agent behaviour and improving prompts.
Single-Agent vs. Multi-Agent Architecture
Single agent: One agent handles the entire task with a set of tools. Simple to build and debug. Works well for tasks with clear, sequential steps.
Multi-agent: Multiple specialised agents collaborate. An orchestrator agent breaks the task into sub-tasks and delegates to specialist agents.
Multi-agent is appropriate when:
- The task is too complex or too long for a single context window
- Different sub-tasks require different specialisations
- Parallel execution of sub-tasks would save significant time
Key Takeaways
- An AI agent has four core components: the LLM (reasoning engine), tools (action capabilities), memory (context and history), and the orchestration layer (the loop manager).
- Tool descriptions are critical -- the LLM selects tools based entirely on reading their descriptions, so write them precisely.
- Three types of memory exist: working (in-context), external (database), and semantic (vector database).
- Always set iteration limits in your orchestration loop to prevent infinite loops and runaway costs.
- Multi-agent architectures are appropriate when tasks are too complex for a single agent's context window or when parallel execution is beneficial.
Try it yourself
Key Takeaways
- An AI agent has four core components: the LLM (reasoning), tools (action capabilities), memory (context and history), and an orchestration layer (loop manager).
- Tool descriptions are the most critical design element -- the LLM selects tools solely based on reading their descriptions.
- Always set maximum iteration limits to prevent infinite loops and runaway API costs.
- Three memory types exist: working (in-context history), external (database storage), and semantic (vector database for meaning-based retrieval).
- Multi-agent architectures suit complex tasks requiring specialisation or parallel execution; single agents are simpler for sequential tasks.
Quick Quiz
1.What determines which tool an AI agent selects to use?
2.What is the purpose of setting a maximum iteration limit in an agent loop?
3.What is the difference between working memory and external memory in an AI agent?
4.When is a multi-agent architecture more appropriate than a single-agent approach?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx