How to Build Small Focused Agents vs Monolithic Agent Systems

The 12-Factor Agents guide recommends treating agents as composable building blocks rather than monolithic all-in-one systems, with each small agent handling 3-20 steps to maintain manageable context windows and clear responsibilities.

The humanlayer/12-factor-agents repository defines architectural patterns for production-grade AI systems. When evaluating how to build small focused agents vs monolith agent systems, the methodology advocates for modular designs that prevent context window exhaustion and enable reliable debugging.

Why Monolithic Agents Break Down

Monolithic agents attempt to handle every possible task within a single LLM loop. According to content/factor-03-own-your-context-window.md, this approach quickly exhausts the context window, makes debugging difficult, and reduces overall reliability. When an LLM must track dozens of steps simultaneously, error rates increase and root-cause analysis becomes expensive.

Benefits of Small Focused Agents

Manageable Context Windows

Small agents perform 3-20 steps per execution, keeping prompts concise and the LLM's attention focused on a single domain. As documented in content/factor-10-small-focused-agents.md and illustrated in img/1a0-small-focused-agents.png, this constraint prevents the attention fragmentation that plagues long-prompt architectures.

Clear Responsibilities

Each dedicated agent serves a well-defined purpose, such as "search-and-summarize" or "extract email addresses." This single-responsibility principle, emphasized in the 12-Factor Agents guide, eliminates the mixed concerns that make monolithic systems difficult to reason about.

Isolated Failures and Debugging

When agents remain small, failures isolate to specific components. The content/factor-09-compact-errors.md file explains how compact error contexts enable simple retries and targeted tests, whereas monolithic errors cascade through entire workflows.

Composable DAG Pipelines

Small agents chain into directed acyclic graphs (DAGs), passing structured output between nodes. This composition pattern, illustrated in img/025-agent-dag.png, allows each agent to own its specific transformation while the orchestration layer manages control flow.

Architectural Pattern for Implementation

To build small focused agents effectively, follow the four-step pattern from the 12-Factor Agents guide:

  1. Define a narrow goal for each agent (e.g., "extract email addresses", "translate text").
  2. Implement a deterministic wrapper that calls the LLM with a concise prompt and expects structured JSON output.
  3. Compose agents into a DAG where output from one agent becomes input for the next.
  4. Iterate until the final result is produced or a termination signal returns.

Code Comparison: Small Agents vs Monolith

The repository provides TypeScript examples demonstrating both approaches. In the small agent pattern, each function invokes the LLM with a task-specific prompt:

// Small focused agent: Extract emails only
async function extractEmails(input: string): Promise<{emails: string[]}> {
  const prompt = `Extract all email addresses from the following text and return JSON:
{
  "emails": [...]
}
Text:
${input}`;
  const response = await llm.call(prompt);
  return JSON.parse(response);
}

// Composed in a DAG pipeline
async function pipeline(text: string) {
  const {emails} = await extractEmails(text);
  const summary = await summarize(emails.join(', '));
  return summary;
}

Contrast this with the monolithic approach that burdens a single LLM call with multiple concerns:

// Monolithic agent: Handles everything in one call
async function monolithAgent(payload: string) {
  const prompt = `You are an AI assistant. Perform the following steps in order:
1. Extract email addresses from the payload.
2. Summarize the extracted emails.
Return a JSON object:
{
  "emails": [...],
  "summary": "..."
}`;
  const response = await llm.call(prompt);
  return JSON.parse(response);
}

The small agent version invokes the LLM twice with short, focused prompts, while the monolith requires the model to maintain context across extraction and summarization simultaneously, degrading performance.

Why Small Agents Remain the Future-Proof Baseline

Even as LLM capabilities expand—supporting larger context windows and longer reasoning chains—the principle of modularity retains value. As noted in content/factor-10-small-focused-agents.md, you can always combine small agents into larger sub-DAGs when models improve, but starting with a monolith locks you into an untestable, unobservable architecture. The guide treats the small-agent approach as the future-proof baseline rather than a temporary constraint.

Summary

  • Monolithic agents exhaust context windows and complicate debugging by handling multiple concerns in single LLM loops.
  • Small focused agents limit execution to 3-20 steps, maintaining clear responsibilities and manageable prompts according to the 12-Factor Agents methodology.
  • DAG composition enables reliable pipelines where agents pass structured outputs through deterministic wrappers.
  • Source files including content/factor-10-small-focused-agents.md and content/factor-03-own-your-context-window.md provide the complete rationale for modular architectures.
  • Even with advancing LLM capabilities, small agents remain the preferred baseline for testability and observability.

Frequently Asked Questions

What is the optimal number of steps for a small focused agent?

The 12-Factor Agents guide recommends limiting each agent to 3-20 steps. This range keeps the context window focused while providing sufficient complexity for meaningful tasks. Reference content/factor-10-small-focused-agents.md for the complete rationale behind these constraints.

How do small agents handle complex workflows that require multiple operations?

Small agents compose into directed acyclic graphs (DAGs) where the output of one agent becomes the input of the next. As shown in img/025-agent-dag.png, this pattern allows complex workflows to emerge from simple, testable components rather than internal control flow.

When might a monolithic agent architecture be appropriate?

A monolithic approach might appear viable if LLMs could reliably handle 100+ steps without losing context. However, the humanlayer/12-factor-agents repository recommends maintaining modular boundaries even then, as the principle of modularity provides essential testability and observability benefits that monoliths sacrifice.

How does the small agent pattern affect error handling?

Small agents produce compact errors that isolate failures to specific components. According to content/factor-09-compact-errors.md, this isolation enables simple retry logic and targeted debugging, whereas monolithic architectures allow errors to cascade through entire workflows making root-cause analysis expensive.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →