# How to Build Small Focused Agents vs Monolithic Agent Systems

> Learn to build small focused agents versus monolithic agent systems. Discover how composable agents improve manageability and context with the 12-factor approach.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: architecture
- Published: 2026-05-19

---

**The 12-Factor Agents guide recommends treating agents as composable building blocks rather than monolithic all-in-one systems, with each small agent handling 3-20 steps to maintain manageable context windows and clear responsibilities.**

The `humanlayer/12-factor-agents` repository defines architectural patterns for production-grade AI systems. When evaluating how to build small focused agents vs monolith agent systems, the methodology advocates for modular designs that prevent context window exhaustion and enable reliable debugging.

## Why Monolithic Agents Break Down

Monolithic agents attempt to handle every possible task within a single LLM loop. According to [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md), this approach quickly exhausts the context window, makes debugging difficult, and reduces overall reliability. When an LLM must track dozens of steps simultaneously, error rates increase and root-cause analysis becomes expensive.

## Benefits of Small Focused Agents

### Manageable Context Windows

Small agents perform **3-20 steps** per execution, keeping prompts concise and the LLM's attention focused on a single domain. As documented in [`content/factor-10-small-focused-agents.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md) and illustrated in `img/1a0-small-focused-agents.png`, this constraint prevents the attention fragmentation that plagues long-prompt architectures.

### Clear Responsibilities

Each dedicated agent serves a well-defined purpose, such as "search-and-summarize" or "extract email addresses." This single-responsibility principle, emphasized in the 12-Factor Agents guide, eliminates the mixed concerns that make monolithic systems difficult to reason about.

### Isolated Failures and Debugging

When agents remain small, failures isolate to specific components. The [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md) file explains how compact error contexts enable simple retries and targeted tests, whereas monolithic errors cascade through entire workflows.

### Composable DAG Pipelines

Small agents chain into **directed acyclic graphs (DAGs)**, passing structured output between nodes. This composition pattern, illustrated in `img/025-agent-dag.png`, allows each agent to own its specific transformation while the orchestration layer manages control flow.

## Architectural Pattern for Implementation

To build small focused agents effectively, follow the four-step pattern from the 12-Factor Agents guide:

1. **Define a narrow goal** for each agent (e.g., "extract email addresses", "translate text").
2. **Implement a deterministic wrapper** that calls the LLM with a concise prompt and expects structured JSON output.
3. **Compose agents into a DAG** where output from one agent becomes input for the next.
4. **Iterate** until the final result is produced or a termination signal returns.

## Code Comparison: Small Agents vs Monolith

The repository provides TypeScript examples demonstrating both approaches. In the small agent pattern, each function invokes the LLM with a task-specific prompt:

```typescript
// Small focused agent: Extract emails only
async function extractEmails(input: string): Promise<{emails: string[]}> {
  const prompt = `Extract all email addresses from the following text and return JSON:
{
  "emails": [...]
}
Text:
${input}`;
  const response = await llm.call(prompt);
  return JSON.parse(response);
}

// Composed in a DAG pipeline
async function pipeline(text: string) {
  const {emails} = await extractEmails(text);
  const summary = await summarize(emails.join(', '));
  return summary;
}

```

Contrast this with the monolithic approach that burdens a single LLM call with multiple concerns:

```typescript
// Monolithic agent: Handles everything in one call
async function monolithAgent(payload: string) {
  const prompt = `You are an AI assistant. Perform the following steps in order:
1. Extract email addresses from the payload.
2. Summarize the extracted emails.
Return a JSON object:
{
  "emails": [...],
  "summary": "..."
}`;
  const response = await llm.call(prompt);
  return JSON.parse(response);
}

```

The small agent version invokes the LLM twice with short, focused prompts, while the monolith requires the model to maintain context across extraction and summarization simultaneously, degrading performance.

## Why Small Agents Remain the Future-Proof Baseline

Even as LLM capabilities expand—supporting larger context windows and longer reasoning chains—the **principle of modularity** retains value. As noted in [`content/factor-10-small-focused-agents.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md), you can always combine small agents into larger sub-DAGs when models improve, but starting with a monolith locks you into an untestable, unobservable architecture. The guide treats the small-agent approach as the **future-proof baseline** rather than a temporary constraint.

## Summary

- **Monolithic agents** exhaust context windows and complicate debugging by handling multiple concerns in single LLM loops.
- **Small focused agents** limit execution to 3-20 steps, maintaining clear responsibilities and manageable prompts according to the 12-Factor Agents methodology.
- **DAG composition** enables reliable pipelines where agents pass structured outputs through deterministic wrappers.
- **Source files** including [`content/factor-10-small-focused-agents.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md) and [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) provide the complete rationale for modular architectures.
- Even with advancing LLM capabilities, small agents remain the preferred baseline for testability and observability.

## Frequently Asked Questions

### What is the optimal number of steps for a small focused agent?

The 12-Factor Agents guide recommends limiting each agent to **3-20 steps**. This range keeps the context window focused while providing sufficient complexity for meaningful tasks. Reference [`content/factor-10-small-focused-agents.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md) for the complete rationale behind these constraints.

### How do small agents handle complex workflows that require multiple operations?

Small agents compose into **directed acyclic graphs (DAGs)** where the output of one agent becomes the input of the next. As shown in `img/025-agent-dag.png`, this pattern allows complex workflows to emerge from simple, testable components rather than internal control flow.

### When might a monolithic agent architecture be appropriate?

A monolithic approach might appear viable if LLMs could reliably handle 100+ steps without losing context. However, the `humanlayer/12-factor-agents` repository recommends maintaining modular boundaries even then, as the **principle of modularity** provides essential testability and observability benefits that monoliths sacrifice.

### How does the small agent pattern affect error handling?

Small agents produce **compact errors** that isolate failures to specific components. According to [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md), this isolation enables simple retry logic and targeted debugging, whereas monolithic architectures allow errors to cascade through entire workflows making root-cause analysis expensive.