How to Build a Claude Code-like AI Coding Agent: A Step-by-Step Technical Guide
Build a Claude Code-like AI coding agent by implementing a core while-loop that invokes LLM tools, then layering capabilities like file editing, task planning, sub-agents, and Git worktree isolation without modifying the original loop.
The shareAI-lab/learn-claude-code repository provides a progressive tutorial on how to build a Claude Code-like AI coding agent from scratch. This guide walks through the architecture, starting from a minimal agent loop and incrementally adding production-grade features like context compression, background tasks, and multi-agent coordination.
The Core Agent Loop (Session s01)
Every Claude Code-like agent starts with a single while-loop that maintains conversation state with an LLM and handles tool invocations. In agents/s01_agent_loop.py【lines 66‑78】, the loop streams responses from the Anthropic API:
while True:
response = client.messages.create(
model=MODEL, system=SYSTEM, messages=messages,
tools=TOOLS, max_tokens=8000,
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
return
# …execute tools, append tool results, loop again…
The system prompt at line 40 establishes the agent's identity: a coding agent operating at a specific working directory that must use tools to act rather than explain. This constraint forces the LLM to invoke functions instead of generating explanatory text.
Extensible Tool Dispatch (Session s02)
Rather than hardcoding tool logic inside the loop, the architecture uses a dispatch map (TOOL_HANDLERS) that registers capabilities dynamically. In agents/s02_tool_use.py【lines 93‑100】, handlers are defined as lambdas wrapping safe file and shell operations:
TOOL_HANDLERS = {
"bash": lambda **kw: run_bash(kw["command"]),
"read_file": lambda **kw: run_read(kw["path"], kw.get("limit")),
"write_file": lambda **kw: run_write(kw["path"], kw["content"]),
"edit_file": lambda **kw: run_edit(kw["path"], kw["old_text"], kw["new_text"]),
}
Adding a new tool requires only two steps: insert a key-value pair into TOOL_HANDLERS and append the corresponding JSON schema to the TOOLS list. The core loop remains unchanged, demonstrating open-closed principle architecture.
Planning and Progress Tracking (Session s03)
Production agents require structured planning. The TodoManager class in agents/s03_todo_write.py【lines 50‑57】maintains a checklist that the LLM updates via a todo tool:
class TodoManager:
def update(self, items: list) -> str: … # validates, renders checklist
The full agent implementation in agents/s_full.py【lines 998‑1002】tracks rounds_without_todo and injects a reminder into the system prompt when the agent has idled for three turns without updating its plan. This prevents the agent from losing track of multi-step objectives.
Sub-Agents for Isolated Execution (Session s04)
Complex tasks benefit from context isolation via sub-agents. The run_subagent function in agents/s04_subagent.py【lines 59‑94】spawns a temporary agent with its own message history and tool set:
def run_subagent(prompt: str, agent_type: str = "Explore") -> str:
# builds a temporary tool set, runs up to 30 turns, returns a summary
Sub-agents execute up to 30 turns before returning a concise summary to the parent agent, preventing context window exhaustion. The parent agent invokes sub-agents through the task tool registered in agents/s_full.py【lines 83‑84】.
Dynamic Skill Loading (Session s05)
Rather than bloating the system prompt with every possible capability, the agent implements on-demand knowledge injection. The SkillLoader class in agents/s05_skill_loading.py【lines 1‑22】reads Markdown skill files from the skills/ directory:
class SkillLoader:
def load(self, name: str) -> str: … # returns <skill …> XML wrapper
Skills are fetched via the load_skill tool only when the LLM recognizes it needs specific expertise (e.g., "React best practices" or "Docker optimization"), keeping the active context minimal.
Context Compression Strategies (Session s06)
Long-running sessions require token management to prevent context window overflow. The repository implements two compression strategies in agents/s06_context_compact.py:
Micro-compact replaces large tool results with [cleared] after accumulating three results (agents/s_full.py【lines 30‑41】). This preserves recent context while discarding older execution noise.
Auto-compact triggers when token count exceeds TOKEN_THRESHOLD (100,000 tokens). The system summarizes the entire conversation history via a dedicated LLM call and stores it as a transcript (agents/s_full.py【lines 42‑58】), effectively resetting the context window while preserving semantic continuity.
Persistent Task Management (Session s07)
Production workflows require durable task tracking beyond ephemeral conversation state. The TaskManager class in agents/s07_task_system.py【lines 62‑85】implements a file-based task board using JSON storage in .tasks/:
class TaskManager:
def create(self, subject, description="") -> str: …
def list_all(self) -> str: …
Tasks support dependency tracking, ownership fields, and lifecycle states. This enables the agent to maintain project backlogs across restarts and coordinate multi-session workflows.
Background Task Execution (Session s08)
Synchronous tool calls block the agent loop, preventing progress during long-running operations. The BackgroundManager in agents/s08_background_tasks.py【lines 27‑44】implements threaded execution with status polling:
class BackgroundManager:
def run(self, command, timeout=120) -> str: …
def drain(self) -> list: …
Commands execute in separate threads with configurable timeouts. The main loop drains notification queues each turn, allowing the agent to issue long-running tests or builds without blocking user interaction.
Multi-Agent Team Coordination (Sessions s09-s11)
Scalable systems require distributed agent teams rather than monolithic instances. The repository implements three coordination layers:
MessageBus (agents/s09_agent_teams.py) provides per-agent mailboxes in inbox/ directories, enabling asynchronous communication between teammates.
TeammateManager (agents/s10_team_protocols.py) spawns persistent agents that monitor their inboxes, claim tasks, and negotiate via request/response protocols.
Autonomous Claiming (agents/s11_autonomous_agents.py) allows idle teammates to scan the task board and claim unassigned work. The logic in agents/s_full.py【lines 496‑511】handles idle state transitions, enabling self-organizing teams without central orchestration.
Worktree-Based Task Isolation (Session s12)
Concurrent development requires filesystem isolation to prevent cross-contamination between tasks. The WorktreeManager in agents/s12_worktree_task_isolation.py【lines 83‑150】leverages Git worktrees for lightweight, branch-based isolation:
class WorktreeManager:
def create(self, name, task_id=None, base_ref="HEAD") -> str: …
def run(self, name, command) -> str: …
def remove(self, name, force=False, complete_task=False) -> str: …
Worktrees are created under .worktrees/, bound to specific task IDs, and execute commands in isolated directories. The EventBus records lifecycle events (creation, execution, removal), providing audit trails for parallel development workflows.
The Full Integration (s_full.py)
The capstone implementation in agents/s_full.py【lines 0‑30】composes all mechanisms—compression, background tasks, inbox management, persistent tasks, sub-agents, teammates, and worktree tools—while preserving the original agent loop unchanged. The top-level REPL provides shortcuts like /tasks, /team, and /inbox for rapid inspection of system state.
Summary
- Start with the loop: Implement a minimal while-loop in
agents/s01_agent_loop.pythat streams LLM responses and handles tool_use stop reasons. - Decouple tools: Use a dispatch map (
TOOL_HANDLERS) inagents/s02_tool_use.pyto add capabilities without modifying core logic. - Add planning: Integrate
TodoManagerfromagents/s03_todo_write.pyto prevent the agent from losing track of multi-step objectives. - Isolate complexity: Spawn sub-agents via
run_subagentinagents/s04_subagent.pyfor exploratory tasks that might exhaust context windows. - Manage context: Implement micro-compact and auto-compact strategies from
agents/s06_context_compact.pyto handle long-running sessions. - Persist state: Use
TaskManagerinagents/s07_task_system.pyandWorktreeManagerinagents/s12_worktree_task_isolation.pyfor durable, isolated workflows. - Scale horizontally: Deploy
MessageBus,TeammateManager, and autonomous claiming from sessions s09-s11 to coordinate multi-agent teams.
Frequently Asked Questions
What is the minimum code needed to start building a Claude Code-like agent?
You need approximately 80 lines of Python to implement the core loop. The file agents/s01_agent_loop.py demonstrates this with a while True block that calls client.messages.create(), appends responses to a message history, and breaks only when stop_reason != "tool_use". This minimal implementation requires only an LLM client, a system prompt, and a basic tool schema.
How does the agent avoid exceeding LLM context window limits?
The repository implements two compression mechanisms in agents/s06_context_compact.py. Micro-compact replaces bulky tool results with [cleared] placeholders after accumulating three results, preserving recent context while discarding older execution noise. Auto-compact triggers when token count exceeds 100,000, summarizing the entire conversation history via a dedicated LLM call and storing it as a transcript, effectively resetting the context window while preserving semantic continuity.
Can multiple agents work on different tasks simultaneously without interfering?
Yes, through the WorktreeManager implemented in agents/s12_worktree_task_isolation.py. This system creates Git worktrees under .worktrees/ that provide filesystem isolation for each task. Combined with the TaskManager from agents/s07_task_system.py, which maintains persistent JSON task boards in .tasks/, agents can bind specific worktrees to task IDs, execute commands in isolated directories, and run background processes without cross-contamination.
How do you add new capabilities to the agent without breaking existing functionality?
The architecture uses a dispatch map pattern defined in agents/s02_tool_use.py. New capabilities are added by extending the TOOL_HANDLERS dictionary with a lambda or function reference, then appending the corresponding JSON schema to the TOOLS list. Because the core loop in agents/s01_agent_loop.py only checks response.stop_reason and dispatches through TOOL_HANDLERS, adding tools requires zero changes to the agent's control flow, following the open-closed principle.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →