Designing an Agent Workbench for AI Development: A Three-File Minimal Framework

An agent workbench provides durable state persistence and task management through three core files—AGENTS.md, agent_state.json, and task_board.json—enabling AI agents to maintain context across sessions without relying on volatile chat history.

Designing an agent workbench for AI development requires solving the fundamental problem of state persistence across sessions. The rohitg00/ai-engineering-from-scratch repository implements a lightweight, file-based framework that eliminates the single point of failure represented by chat history. This approach gives AI agents a durable place to store state, a clear task queue, and a concise routing file, allowing models to focus on problem-solving rather than re-learning coordination patterns.

Core Architecture of the Agent Workbench

The minimal viable workbench consists of exactly three files that abstract repetitive scaffolding. According to the source code in phases/14-agent-engineering/32-minimal-agent-workbench/, these files work together to create a durable system of record.

The Router File (AGENTS.md)

workdir/AGENTS.md serves as a short router that points the agent to the state file, task board, and deeper rule documentation. It replaces massive, monolithic manuals and keeps the model’s context budget low by only loading deeper docs when needed.

The State File (agent_state.json)

workdir/agent_state.json acts as the system of record, implemented via the AgentState dataclass. It holds the current active task ID, files touched during the turn, assumptions, blockers, and the next action the agent should take. This persists across sessions, eliminating reliance on volatile chat history.

The Task Board (task_board.json)

workdir/task_board.json maintains a JSON queue of tasks with fields id, goal, owner, acceptance, and status. The agent pulls the next todo task when the state is empty, providing a persistent view of work progress that survives process restarts.

How the Agent Workbench Works

The reference implementation in phases/14-agent-engineering/32-minimal-agent-workbench/code/main.py demonstrates a deterministic five-step lifecycle:

  1. Bootstrap – The script creates the three files if they don’t exist using write_initial(state_path, board_path, agents_path).
  2. Load – At the start of each turn, the agent reads the current context via load_state(state_path) and load_board(board_path).
  3. Decision – If no active task exists, the agent picks the next todo task, marks it in_progress, and records the task ID in the state. If an active task exists, the agent performs deterministic actions: touching source files, adding tests, running verification, then marking the task done.
  4. Persist – Updated state and board are written back to disk using save_state(state_path, state) and save_board(board_path, board).
  5. Repeat – Re-running the script demonstrates continuity across sessions as the second turn automatically picks up where the first left off.

Why Three Files?

The three-file constraint in rohitg00/ai-engineering-from-scratch provides specific architectural advantages:

  • Simplicity – Anything more complex builds on this foundation, making it easy to reason about, test, and version-control.
  • Portability – The same three files map directly onto conventions of major AI coding assistants (Claude Code, Cursor, Codex) with only naming tweaks.
  • Reliability – By storing the workbench on disk, you remove the single point of failure that chat history represents. The workbench can be audited, linted, and migrated independently of the model.

Extending the Workbench

Later lessons in Phase 14 demonstrate how to layer additional surfaces on top of the base three files. According to phases/14-agent-engineering/34-repo-memory-and-state/docs/en.md, extensions include:

  • Scope contracts – Defining which files the agent may modify
  • Verification gates – Commands that confirm a task’s definition of done
  • Reviewer agents – Automatic checks before a task is marked done
  • Handoff packets – Machine-readable summaries for the next session

All extensions read and write the same three core files, preserving backward compatibility while adding robustness.

Implementation Example

The following example from code/main.py demonstrates the workbench scaffold and a single agent turn:


# Run the workbench scaffold and one agent turn

python3 code/main.py

The script performs these specific operations:


# 1. Ensure the workdir exists

ROOT.mkdir(exist_ok=True)

# 2. Write AGENTS.md, agent_state.json, and task_board.json if missing

write_initial(state_path, board_path, agents_path)

# 3. Load the current state and task board

state = load_state(state_path)
board = load_board(board_path)

# 4. Print the pre-turn status

print("before turn:", state.active_task_id, state.next_action)

# 5. Run a deterministic turn

state, board = run_one_turn(state, board)

# 6. Persist the updated files

save_state(state_path, state)
save_board(board_path, board)

# 7. Print the post-turn status

print("after turn:", state.active_task_id, state.next_action)

Running the script a second time will show the agent picking up the next task automatically, proving the durability of the workbench architecture.

Summary

  • Designing an agent workbench for AI development requires only three files: AGENTS.md for routing, agent_state.json for durable context, and task_board.json for task queue management.
  • The AgentState dataclass in agent_state.json tracks active task IDs, touched files, assumptions, and next actions across sessions.
  • Functions load_state(), save_state(), load_board(), and save_board() provide the persistence layer that eliminates dependency on chat history.
  • The three-file approach balances simplicity with extensibility, allowing later phases to add scope contracts and verification gates without breaking the core schema.
  • This architecture is implemented in phases/14-agent-engineering/32-minimal-agent-workbench/code/main.py in the rohitg00/ai-engineering-from-scratch repository.

Frequently Asked Questions

What makes the three-file workbench architecture reliable for AI development?

The three-file architecture eliminates the single point of failure represented by volatile chat history. By storing state in agent_state.json and the task queue in task_board.json on disk, the workbench survives process restarts, CI runs, and human hand-offs. Files can be version-controlled, audited, and migrated independently of the model, providing deterministic continuity across sessions.

How does the AGENTS.md file reduce context usage?

AGENTS.md acts as a lightweight router rather than a comprehensive manual. It points the agent to the state file, task board, and deeper rule documentation only when needed. This concision keeps the model’s context budget focused on the actual problem rather than coordination mechanics, which is critical when designing an agent workbench for AI development with limited token windows.

Can the minimal workbench handle complex multi-step workflows?

Yes. The task_board.json queue supports complex workflows through its id, goal, owner, acceptance, and status fields. The run_one_turn() function in main.py demonstrates how agents perform deterministic actions—touching files, adding tests, running verification—before marking tasks done. Later lessons in Phase 14 show how to add verification gates and reviewer agents without changing the core three-file structure.

Where can I find the reference implementation for this agent workbench?

The complete reference implementation is located in phases/14-agent-engineering/32-minimal-agent-workbench/code/main.py within the rohitg00/ai-engineering-from-scratch repository. The accompanying documentation in phases/14-agent-engineering/32-minimal-agent-workbench/docs/en.md explains the three-file concept in depth, while phases/14-agent-engineering/34-repo-memory-and-state/docs/en.md covers extending the workbench into a full repository memory system.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →