Proposer-Reviewer Pattern for AI Agent Actions: A Two-Stage Validation Workflow
The Proposer-Reviewer pattern is a two-stage workflow that separates text-based plan generation from multimodal validation, allowing AI agents to iteratively refine outputs through structured feedback loops.
The Proposer-Reviewer pattern is a foundational architecture for building reliable AI agents that need both creativity and quality control. As implemented in the bojieli/ai-agent-book repository, this pattern isolates the generation of structured plans from their validation, enabling iterative refinement without overwhelming context windows.
How the Proposer-Reviewer Pattern Works
The architecture divides responsibilities between two specialized agents: one that proposes solutions using pure text reasoning, and another that reviews the resulting artifacts using vision-capable models.
The Proposer Agent: Text-Based Generation
The Proposer Agent acts as the creative engine, converting natural-language requests into structured execution plans. By design, this agent never processes images, keeping its context lightweight and token-efficient.
In chapter5/video-edit/agents.py, the ProposerAgent class (lines 32-80) implements this logic with two core methods:
parse_request(): Invokes the LLM through the sharedclient()wrapper to emit JSON describing target scenes and effectsrevise_bounds(): Adjusts temporal boundaries based on reviewer feedback while preserving sanity checks (clamping to video length)
The proposer outputs purely textual plans—for example, a JSON structure specifying which video segments to extract and what effects to apply—without ever loading pixel data into the context window.
The Reviewer Agent: Vision-Based Validation
The Reviewer Agent serves as the quality gate, inspecting the actual artifacts produced by the proposer. Unlike the proposer, this agent leverages vision-capable LLMs to evaluate the multimodal output.
Also defined in chapter5/video-edit/agents.py (lines 82-126), the ReviewerAgent class:
- Extracts key frames from the generated video
- Sends frames alongside a concise prompt to a vision-enabled model
- Returns a structured JSON verdict containing a boolean
passflag, a numericscore, and human-readablefeedback
This separation allows the reviewer to scrutinize visual quality, temporal accuracy, and semantic alignment between the generated content and the original request.
The Iterative Feedback Loop
The two agents communicate through a closed-loop orchestration demonstrated in chapter5/video-edit/demo.py (lines 190-227):
- The proposer generates an initial plan using
parse_request() - The system executes the plan to produce an artifact (e.g., a video clip)
- The reviewer inspects the result using its vision capabilities
- If the reviewer returns
"pass": false, its feedback is routed back to the proposer viarevise_bounds() - The proposer adjusts the plan (e.g., tweaking start/end timestamps) without regenerating the entire artifact from scratch
- The loop repeats until the reviewer signals success or a maximum iteration count is reached
Implementation in the ai-agent-book Repository
The following example from the repository demonstrates the complete workflow using the ProposerAgent and ReviewerAgent classes:
from agents import ProposerAgent, ReviewerAgent, TokenMeter
# Optional token metering for diagnostics
proposer_meter = TokenMeter()
reviewer_meter = TokenMeter()
proposer = ProposerAgent(meter=proposer_meter)
reviewer = ReviewerAgent(meter=reviewer_meter)
# 1️⃣ Propose a plan from a natural-language request
nl_req = "剪掉开场的广告,保留包含‘庆祝’字样的片段,并加慢动作效果。"
plan = proposer.parse_request(nl_req)
target_query = plan["target_query"] # e.g. "scene showing the word '庆祝'"
effects = plan["effects"] # e.g. [{"type":"slowmo","factor":2.0}]
# 2️⃣ Locate the target segment (using a helper VideoAnalyzerAgent)
# ... assume (start, end) have been determined ...
# 3️⃣ Apply the effects and produce a clipped video
# 4️⃣ Reviewer checks the result
review = reviewer.review("output/clip.mp4", target_query)
print("Reviewer verdict:", review["pass"])
print("Score:", review["score"])
print("Feedback:", review["feedback"])
# 5️⃣ If not satisfactory, ask the proposer to revise the bounds
if not review["pass"]:
start, end = proposer.revise_bounds(
start, end,
review["feedback"],
duration=video_duration
)
# Repeat clipping with the new bounds ...
This implementation showcases how the pattern maintains clean separation: the proposer handles chapter5/video-edit/agents.py lines 32-80 using text-only models, while the reviewer operates on lines 82-126 with vision capabilities.
Why Use the Proposer-Reviewer Pattern?
The bojieli/ai-agent-book repository employs this pattern for three architectural advantages:
- Isolation of token budgets: Keeping the proposer text-only prevents vision model costs and context limits from constraining the planning phase
- Modular reasoning: Generation and verification can use different model families optimized for their specific tasks (e.g., GPT-4 for planning, GPT-4V for visual inspection)
- Precise iterative refinement: The reviewer’s structured JSON feedback enables surgical adjustments through
revise_bounds()rather than expensive full regeneration
Summary
- The Proposer-Reviewer pattern separates AI agent workflows into generation and validation stages
- The Proposer Agent (
ProposerAgentinchapter5/video-edit/agents.py) handles text-only plan creation viaparse_request()and boundary adjustment viarevise_bounds() - The Reviewer Agent (
ReviewerAgentinchapter5/video-edit/agents.py) validates multimodal outputs using vision-capable models and returns structured feedback - The iterative loop demonstrated in
chapter5/video-edit/demo.pyenables refinement without regenerating entire artifacts - This architecture optimizes token usage by restricting vision processing to the validation stage only
Frequently Asked Questions
What is the main benefit of separating the Proposer and Reviewer agents?
Separating these roles allows each agent to use the most cost-effective model for its task. The proposer can run on cheaper text-only models, while the reviewer invokes vision-capable models only when necessary. According to the ai-agent-book source code, this isolation prevents the planning context from being polluted by large image payloads, significantly reducing token consumption.
How does the Proposer Agent handle feedback from the Reviewer Agent?
When the reviewer returns a failing verdict, the proposer invokes the revise_bounds() method (implemented in chapter5/video-edit/agents.py lines 32-80). This method takes the current temporal boundaries, the reviewer’s human-readable feedback, and the total video duration, then produces adjusted start and end timestamps clamped to valid ranges. This surgical adjustment avoids the cost of regenerating the entire plan from scratch.
Can the Proposer-Reviewer pattern be used for non-video AI tasks?
Yes. While the bojieli/ai-agent-book repository demonstrates this pattern for video editing, the architecture generalizes to any multimodal workflow. You can replace the proposer with a code-generation agent, slide-deck creator, or prompt engineer, and pair it with a reviewer that understands the target modality—whether that involves parsing JSON structures, analyzing audio waveforms, or validating API responses. The key is maintaining the separation between the text-based planning phase and the artifact validation phase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →