How Open-SWE Assists with Code Review: Automated PR Analysis and Response Generation
Open-SWE assists with code review by treating pull requests as structured conversations, automatically collecting and sanitizing GitHub comments, and equipping an LLM agent with read-only guidelines to generate contextual responses to reviewer feedback.
Open-SWE is an open-source AI agent framework designed to automate software engineering workflows. When assisting with code review, the system transforms GitHub pull requests into manageable dialogue contexts, enabling autonomous agents to understand reviewer feedback and respond appropriately without requiring write permissions until explicitly authorized.
Core Architecture for Code Review Assistance
System Prompt Guidelines
The foundation of Open-SWE’s code review capability lies in its system prompt engineering. In agent/prompt.py (lines 173-192), the CODE_REVIEW_GUIDELINES_SECTION constant provides the LLM with explicit instructions:
- Use read-only operations only
- Inspect diffs using
git diff - Focus on the most relevant files
- Avoid making changes until explicitly requested
These constraints ensure the agent acts as a careful reviewer rather than an autonomous editor, maintaining security boundaries while analyzing code.
Comment Collection Pipeline
Open-SWE implements a sophisticated comment aggregation system in agent/utils/github_comments.py. The fetch_pr_comments_since_last_tag function (lines 214-233) queries three distinct GitHub data sources:
- Issue-style comments on the PR
- Inline review comments attached to specific lines
- Review bodies submitted by reviewers
The function merges these chronologically and filters for comments appearing after the latest @open-swe tag, ensuring the agent only processes new feedback since its last interaction.
Processing Reviewer Feedback
Fetching Comments Since Last Tag
The temporal filtering mechanism prevents the agent from re-processing historical comments. When fetch_pr_comments_since_last_tag executes, it:
- Retrieves the repository configuration
- Identifies the PR number
- Locates the most recent
@open-swetag in the comment history - Returns only subsequent comments with full metadata
# agent/utils/github_comments.py
comments = await fetch_pr_comments_since_last_tag(
repo_config={"owner": "langchain-ai", "name": "open-swe"},
pr_number=42,
token=github_token,
)
Sanitizing Untrusted Input
Security considerations are paramount when processing external comments. Open-SWE implements two-stage sanitization in agent/utils/github_comments.py:
Stage 1: Dangerous Tag Removal
The sanitize_github_comment_body function (lines 60-71) strips special "dangerous" tags that could affect downstream processing or prompt injection attacks.
Stage 2: Untrusted Content Marking
The format_github_comment_body_for_prompt function (lines 74-84) wraps sanitized bodies in visible markers, ensuring the LLM recognizes external content versus internal system instructions.
# agent/utils/github_comments.py
clean_body = sanitize_github_comment_body(raw_body)
marked_body = format_github_comment_body_for_prompt(clean_body)
Building the PR Prompt Context
Once comments are collected and sanitized, build_pr_prompt in agent/prompt.py (lines 86-101) constructs the final LLM context:
- Formats each comment with author attribution
- Adds file paths and line numbers for inline comments
- Chronologically stitches comments into a coherent conversation thread
- Outputs a single user-message for LLM reasoning
# agent/prompt.py
prompt = build_pr_prompt(comments, pr_url="https://github.com/.../pull/42")
Interactive Review Responses
Reacting to Review Comments
Open-SWE provides immediate visual feedback to reviewers through the react_to_github_comment function in agent/utils/github_comments.py (lines 87-105). When processing a pull_request_review webhook payload, the agent:
- Issues a 👀 reaction via GitHub's REST API for regular comments
- Uses GraphQL mutations for review bodies requiring node IDs
- Signals active engagement while processing the review request
# agent/utils/github_comments.py
await react_to_github_comment(
repo_config={"owner": "langchain-ai", "name": "open-swe"},
comment_id=12345678,
event_type="pull_request_review",
token=github_token,
node_id="MDU6UmV2aWV3MTIzNDU2Nzg5"
)
Webhook Event Routing
The entry point for code review assistance resides in agent/webapp.py (lines 1073-1076). The webhook handler:
- Identifies supported GitHub events:
issue_comment,pull_request_review_comment, andpull_request_review - Routes PR-related events to the comment-fetching and prompt-building pipeline
- Triggers the LLM agent with the constructed context
Once the agent generates a response or code changes, agent/tools/commit_and_open_pr.py creates a new pull request with the suggested modifications, completing the review loop.
Summary
Open-SWE assists with code review through a structured pipeline that prioritizes security, context awareness, and automated responsiveness:
- System prompt guidelines in
agent/prompt.pyenforce read-only review policies and safe operations - Automated comment collection via
fetch_pr_comments_since_last_tagaggregates issue comments, inline reviews, and review bodies since the last agent interaction - Input sanitization through
sanitize_github_comment_bodyandformat_github_comment_body_for_promptprevents prompt injection and marks untrusted content - Context construction via
build_pr_promptformats reviewer feedback into chronological, metadata-rich LLM prompts - Interactive acknowledgment through
react_to_github_commentprovides immediate visual feedback to human reviewers - Webhook integration in
agent/webapp.pyroutes GitHub events to trigger the review pipeline automatically
Frequently Asked Questions
How does Open-SWE prevent unauthorized code changes during code review?
Open-SWE embeds explicit read-only constraints in the system prompt through the CODE_REVIEW_GUIDELINES_SECTION in agent/prompt.py (lines 173-192). These instructions direct the LLM to use only inspection commands like git diff and prohibit file modifications until explicitly authorized. Additionally, the sanitize_github_comment_body function strips potentially dangerous tags from external comments to prevent prompt injection attacks that could bypass these constraints.
What types of GitHub comments does Open-SWE collect for code review?
The fetch_pr_comments_since_last_tag function in agent/utils/github_comments.py (lines 214-233) queries three distinct GitHub data sources: issue-style comments posted on the PR conversation tab, inline review comments attached to specific lines of code, and review bodies submitted as part of formal GitHub reviews. The function merges these chronologically and filters for comments appearing after the latest @open-swe tag, ensuring the agent processes only new feedback.
How does Open-SWE handle potentially malicious content in review comments?
Open-SWE implements a two-stage sanitization pipeline in agent/utils/github_comments.py. First, sanitize_github_comment_body (lines 60-71) strips special "dangerous" tags that could affect downstream processing or enable prompt injection. Second, format_github_comment_body_for_prompt (lines 74-84) wraps the sanitized content in visible markers that clearly distinguish untrusted external input from internal system instructions, ensuring the LLM maintains appropriate context boundaries.
Can Open-SWE acknowledge receipt of review comments automatically?
Yes, through the react_to_github_comment function in agent/utils/github_comments.py (lines 87-105), Open-SWE can issue immediate visual feedback when processing review events. When the webhook payload indicates a pull_request_review event, the function adds a 👀 reaction to the comment via GitHub's REST API (for regular comments) or GraphQL mutations (for review bodies), signaling to human reviewers that the agent has received and is processing their feedback.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →