How Open-SWE Assists with Code Review: Automated PR Analysis and Response Generation

Open-SWE assists with code review by treating pull requests as structured conversations, automatically collecting and sanitizing GitHub comments, and equipping an LLM agent with read-only guidelines to generate contextual responses to reviewer feedback.

Open-SWE is an open-source AI agent framework designed to automate software engineering workflows. When assisting with code review, the system transforms GitHub pull requests into manageable dialogue contexts, enabling autonomous agents to understand reviewer feedback and respond appropriately without requiring write permissions until explicitly authorized.

Core Architecture for Code Review Assistance

System Prompt Guidelines

The foundation of Open-SWE’s code review capability lies in its system prompt engineering. In agent/prompt.py (lines 173-192), the CODE_REVIEW_GUIDELINES_SECTION constant provides the LLM with explicit instructions:

  • Use read-only operations only
  • Inspect diffs using git diff
  • Focus on the most relevant files
  • Avoid making changes until explicitly requested

These constraints ensure the agent acts as a careful reviewer rather than an autonomous editor, maintaining security boundaries while analyzing code.

Comment Collection Pipeline

Open-SWE implements a sophisticated comment aggregation system in agent/utils/github_comments.py. The fetch_pr_comments_since_last_tag function (lines 214-233) queries three distinct GitHub data sources:

  1. Issue-style comments on the PR
  2. Inline review comments attached to specific lines
  3. Review bodies submitted by reviewers

The function merges these chronologically and filters for comments appearing after the latest @open-swe tag, ensuring the agent only processes new feedback since its last interaction.

Processing Reviewer Feedback

Fetching Comments Since Last Tag

The temporal filtering mechanism prevents the agent from re-processing historical comments. When fetch_pr_comments_since_last_tag executes, it:

  • Retrieves the repository configuration
  • Identifies the PR number
  • Locates the most recent @open-swe tag in the comment history
  • Returns only subsequent comments with full metadata

# agent/utils/github_comments.py

comments = await fetch_pr_comments_since_last_tag(
    repo_config={"owner": "langchain-ai", "name": "open-swe"},
    pr_number=42,
    token=github_token,
)

Sanitizing Untrusted Input

Security considerations are paramount when processing external comments. Open-SWE implements two-stage sanitization in agent/utils/github_comments.py:

Stage 1: Dangerous Tag Removal The sanitize_github_comment_body function (lines 60-71) strips special "dangerous" tags that could affect downstream processing or prompt injection attacks.

Stage 2: Untrusted Content Marking The format_github_comment_body_for_prompt function (lines 74-84) wraps sanitized bodies in visible markers, ensuring the LLM recognizes external content versus internal system instructions.


# agent/utils/github_comments.py

clean_body = sanitize_github_comment_body(raw_body)
marked_body = format_github_comment_body_for_prompt(clean_body)

Building the PR Prompt Context

Once comments are collected and sanitized, build_pr_prompt in agent/prompt.py (lines 86-101) constructs the final LLM context:

  • Formats each comment with author attribution
  • Adds file paths and line numbers for inline comments
  • Chronologically stitches comments into a coherent conversation thread
  • Outputs a single user-message for LLM reasoning

# agent/prompt.py

prompt = build_pr_prompt(comments, pr_url="https://github.com/.../pull/42")

Interactive Review Responses

Reacting to Review Comments

Open-SWE provides immediate visual feedback to reviewers through the react_to_github_comment function in agent/utils/github_comments.py (lines 87-105). When processing a pull_request_review webhook payload, the agent:

  • Issues a 👀 reaction via GitHub's REST API for regular comments
  • Uses GraphQL mutations for review bodies requiring node IDs
  • Signals active engagement while processing the review request

# agent/utils/github_comments.py

await react_to_github_comment(
    repo_config={"owner": "langchain-ai", "name": "open-swe"},
    comment_id=12345678,
    event_type="pull_request_review",
    token=github_token,
    node_id="MDU6UmV2aWV3MTIzNDU2Nzg5"
)

Webhook Event Routing

The entry point for code review assistance resides in agent/webapp.py (lines 1073-1076). The webhook handler:

  • Identifies supported GitHub events: issue_comment, pull_request_review_comment, and pull_request_review
  • Routes PR-related events to the comment-fetching and prompt-building pipeline
  • Triggers the LLM agent with the constructed context

Once the agent generates a response or code changes, agent/tools/commit_and_open_pr.py creates a new pull request with the suggested modifications, completing the review loop.

Summary

Open-SWE assists with code review through a structured pipeline that prioritizes security, context awareness, and automated responsiveness:

  • System prompt guidelines in agent/prompt.py enforce read-only review policies and safe operations
  • Automated comment collection via fetch_pr_comments_since_last_tag aggregates issue comments, inline reviews, and review bodies since the last agent interaction
  • Input sanitization through sanitize_github_comment_body and format_github_comment_body_for_prompt prevents prompt injection and marks untrusted content
  • Context construction via build_pr_prompt formats reviewer feedback into chronological, metadata-rich LLM prompts
  • Interactive acknowledgment through react_to_github_comment provides immediate visual feedback to human reviewers
  • Webhook integration in agent/webapp.py routes GitHub events to trigger the review pipeline automatically

Frequently Asked Questions

How does Open-SWE prevent unauthorized code changes during code review?

Open-SWE embeds explicit read-only constraints in the system prompt through the CODE_REVIEW_GUIDELINES_SECTION in agent/prompt.py (lines 173-192). These instructions direct the LLM to use only inspection commands like git diff and prohibit file modifications until explicitly authorized. Additionally, the sanitize_github_comment_body function strips potentially dangerous tags from external comments to prevent prompt injection attacks that could bypass these constraints.

What types of GitHub comments does Open-SWE collect for code review?

The fetch_pr_comments_since_last_tag function in agent/utils/github_comments.py (lines 214-233) queries three distinct GitHub data sources: issue-style comments posted on the PR conversation tab, inline review comments attached to specific lines of code, and review bodies submitted as part of formal GitHub reviews. The function merges these chronologically and filters for comments appearing after the latest @open-swe tag, ensuring the agent processes only new feedback.

How does Open-SWE handle potentially malicious content in review comments?

Open-SWE implements a two-stage sanitization pipeline in agent/utils/github_comments.py. First, sanitize_github_comment_body (lines 60-71) strips special "dangerous" tags that could affect downstream processing or enable prompt injection. Second, format_github_comment_body_for_prompt (lines 74-84) wraps the sanitized content in visible markers that clearly distinguish untrusted external input from internal system instructions, ensuring the LLM maintains appropriate context boundaries.

Can Open-SWE acknowledge receipt of review comments automatically?

Yes, through the react_to_github_comment function in agent/utils/github_comments.py (lines 87-105), Open-SWE can issue immediate visual feedback when processing review events. When the webhook payload indicates a pull_request_review event, the function adds a 👀 reaction to the comment via GitHub's REST API (for regular comments) or GraphQL mutations (for review bodies), signaling to human reviewers that the agent has received and is processing their feedback.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →