# Can Open SWE Refactor Existing Codebases? A Deep Dive into Automated Code Transformation

> Yes Open SWE refactors existing codebases with a strict prompt precise tools and automated PRs for safe code transformation from discovery to review.

- Repository: [LangChain/open-swe](https://github.com/langchain-ai/open-swe)
- Tags: deep-dive
- Published: 2026-03-19

---

**Yes, Open SWE can refactor existing codebases by combining a strict system prompt, precision file-editing tools, and automated PR workflows to safely transform code from discovery to review.**

Open SWE, developed by LangChain AI, is an open-source agentic framework explicitly built to automate code-change tasks within existing repositories. Whether you need to migrate legacy patterns, enforce new style guides, or restructure entire packages, Open SWE provides the architectural foundation to **refactor existing codebases** reliably and reviewably.

## Core Architecture for Refactoring Existing Codebases

The ability to safely transform production code stems from three integrated components: a disciplined system prompt, a precision toolset, and mandatory review workflows.

### System Prompt: The Refactoring Philosophy

In [`agent/prompt.py`](https://github.com/langchain-ai/open-swe/blob/main/agent/prompt.py), the system prompt establishes strict ground rules that govern how the agent approaches modifications. Lines 13-21 instruct the agent to *"Read files before modifying them"* and to *"Fix root causes, not symptoms"* while preserving the existing code style. This constraint prevents superficial changes and ensures refactoring addresses architectural debt rather than masking it.

### Deep Agents Toolset: Precision Editing

Open SWE imports the **Deep Agents toolset** automatically, exposing functions that allow surgical code manipulation:

- `read_file`: Inspects current implementation before changes.
- `edit_file`: Modifies specific line ranges without rewriting entire files.
- `write_file`: Creates new modules when restructuring requires new files.

These tools operate within sandboxed environments, ensuring the agent can experiment with transformations without risking production stability.

## The Four-Stage Refactoring Workflow

Open SWE executes refactoring through a deterministic pipeline that mirrors professional software engineering practices.

### Stage 1: Discovery and Analysis

The agent begins by exploring the codebase to identify refactoring targets. Using `execute` to run `grep`, `find`, or repository-specific scripts, the agent maps dependencies and locates legacy patterns. This discovery phase ensures the agent understands the full scope of changes before modifying any files.

### Stage 2: Applying Code Transformations

With targets identified, the agent invokes `edit_file` or `write_file` to apply transformations. For example, converting legacy classes to `@dataclass` decorators or migrating callback-based code to async/await patterns. The system prompt constraints ensure these changes maintain existing style conventions and do not introduce unrelated modifications.

### Stage 3: Validation with Linters and Tests

Before finalizing, the agent validates changes within the sandbox environment. The `execute` tool runs formatters like **Black** or **isort**, followed by the repository's test suite. This validation step catches syntax errors or regressions introduced during the automated refactor, ensuring only working code reaches review.

### Stage 4: Committing and Opening Pull Requests

Once validation passes, the agent calls `commit_and_open_pr` from [`agent/tools/commit_and_open_pr.py`](https://github.com/langchain-ai/open-swe/blob/main/agent/tools/commit_and_open_pr.py). This tool:

1. Stages all modified files.
2. Creates a commit with a descriptive message.
3. Pushes to a dedicated branch named `open-swe/<thread_id>`.
4. Opens a GitHub PR using the repository's required template conventions (lines 27-36 and 42-55).

Additionally, the `open_pr` middleware in [`agent/middleware/open_pr.py`](https://github.com/langchain-ai/open-swe/blob/main/agent/middleware/open_pr.py) acts as a safety net, automatically triggering the PR creation if the LLM neglects to call the tool explicitly (lines 1-20).

## Sandbox Environments for Safe Refactoring

Open SWE supports multiple **sandbox backends** that provide isolated Linux environments for refactoring operations. According to the README (lines 52-57), supported platforms include:

- **Modal**
- **Daytona**
- **Runloop**
- **LangSmith**

These sandboxes allow the agent to run destructive operations—such as mass find-and-replace scripts or dependency upgrades—without affecting the production repository. The isolated environment ensures that refactoring experiments remain contained until explicitly committed and pushed via the PR workflow.

## Practical Examples: Refactoring Code with Open SWE

### Example 1: Triggering Refactor via Slack Integration

You can initiate a refactoring task through natural language commands:

```text
@openswe Please refactor the `utils` package to use type‑annotated dataclasses.

```

The agent executes the following sequence:

1. `execute` to search for class definitions in the `utils` directory.
2. `edit_file` to replace legacy classes with `@dataclass` decorators and type annotations.
3. `execute` to run `black . && isort .` for formatting.
4. `commit_and_open_pr` to create a reviewable pull request.

### Example 2: Programmatic Refactoring with the Python API

For automated pipelines, invoke the `commit_and_open_pr` tool directly:

```python
from openswe.agent.tools import commit_and_open_pr

result = commit_and_open_pr(
    title="refactor: convert utils to dataclasses [closes PROJ-123]",
    body="""

## Description

Converted legacy utility classes to `@dataclass` for clearer type contracts.

## Test Plan

- [ ] Run the existing unit‑test suite (passes) 
""",
    commit_message="refactor utils to dataclasses"
)

if result["success"]:
    print("PR opened:", result["pr_url"])
else:
    print("Error:", result["error"])

```

This function handles all git operations, pushes to the `open-swe/<thread_id>` branch, and creates the PR using GitHub's API according to the template conventions defined in [`agent/tools/commit_and_open_pr.py`](https://github.com/langchain-ai/open-swe/blob/main/agent/tools/commit_and_open_pr.py) (lines 84-100).

### Example 3: Validating Refactors with Linters

Before submitting changes, validate code quality within the sandbox:

```python

# The agent executes formatting and linting

execute("black . && isort .")

```

After confirming clean output, the agent proceeds to `commit_and_open_pr`, ensuring that only properly formatted code reaches the pull request.

## Summary

Open SWE provides a complete, automated solution to **refactor existing codebases** through:

- **Disciplined system prompts** in [`agent/prompt.py`](https://github.com/langchain-ai/open-swe/blob/main/agent/prompt.py) that enforce root-cause fixes and style preservation.
- **Precision editing tools** (`edit_file`, `write_file`, `read_file`) that enable surgical code transformations.
- **Automated validation** via sandboxed execution of linters and test suites.
- **Guaranteed review workflows** through `commit_and_open_pr` and `open_pr` middleware that ensure every refactor becomes a pull request.

Whether modernizing legacy patterns or restructuring entire packages, Open SWE executes refactoring tasks with the rigor of a senior engineer while maintaining full auditability through GitHub's standard review process.

## Frequently Asked Questions

### Can Open SWE handle large-scale refactoring across multiple packages?

Yes. Open SWE uses the `execute` tool with `grep` and `find` to map dependencies across packages, then applies changes incrementally using `edit_file`. The sandbox environment supports running repository-wide scripts, and the `commit_and_open_pr` tool captures all modifications in a single, reviewable pull request regardless of scope.

### How does Open SWE ensure code style consistency during refactoring?

The system prompt in [`agent/prompt.py`](https://github.com/langchain-ai/open-swe/blob/main/agent/prompt.py) explicitly instructs the agent to preserve existing code style and fix root causes rather than symptoms. Additionally, the agent runs formatters like **Black** and **isort** via the `execute` tool before committing, ensuring all refactored code matches the repository's formatting standards.

### What happens if the agent forgets to open a pull request after refactoring?

The `open_pr` middleware in [`agent/middleware/open_pr.py`](https://github.com/langchain-ai/open-swe/blob/main/agent/middleware/open_pr.py) acts as a safety mechanism. If the LLM completes the refactoring workflow without explicitly calling `commit_and_open_pr`, this middleware automatically triggers the PR creation, ensuring that no code changes are lost and every refactor remains visible for human review.

### Is it safe to run Open SWE on production repositories?

Yes, because Open SWE operates entirely within **sandboxed environments** provided by backends like Modal, Daytona, Runloop, or LangSmith. The agent cannot directly modify production branches; instead, it creates isolated branches (prefixed `open-swe/<thread_id>`) and opens pull requests, requiring standard human approval before any changes reach production.