# How DOCX Editing via XML Surgical Edits Works in DesktopCommanderMCP

> Discover how DesktopCommanderMCP performs DOCX editing via XML surgical edits by treating DOCX as ZIP archives for precise string replacements without full document parsing.

- Repository: [Eduard Ruzga/DesktopCommanderMCP](https://github.com/wonderwhy-er/DesktopCommanderMCP)
- Tags: internals
- Published: 2026-07-22

---

**DesktopCommanderMCP treats DOCX files as ZIP archives and performs precise string replacements on pretty-printed XML parts, enabling surgical text edits without requiring full document parsing libraries.**

DesktopCommanderMCP provides a lightweight approach to Word document manipulation by exposing **.docx** files as editable ZIP archives containing XML parts. The `DocxFileHandler` class in [`src/utils/files/docx.ts`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/src/utils/files/docx.ts) implements a surgical editing strategy that allows precise modifications to document content, headers, and footers through literal string replacement rather than complex object models.

## Understanding the DOCX File Structure

A **DOCX** file is essentially a ZIP archive containing multiple XML documents and supporting assets. The primary text content resides in [`word/document.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/word/document.xml), while headers and footers live in separate relationship-based files like [`header1.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/header1.xml) and [`footer2.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/footer2.xml). DesktopCommanderMCP leverages the **PizZip** library to extract these components, maintaining archive integrity while enabling direct XML manipulation. This architecture allows the system to preserve styles, relationships, and media files untouched while modifying only specific text fragments.

## Core Operations in DocxFileHandler

The `DocxFileHandler` class implements three fundamental operations defined in [`src/utils/files/docx.ts`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/src/utils/files/docx.ts): reading documents with dual-mode output, performing validated string replacements, and extracting document metadata.

### Reading DOCX Files (Outline vs. Raw XML)

The `read()` method (lines 18-33) operates in two distinct modes based on the `offset` parameter. 

When `offset` is 0 (default), the handler returns a human-readable **outline** generated by `extractOutline`, displaying paragraph text, table rows, headings, and image hints extracted from the `<w:body>` element. This mode facilitates high-level document comprehension for LLM agents.

When `offset` exceeds 0 (lines 24-31), the method enters **raw XML mode**, returning pretty-printed XML paginated line-by-line according to the `length` parameter. This enables clients to request specific structural sections of [`document.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/document.xml) for precise editing.

```typescript
// Retrieve human-readable outline (offset = 0 default)
const outline = await read_file({
  path: "/tmp/report.docx",
});
// Returns structured view of paragraphs, tables, and headings

// Retrieve raw XML starting at line 1
const rawXml = await read_file({
  path: "/tmp/report.docx",
  offset: 1,
  length: 200,
});
// Returns first 200 lines of pretty-printed document.xml

```

### Surgical Editing with editRange

The `editRange` method (lines 118-128) handles `edit_block` operations by performing literal string replacements on XML content. The workflow follows this sequence:

1. **Pretty-print** the target XML file using `prettyPrintXml` to ensure deterministic formatting
2. **Count occurrences** of `old_string` and validate against `expected_replacements`
3. **Execute** `String.replace` (or single-instance splice) to substitute `new_string`
4. **Compact** the modified XML via `compactXml` to remove extra whitespace
5. **Write back** to the ZIP entry using `zip.file(targetFile, compacted)`

This approach ensures that only the intended XML fragments change while the surrounding document structure remains intact.

### Header and Footer Support

When `old_string` is not found in [`word/document.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/word/document.xml), the handler automatically scans common header and footer parts (lines 62-75). The system checks [`header1.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/header1.xml), [`footer2.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/footer2.xml), and related relationship files, applying the same replacement logic to these ancillary document parts. This enables the `edit_block` API to modify running headers and footers without requiring explicit XML file targeting.

## The XML Surgical Edit Workflow

The architectural flow through [`src/utils/files/factory.ts`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/src/utils/files/factory.ts) and [`src/utils/files/docx.ts`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/src/utils/files/docx.ts) follows a precise pattern:

1. **Factory Routing**: `getFileHandler()` checks file extensions and routes `.docx` files to `DocxFileHandler` before generic text handlers process the request
2. **Deterministic Preparation**: The `prettyPrintXml` function normalizes XML formatting to create line-stable output that matches exactly between read and edit operations
3. **Validation and Replacement**: The system validates fragment occurrences before executing the surgical edit, then regenerates the ZIP archive preserving all non-modified parts (styles, relationships, media)

## Practical Code Examples

### Retrieving Document Structure

Request the default outline view to understand document organization before editing:

```typescript
const outline = await read_file({
  path: "/tmp/contract.docx",
});
console.log(outline);
// Output shows hierarchical structure:
// - Heading: Introduction
// - Paragraph: This agreement...
// - Table: [2 rows]
// - Image: [image1.png]

```

### Reading Raw XML for Editing

Access the underlying XML to copy exact fragments for surgical editing:

```typescript
const xmlFragment = await read_file({
  file_path: "/tmp/contract.docx",
  offset: 150,
  length: 50,
});
// Returns lines 150-200 of pretty-printed document.xml
// <w:t>This is the target text for editing</w:t>

```

### Performing Surgical Edits

Replace exact XML fragments using `edit_block` with `old_string` and `new_string`:

```typescript
await edit_block({
  file_path: "/tmp/contract.docx",
  old_string: '<w:t>This is the target text for editing</w:t>',
  new_string: '<w:t>This text has been surgically replaced</w:t>',
  expected_replacements: 1,
});

```

### Editing Headers and Footers

The same API modifies headers and footers when the target text resides outside the main document body:

```typescript
await edit_block({
  file_path: "/tmp/contract.docx",
  old_string: '<w:t>Old Company Name</w:t>',
  new_string: '<w:t>Acme Corporation</w:t>',
});
// Automatically finds and updates content in header1.xml if not in document.xml

```

### Creating New DOCX Files

Generate new Word documents from Markdown-style content:

```typescript
await write_file({
  path: "/tmp/new.docx",
  content: `

# Executive Summary

## Key Findings

The analysis reveals significant improvements in processing speed.

## Recommendations

Implement the surgical editing approach for all document modifications.
`,
});
// Converts # headings to Word headings and lines to paragraphs

```

## Summary

- DesktopCommanderMCP manipulates DOCX files by treating them as ZIP archives containing XML parts rather than binary blobs
- The `DocxFileHandler` in [`src/utils/files/docx.ts`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/src/utils/files/docx.ts) provides dual-mode reading via `read()`: human-readable outlines (offset 0) or raw pretty-printed XML (offset > 0)
- Surgical editing uses `prettyPrintXml` to normalize whitespace, executes literal string replacement via `editRange`, and compacts results with `compactXml` before repackaging
- Header and footer editing occurs automatically when `old_string` is not found in the main [`document.xml`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/document.xml) file
- The architecture preserves all non-modified document components (styles, relationships, embedded media) during the editing process

## Frequently Asked Questions

### What makes XML surgical editing different from traditional DOCX libraries?

Traditional DOCX libraries parse documents into complex object models requiring extensive APIs to navigate and modify content. DesktopCommanderMCP uses direct string replacement on pretty-printed XML fragments, eliminating heavy dependencies while providing LLM-friendly interfaces. This approach treats the document as editable text rather than an object hierarchy, simplifying the mental model for automated editing tasks.

### How does DesktopCommanderMCP handle formatting consistency across different DOCX generators?

The handler calls `prettyPrintXml` before any editing operation to normalize XML formatting into a deterministic, line-stable structure. This standardization ensures that `old_string` values copied from the raw XML output match exactly with the document's internal representation, regardless of whether the DOCX was generated by Microsoft Word, Google Docs, or LibreOffice.

### Can this method modify tables and images within DOCX files?

Table text content is fully editable because it exists as text nodes within `<w:t>` tags in the XML structure. However, image manipulation requires updating relationship files (`/_rels/*.rels`) and binary data in the `word/media/` directory, which falls outside the scope of string-based surgical edits currently implemented in [`src/utils/files/docx.ts`](https://github.com/wonderwhy-er/DesktopCommanderMCP/blob/main/src/utils/files/docx.ts). The `getInfo()` method (lines 131-140) can report image counts, but modifying them requires additional binary handling not present in the current XML surgical edit workflow.

### Why does the system validate expected_replacements before editing?

The `editRange` method counts occurrences of `old_string` in the pretty-printed XML and validates this count against the `expected_replacements` parameter (default 1). This safety mechanism prevents accidental global replacements that could corrupt document structure. By requiring explicit confirmation of replacement scope, the system ensures surgical precision—modifying only the intended instances while leaving similar text elsewhere in the document untouched.