# How nGPT Generates Git Commit Messages from Staged Changes: Chunking Strategies Explained

> Learn how nGPT generates Git commit messages. Discover its three-phase pipeline, including chunking strategies for analyzing staged changes and creating Conventional Commit messages.

- Repository: [nazDridoy/ngpt](https://github.com/nazdridoy/ngpt)
- Tags: how-to-guide
- Published: 2026-03-07

---

**nGPT uses a three-phase pipeline to generate git commit messages from staged changes: it first gathers the diff via `git diff --staged`, optionally splits large diffs into line-based chunks for recursive analysis, and finally synthesizes a Conventional Commit-style message with post-processing optimization.**

The `nazdridoy/ngpt` repository provides a dedicated **git-commit-message mode** (`-g` or `--gitcommsg`) that transforms your staged changes into structured commit messages. This article breaks down exactly how nGPT generates git commit messages from staged changes, with specific focus on the chunking strategies that handle large diffs.

## Phase 1: Gathering the Diff from Staged Changes

The process begins in [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py) with the `get_diff_content()` function (lines 16-53). This utility collects the raw diff text through three possible methods:

1.  **File input**: If `--diff <file>` is specified, nGPT reads the diff directly from the provided file path.
2.  **Git staged changes**: By default, nGPT executes `git diff --staged` via subprocess to capture currently staged modifications.
3.  **Stdin pipe**: When `--pipe` is enabled, the tool reads the diff from `sys.stdin`, enabling integration with CI pipelines or other git hooks.

If no staged changes are detected, the function returns `None` and the CLI exits with an informative message.

## Phase 2: Chunking Strategies for Large Diffs

Once the diff is collected, nGPT evaluates whether to process it as a single block or split it into manageable chunks. This decision point occurs in `process_with_chunking()` (lines 71-87) and depends on whether the `--rec-chunk` flag is enabled.

### Simple Processing Without Chunking

When `--rec-chunk` is **not** set, the entire diff is passed directly to the LLM with the commit-message system prompt (`create_system_prompt`). This approach works efficiently for typical diffs under approximately 2,000 lines and minimizes API calls.

### Line-Based Chunking with `--chunk-size`

For recursive processing, nGPT first splits the diff using `split_into_chunks()` (lines 55-64):

```python
def split_into_chunks(content, chunk_size=200):
    lines = content.splitlines()
    return ["\n".join(lines[i:i+chunk_size])
            for i in range(0, len(lines), chunk_size)]

```

By default, the **chunk size is 200 lines**, configurable via `--chunk-size` or the `chunk-size` setting in [`ngpt/core/cli_config.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/cli_config.py). Each chunk receives its own **technical analysis prompt** (`create_technical_analysis_system_prompt`), producing a human-readable summary of that specific code section.

### Recursive Chunking with `--rec-chunk`

The recursive strategy activates when the combined technical analyses still exceed `--analyses-chunk-size` (default 200 lines). The `recursive_chunk_analysis()` function handles this by:

1.  Splitting the combined analyses into new chunks
2.  Sending each to the LLM for further summarization
3.  Recombining the results
4.  Recursing until the text fits within the chunk size limit or `--max-recursion-depth` (default 3) is reached

This hierarchical summarization ensures that even massive diffs (10,000+ lines) can be distilled into a coherent commit message without exceeding context window limits.

### Combining Chunk Analyses into Final Messages

After all chunks are processed, `create_combine_prompt()` assembles a synthesis prompt:

```python
def create_combine_prompt(partial_analyses):
    return f"""You have the following technical analyses of a git diff:\n\n{'\n\n'.join(partial_analyses)}\n\nWrite a concise conventional‑commit message…"""

```

The LLM then generates the final Conventional Commit-style message using the standard commit-message system prompt, ensuring the output follows the `type(scope): description` format with proper body and footer sections.

## Phase 3: Generating and Post-Processing Commit Messages

The final phase ensures the generated message meets quality and length constraints through two specialized functions in [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py).

### Condensing Long Messages

If the initial generation exceeds `--max-msg-lines` (default 20 lines), `condense_commit_message()` (lines 810-830) recursively sends the message back to the LLM with a condensation prompt. This process repeats up to `--max-recursion-depth` times, each iteration requesting a more concise version while preserving the Conventional Commit structure and all critical change details.

### Optimizing File References

The `optimize_file_references()` function performs final cleanup by:

-   Removing redundant path prefixes when files are mentioned multiple times
-   Ensuring the first mention uses the full relative path
-   Stripping extraneous Markdown code-block markers
-   Standardizing bullet list formatting

This optimization ensures the commit message remains readable while accurately reflecting the scope of changes.

## Configuration Options and CLI Flags

The following flags control how nGPT generates git commit messages from staged changes:

| Flag | Default | Description |
|------|---------|-------------|
| `-g` / `--gitcommsg` | — | Enable commit message generation mode |
| `--diff [FILE]` | — | Read diff from file instead of staged changes |
| `--pipe` | — | Accept diff via stdin |
| `--rec-chunk` | False | Enable recursive chunking for large diffs |
| `--chunk-size` | 200 | Lines per chunk when splitting diffs |
| `--analyses-chunk-size` | 200 | Maximum lines before triggering recursive analysis |
| `--max-msg-lines` | 20 | Maximum lines in final commit message |
| `--max-recursion-depth` | 3 | Maximum recursion levels for chunking and condensation |

Configuration defaults are stored in [`ngpt/core/cli_config.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/cli_config.py) and can be overridden via the CLI or environment variables.

## Example Usage Commands

Generate a commit message from currently staged changes:

```bash
ngpt -g

```

Process a specific diff file with custom chunking:

```bash
ngpt -g --diff changes.patch --chunk-size 150

```

Handle massive diffs with recursive analysis:

```bash
ngpt -g --rec-chunk --chunk-size 200 --max-recursion-depth 4

```

Integrate with CI pipelines using stdin:

```bash
git diff --staged | ngpt -g --pipe

```

Limit output length for concise messages:

```bash
ngpt -g --max-msg-lines 10

```

## Summary

- **nGPT** generates git commit messages from staged changes through a three-phase pipeline implemented in [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py).
- **Diff collection** supports staged changes, file input, or stdin pipes via `get_diff_content()`.
- **Chunking strategies** include simple processing (no chunking), line-based splitting (default 200 lines), and recursive chunking (`--rec-chunk`) for massive diffs.
- **Technical analysis** prompts process individual chunks before synthesis into a final Conventional Commit message.
- **Post-processing** includes recursive condensation (`condense_commit_message()`) for length limits and file reference optimization (`optimize_file_references()`).
- **Configuration** is controlled via CLI flags like `--chunk-size`, `--max-msg-lines`, and `--max-recursion-depth`.

## Frequently Asked Questions

### How does nGPT handle very large diffs that exceed the LLM context window?

nGPT handles large diffs through **recursive chunking** activated by the `--rec-chunk` flag. The diff is first split into 200-line chunks (configurable via `--chunk-size`). Each chunk receives a technical analysis prompt to summarize its contents. If the combined analyses still exceed the `--analyses-chunk-size` limit (default 200 lines), the `recursive_chunk_analysis()` function recursively summarizes the summaries up to `--max-recursion-depth` levels (default 3). This hierarchical approach ensures even 10,000+ line diffs can be processed within context limits.

### What is the difference between `--chunk-size` and `--analyses-chunk-size`?

The `--chunk-size` parameter controls how many lines each **initial diff chunk** contains when splitting the raw git diff (default 200 lines). This determines how the original code changes are divided for parallel processing. The `--analyses-chunk-size` parameter controls the threshold for **recursive analysis** of the combined technical summaries. If the concatenated analyses from all chunks exceed this limit (default 200 lines), nGPT triggers another round of chunking and summarization. Essentially, `--chunk-size` governs the input segmentation, while `--analyses-chunk-size` governs the intermediate summary recursion threshold.

### Can I use nGPT to generate commit messages for unstaged changes or specific files?

Yes, nGPT supports multiple input methods beyond staged changes. You can use the `--diff <file>` flag to read a diff from a specific patch file, allowing you to generate messages for arbitrary changes. Alternatively, you can use the `--pipe` flag to stream diff content via stdin, enabling workflows like `git diff HEAD~1 | ngpt -g --pipe` to generate messages for unstaged changes or specific commit ranges. The `get_diff_content()` function in [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py) handles all three input methods: staged changes (default), file input, and stdin pipes.

### How does nGPT ensure the generated commit message follows the Conventional Commit specification?

nGPT enforces the Conventional Commit format through a specialized **system prompt** created by `create_system_prompt()` in [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py). This prompt explicitly instructs the LLM to output messages in the `type(scope): description` format with optional body and footer sections. Additionally, the `condense_commit_message()` function recursively refines messages that exceed `--max-msg-lines`, ensuring the output remains concise while preserving the required structure. Finally, `optimize_file_references()` cleans up file paths and removes formatting artifacts, ensuring the final message adheres to clean Conventional Commit standards suitable for automated changelog generation.