# How nGPT Processes Large Diffs for Git Commit Message Generation

> Discover how nGPT processes large diffs for Git commit message generation. Learn its unique chunking and recursive summarization technique for efficient analysis.

- Repository: [nazDridoy/ngpt](https://github.com/nazdridoy/ngpt)
- Tags: internals
- Published: 2026-03-07

---

**nGPT splits large diffs into fixed-size chunks of 200 lines, processes each chunk with a technical analysis prompt to generate intermediate summaries, recursively re-chunks those summaries if they exceed safe token limits, and finally condenses the combined analysis into a conventional commit message that respects the configured line budget.**

When working with extensive code changes in the `nazdridoy/ngpt` repository, feeding entire raw diffs directly to an LLM quickly exhausts token limits. Instead, nGPT implements a sophisticated chunking pipeline in [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py) that preserves semantic context across massive changes while respecting model constraints.

## The Challenge of Token Limits with Large Diffs

Standard LLM APIs impose strict context windows that make submitting entire repository diffs impossible for substantial changes. Raw diff content grows linearly with modified lines, often exceeding 8k–32k token windows when hundreds of files change. nGPT solves this by treating the diff as a stream that must be processed in segments rather than a monolithic block, ensuring no single request exceeds safe limits.

## The nGPT Chunking Pipeline

The workflow centers on [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py), which orchestrates acquisition, splitting, analysis, and condensation through a series of specialized functions.

### Acquiring and Logging the Diff

The process starts with `get_diff_content()` at line 16 of [`gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/gitcommsg.py), which either reads a provided file path or executes `git diff --staged` to capture staged changes. Immediately after acquisition, `logger.log_diff()`—defined in [`ngpt/core/log.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/log.py) at line 390—records the raw content to `~/.ngpt/logs` for debugging purposes.

### Splitting Diffs into Line-Bounded Chunks

To ensure each request stays comfortably under token budgets, `split_into_chunks()` at line 55 divides the diff using a default **chunk_size of 200 lines**. This line-bounded approach maintains logical boundaries better than character-based splitting, preserving file header contexts and hunk structures within each chunk.

### Technical Analysis of Individual Chunks

The `process_with_chunking()` function iterates over each chunk, building a **technical analysis system prompt** via `create_technical_analysis_system_prompt()`. Each chunk is sent to the LLM through `handle_api_call()`, which wraps the `NGPTClient` from [`ngpt/api/client.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/api/client.py). During processing, `spinner()` from [`ngpt/ui/tui.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/ui/tui.py) provides visual feedback while the code respects rate limits between API calls. The resulting partial analyses are collected in a list and joined with double newlines (`"\n\n"`) to maintain clear separation between segment summaries.

## Recursive Summarization and Final Generation

After initial chunk processing, the system must ensure the combined analysis itself does not exceed limits before generating the final commit message.

### Recursive Chunk Analysis for Oversized Results

If the merged analyses exceed the safe threshold (`analyses_chunk_size`), `recursive_chunk_analysis()` at line 69 triggers automatically. This function treats the combined technical summaries as new input text, re-applying the chunking and analysis logic recursively until the total line count fits within the model's context window. This progressive summarization preserves the semantic essence of the entire diff while shrinking the payload to manageable size.

### Final Commit Message Generation

Once the combined analysis fits within limits, `create_combine_prompt()` assembles the final prompt and sends it with the **commit message system prompt** (`create_system_prompt()`). The LLM returns a draft conventional commit message based on the aggregated technical summaries rather than the raw diff text.

### Enforcing Line Budget Constraints

If the generated message exceeds `max_msg_lines` (defaulting to 20 lines), `condense_commit_message()` at line 808 recursively shortens the text. On the final recursion depth, the function forces the model to respect the exact line count, ensuring the output adheres to repository standards regardless of initial verbosity. Finally, `optimize_file_references()` cleans redundant path repetitions from the result.

## Complete Workflow Example

The following implementation demonstrates the full pipeline for processing massive diffs:

```python
from ngpt.cli.modes.gitcommsg import (
    get_diff_content,
    process_with_chunking,
    create_gitcommsg_logger,
)
from ngpt.api.client import NGPTClient

# 1. Gather the diff (staged changes)

diff = get_diff_content()  # Returns None if no staged changes

# 2. Set up a logger (writes to ~/.ngpt/logs by default)

logger = create_gitcommsg_logger()

# 3. Instantiate the NGPT API client

client = NGPTClient()

# 4. Run the full chunked pipeline

commit_msg = process_with_chunking(
    client=client,
    diff_content=diff,
    preprompt=None,          # optional user-supplied pre-prompt

    chunk_size=200,          # lines per chunk

    recursive=True,          # enable secondary chunking of analyses

    logger=logger,
    max_msg_lines=20,        # max lines for final message

    max_recursion_depth=3,   # safety guard for condensing

)

print("\nGenerated conventional commit message:\n")
print(commit_msg)

```

When executed against a repository with extensive changes, this code automatically splits the diff, generates technical summaries per segment, recursively compresses intermediate results if necessary, and outputs a polished conventional commit message satisfying the line-count rule.

## Summary

- **nGPT manages large diffs** by splitting them into 200-line chunks in [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py) to respect LLM token limits.
- **Each chunk undergoes independent technical analysis** using `process_with_chunking()`, with visual feedback provided by `spinner()` from [`ngpt/ui/tui.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/ui/tui.py).
- **Recursive summarization** via `recursive_chunk_analysis()` handles cases where combined intermediate analyses grow too large for the context window.
- **Strict output constraints** are enforced by `condense_commit_message()`, which recursively shortens text until it fits within `max_msg_lines` (default 20).
- **Full traceability** is maintained through `logger.log_diff()` in [`ngpt/core/log.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/log.py), recording every stage of the pipeline for debugging.

## Frequently Asked Questions

### How does nGPT determine the chunk size for large diffs?

nGPT uses a default `chunk_size` of 200 lines defined in `split_into_chunks()` at line 55 of [`ngpt/cli/modes/gitcommsg.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/modes/gitcommsg.py). This value keeps each request safely under typical token limits while preserving enough context for meaningful technical analysis. Users can override this parameter when calling `process_with_chunking()`.

### What happens if the technical analysis of all chunks combined is still too large?

When the combined line count of partial analyses exceeds `analyses_chunk_size`, the system invokes `recursive_chunk_analysis()` at line 69. This function re-chunks the intermediate summaries and re-analyzes them, progressively condensing the content until it fits within the model's context window.

### How does nGPT ensure the final commit message isn't excessively long?

The `condense_commit_message()` function at line 808 recursively rewrites the message until it respects the `max_msg_lines` parameter, which defaults to 20 lines. At the final recursion depth, the prompt explicitly forces the LLM to adhere to the exact line budget, ensuring concise output regardless of the diff size.

### Where does nGPT log the intermediate processing steps for debugging?

All diff content, chunking details, and template prompts are logged through functions in [`ngpt/core/log.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/log.py). Specifically, `logger.log_diff()` at line 390 records the raw diff, while `logger.log_chunks()` and `logger.log_template()` track segmentation and prompt construction throughout the pipeline.