How nGPT Generates Git Commit Messages from Staged Changes: Chunking Strategies Explained

nGPT uses a three-phase pipeline to generate git commit messages from staged changes: it first gathers the diff via git diff --staged, optionally splits large diffs into line-based chunks for recursive analysis, and finally synthesizes a Conventional Commit-style message with post-processing optimization.

The nazdridoy/ngpt repository provides a dedicated git-commit-message mode (-g or --gitcommsg) that transforms your staged changes into structured commit messages. This article breaks down exactly how nGPT generates git commit messages from staged changes, with specific focus on the chunking strategies that handle large diffs.

Phase 1: Gathering the Diff from Staged Changes

The process begins in ngpt/cli/modes/gitcommsg.py with the get_diff_content() function (lines 16-53). This utility collects the raw diff text through three possible methods:

  1. File input: If --diff <file> is specified, nGPT reads the diff directly from the provided file path.
  2. Git staged changes: By default, nGPT executes git diff --staged via subprocess to capture currently staged modifications.
  3. Stdin pipe: When --pipe is enabled, the tool reads the diff from sys.stdin, enabling integration with CI pipelines or other git hooks.

If no staged changes are detected, the function returns None and the CLI exits with an informative message.

Phase 2: Chunking Strategies for Large Diffs

Once the diff is collected, nGPT evaluates whether to process it as a single block or split it into manageable chunks. This decision point occurs in process_with_chunking() (lines 71-87) and depends on whether the --rec-chunk flag is enabled.

Simple Processing Without Chunking

When --rec-chunk is not set, the entire diff is passed directly to the LLM with the commit-message system prompt (create_system_prompt). This approach works efficiently for typical diffs under approximately 2,000 lines and minimizes API calls.

Line-Based Chunking with --chunk-size

For recursive processing, nGPT first splits the diff using split_into_chunks() (lines 55-64):

def split_into_chunks(content, chunk_size=200):
    lines = content.splitlines()
    return ["\n".join(lines[i:i+chunk_size])
            for i in range(0, len(lines), chunk_size)]

By default, the chunk size is 200 lines, configurable via --chunk-size or the chunk-size setting in ngpt/core/cli_config.py. Each chunk receives its own technical analysis prompt (create_technical_analysis_system_prompt), producing a human-readable summary of that specific code section.

Recursive Chunking with --rec-chunk

The recursive strategy activates when the combined technical analyses still exceed --analyses-chunk-size (default 200 lines). The recursive_chunk_analysis() function handles this by:

  1. Splitting the combined analyses into new chunks
  2. Sending each to the LLM for further summarization
  3. Recombining the results
  4. Recursing until the text fits within the chunk size limit or --max-recursion-depth (default 3) is reached

This hierarchical summarization ensures that even massive diffs (10,000+ lines) can be distilled into a coherent commit message without exceeding context window limits.

Combining Chunk Analyses into Final Messages

After all chunks are processed, create_combine_prompt() assembles a synthesis prompt:

def create_combine_prompt(partial_analyses):
    return f"""You have the following technical analyses of a git diff:\n\n{'\n\n'.join(partial_analyses)}\n\nWrite a concise conventional‑commit message…"""

The LLM then generates the final Conventional Commit-style message using the standard commit-message system prompt, ensuring the output follows the type(scope): description format with proper body and footer sections.

Phase 3: Generating and Post-Processing Commit Messages

The final phase ensures the generated message meets quality and length constraints through two specialized functions in ngpt/cli/modes/gitcommsg.py.

Condensing Long Messages

If the initial generation exceeds --max-msg-lines (default 20 lines), condense_commit_message() (lines 810-830) recursively sends the message back to the LLM with a condensation prompt. This process repeats up to --max-recursion-depth times, each iteration requesting a more concise version while preserving the Conventional Commit structure and all critical change details.

Optimizing File References

The optimize_file_references() function performs final cleanup by:

  • Removing redundant path prefixes when files are mentioned multiple times
  • Ensuring the first mention uses the full relative path
  • Stripping extraneous Markdown code-block markers
  • Standardizing bullet list formatting

This optimization ensures the commit message remains readable while accurately reflecting the scope of changes.

Configuration Options and CLI Flags

The following flags control how nGPT generates git commit messages from staged changes:

Flag Default Description
-g / --gitcommsg — Enable commit message generation mode
--diff [FILE] — Read diff from file instead of staged changes
--pipe — Accept diff via stdin
--rec-chunk False Enable recursive chunking for large diffs
--chunk-size 200 Lines per chunk when splitting diffs
--analyses-chunk-size 200 Maximum lines before triggering recursive analysis
--max-msg-lines 20 Maximum lines in final commit message
--max-recursion-depth 3 Maximum recursion levels for chunking and condensation

Configuration defaults are stored in ngpt/core/cli_config.py and can be overridden via the CLI or environment variables.

Example Usage Commands

Generate a commit message from currently staged changes:

ngpt -g

Process a specific diff file with custom chunking:

ngpt -g --diff changes.patch --chunk-size 150

Handle massive diffs with recursive analysis:

ngpt -g --rec-chunk --chunk-size 200 --max-recursion-depth 4

Integrate with CI pipelines using stdin:

git diff --staged | ngpt -g --pipe

Limit output length for concise messages:

ngpt -g --max-msg-lines 10

Summary

  • nGPT generates git commit messages from staged changes through a three-phase pipeline implemented in ngpt/cli/modes/gitcommsg.py.
  • Diff collection supports staged changes, file input, or stdin pipes via get_diff_content().
  • Chunking strategies include simple processing (no chunking), line-based splitting (default 200 lines), and recursive chunking (--rec-chunk) for massive diffs.
  • Technical analysis prompts process individual chunks before synthesis into a final Conventional Commit message.
  • Post-processing includes recursive condensation (condense_commit_message()) for length limits and file reference optimization (optimize_file_references()).
  • Configuration is controlled via CLI flags like --chunk-size, --max-msg-lines, and --max-recursion-depth.

Frequently Asked Questions

How does nGPT handle very large diffs that exceed the LLM context window?

nGPT handles large diffs through recursive chunking activated by the --rec-chunk flag. The diff is first split into 200-line chunks (configurable via --chunk-size). Each chunk receives a technical analysis prompt to summarize its contents. If the combined analyses still exceed the --analyses-chunk-size limit (default 200 lines), the recursive_chunk_analysis() function recursively summarizes the summaries up to --max-recursion-depth levels (default 3). This hierarchical approach ensures even 10,000+ line diffs can be processed within context limits.

What is the difference between --chunk-size and --analyses-chunk-size?

The --chunk-size parameter controls how many lines each initial diff chunk contains when splitting the raw git diff (default 200 lines). This determines how the original code changes are divided for parallel processing. The --analyses-chunk-size parameter controls the threshold for recursive analysis of the combined technical summaries. If the concatenated analyses from all chunks exceed this limit (default 200 lines), nGPT triggers another round of chunking and summarization. Essentially, --chunk-size governs the input segmentation, while --analyses-chunk-size governs the intermediate summary recursion threshold.

Can I use nGPT to generate commit messages for unstaged changes or specific files?

Yes, nGPT supports multiple input methods beyond staged changes. You can use the --diff <file> flag to read a diff from a specific patch file, allowing you to generate messages for arbitrary changes. Alternatively, you can use the --pipe flag to stream diff content via stdin, enabling workflows like git diff HEAD~1 | ngpt -g --pipe to generate messages for unstaged changes or specific commit ranges. The get_diff_content() function in ngpt/cli/modes/gitcommsg.py handles all three input methods: staged changes (default), file input, and stdin pipes.

How does nGPT ensure the generated commit message follows the Conventional Commit specification?

nGPT enforces the Conventional Commit format through a specialized system prompt created by create_system_prompt() in ngpt/cli/modes/gitcommsg.py. This prompt explicitly instructs the LLM to output messages in the type(scope): description format with optional body and footer sections. Additionally, the condense_commit_message() function recursively refines messages that exceed --max-msg-lines, ensuring the output remains concise while preserving the required structure. Finally, optimize_file_references() cleans up file paths and removes formatting artifacts, ensuring the final message adheres to clean Conventional Commit standards suitable for automated changelog generation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →