How nGPT Processes Large Diffs for Git Commit Message Generation
nGPT splits large diffs into fixed-size chunks of 200 lines, processes each chunk with a technical analysis prompt to generate intermediate summaries, recursively re-chunks those summaries if they exceed safe token limits, and finally condenses the combined analysis into a conventional commit message that respects the configured line budget.
When working with extensive code changes in the nazdridoy/ngpt repository, feeding entire raw diffs directly to an LLM quickly exhausts token limits. Instead, nGPT implements a sophisticated chunking pipeline in ngpt/cli/modes/gitcommsg.py that preserves semantic context across massive changes while respecting model constraints.
The Challenge of Token Limits with Large Diffs
Standard LLM APIs impose strict context windows that make submitting entire repository diffs impossible for substantial changes. Raw diff content grows linearly with modified lines, often exceeding 8k–32k token windows when hundreds of files change. nGPT solves this by treating the diff as a stream that must be processed in segments rather than a monolithic block, ensuring no single request exceeds safe limits.
The nGPT Chunking Pipeline
The workflow centers on ngpt/cli/modes/gitcommsg.py, which orchestrates acquisition, splitting, analysis, and condensation through a series of specialized functions.
Acquiring and Logging the Diff
The process starts with get_diff_content() at line 16 of gitcommsg.py, which either reads a provided file path or executes git diff --staged to capture staged changes. Immediately after acquisition, logger.log_diff()—defined in ngpt/core/log.py at line 390—records the raw content to ~/.ngpt/logs for debugging purposes.
Splitting Diffs into Line-Bounded Chunks
To ensure each request stays comfortably under token budgets, split_into_chunks() at line 55 divides the diff using a default chunk_size of 200 lines. This line-bounded approach maintains logical boundaries better than character-based splitting, preserving file header contexts and hunk structures within each chunk.
Technical Analysis of Individual Chunks
The process_with_chunking() function iterates over each chunk, building a technical analysis system prompt via create_technical_analysis_system_prompt(). Each chunk is sent to the LLM through handle_api_call(), which wraps the NGPTClient from ngpt/api/client.py. During processing, spinner() from ngpt/ui/tui.py provides visual feedback while the code respects rate limits between API calls. The resulting partial analyses are collected in a list and joined with double newlines ("\n\n") to maintain clear separation between segment summaries.
Recursive Summarization and Final Generation
After initial chunk processing, the system must ensure the combined analysis itself does not exceed limits before generating the final commit message.
Recursive Chunk Analysis for Oversized Results
If the merged analyses exceed the safe threshold (analyses_chunk_size), recursive_chunk_analysis() at line 69 triggers automatically. This function treats the combined technical summaries as new input text, re-applying the chunking and analysis logic recursively until the total line count fits within the model's context window. This progressive summarization preserves the semantic essence of the entire diff while shrinking the payload to manageable size.
Final Commit Message Generation
Once the combined analysis fits within limits, create_combine_prompt() assembles the final prompt and sends it with the commit message system prompt (create_system_prompt()). The LLM returns a draft conventional commit message based on the aggregated technical summaries rather than the raw diff text.
Enforcing Line Budget Constraints
If the generated message exceeds max_msg_lines (defaulting to 20 lines), condense_commit_message() at line 808 recursively shortens the text. On the final recursion depth, the function forces the model to respect the exact line count, ensuring the output adheres to repository standards regardless of initial verbosity. Finally, optimize_file_references() cleans redundant path repetitions from the result.
Complete Workflow Example
The following implementation demonstrates the full pipeline for processing massive diffs:
from ngpt.cli.modes.gitcommsg import (
get_diff_content,
process_with_chunking,
create_gitcommsg_logger,
)
from ngpt.api.client import NGPTClient
# 1. Gather the diff (staged changes)
diff = get_diff_content() # Returns None if no staged changes
# 2. Set up a logger (writes to ~/.ngpt/logs by default)
logger = create_gitcommsg_logger()
# 3. Instantiate the NGPT API client
client = NGPTClient()
# 4. Run the full chunked pipeline
commit_msg = process_with_chunking(
client=client,
diff_content=diff,
preprompt=None, # optional user-supplied pre-prompt
chunk_size=200, # lines per chunk
recursive=True, # enable secondary chunking of analyses
logger=logger,
max_msg_lines=20, # max lines for final message
max_recursion_depth=3, # safety guard for condensing
)
print("\nGenerated conventional commit message:\n")
print(commit_msg)
When executed against a repository with extensive changes, this code automatically splits the diff, generates technical summaries per segment, recursively compresses intermediate results if necessary, and outputs a polished conventional commit message satisfying the line-count rule.
Summary
- nGPT manages large diffs by splitting them into 200-line chunks in
ngpt/cli/modes/gitcommsg.pyto respect LLM token limits. - Each chunk undergoes independent technical analysis using
process_with_chunking(), with visual feedback provided byspinner()fromngpt/ui/tui.py. - Recursive summarization via
recursive_chunk_analysis()handles cases where combined intermediate analyses grow too large for the context window. - Strict output constraints are enforced by
condense_commit_message(), which recursively shortens text until it fits withinmax_msg_lines(default 20). - Full traceability is maintained through
logger.log_diff()inngpt/core/log.py, recording every stage of the pipeline for debugging.
Frequently Asked Questions
How does nGPT determine the chunk size for large diffs?
nGPT uses a default chunk_size of 200 lines defined in split_into_chunks() at line 55 of ngpt/cli/modes/gitcommsg.py. This value keeps each request safely under typical token limits while preserving enough context for meaningful technical analysis. Users can override this parameter when calling process_with_chunking().
What happens if the technical analysis of all chunks combined is still too large?
When the combined line count of partial analyses exceeds analyses_chunk_size, the system invokes recursive_chunk_analysis() at line 69. This function re-chunks the intermediate summaries and re-analyzes them, progressively condensing the content until it fits within the model's context window.
How does nGPT ensure the final commit message isn't excessively long?
The condense_commit_message() function at line 808 recursively rewrites the message until it respects the max_msg_lines parameter, which defaults to 20 lines. At the final recursion depth, the prompt explicitly forces the LLM to adhere to the exact line budget, ensuring concise output regardless of the diff size.
Where does nGPT log the intermediate processing steps for debugging?
All diff content, chunking details, and template prompts are logged through functions in ngpt/core/log.py. Specifically, logger.log_diff() at line 390 records the raw diff, while logger.log_chunks() and logger.log_template() track segmentation and prompt construction throughout the pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →