How SkillSpector Handles Large Codebases: Chunking and Token Management Strategy

TLDR: SkillSpector processes arbitrarily large repositories by splitting oversized files into overlapping line-based chunks and treating each chunk as an independent LLM call, ensuring no request exceeds the model's token limit while preserving precise line-level traceability.

Modern AI code analysis tools often fail when confronted with enterprise-scale repositories containing multi-megabyte source files or monorepos with tens of thousands of lines. NVIDIA's SkillSpector solves this scalability challenge through a sophisticated chunking architecture implemented in src/skillspector/llm_analyzer_base.py. This system dynamically splits files into token-safe batches with overlapping context boundaries, allowing the tool to analyze codebases of any size without truncation or loss of semantic context.

Token Budget Calculation and Estimation

Before chunking begins, SkillSpector determines the available context window for the selected model. The get_max_input_tokens function in src/skillspector/model_info.py retrieves the model-specific token limit from the internal registry. This value sets the hard ceiling for each LLM request.

To avoid expensive tokenization calls during the splitting phase, SkillSpector uses a fast heuristic defined in the analyzer base. The system assumes approximately 4 characters per token (CHARS_PER_TOKEN = 4), allowing the estimate_tokens helper to quickly convert string lengths into approximate token counts. This lightweight estimation enables rapid file segmentation without loading heavy tokenization libraries for every chunking decision.

Line-Based Chunking with Overlap

The core scaling mechanism resides in chunk_file_by_lines within src/skillspector/llm_analyzer_base.py. This method processes files line-by-line, accumulating token counts until adding the next line would exceed the budget determined by the model's maximum input tokens.

When a chunk reaches capacity, the system emits it alongside its start and end line numbers. To prevent findings from being lost at chunk boundaries, consecutive chunks share 50 lines of context by default (CHUNK_OVERLAP_LINES = 50). This overlap ensures that code patterns spanning the boundary retain sufficient surrounding context for accurate analysis.

Batch Construction and Prompt Engineering

The LLMAnalyzerBase.get_batches method constructs a Batch object for every file. If the file fits within the token budget, it becomes a single batch containing the complete source. Otherwise, the method splits the file using the line-based chunking strategy, creating multiple batches that each record the original file path, chunk text, and precise line range (start_line/end_line).

During prompt construction, build_prompt uses number_lines to prefix each line with its real line number in L<nn> format. This numbering allows the LLM to report findings with exact line references that map directly back to the original source file, maintaining traceability across the chunked analysis.

Asynchronous Execution and Result Merging

The run_batches method (and its async counterpart arun_batches) sends each batch to the LLM independently. Because every batch respects the model's token limit, no request is truncated, even for files containing thousands of lines of code. After all batches return, the analyzer reassembles the findings, preserving the correct file and line numbers, and passes them downstream to the meta-analyzer or reporting nodes.

Implementation Example

The following example demonstrates how to initialize an analyzer and process large files using the batching system:

from skillspector.llm_analyzer_base import LLMAnalyzerBase

# Initialise the analyzer with a prompt and the target model

analyzer = LLMAnalyzerBase(
    base_prompt="Find security issues in the following code.",
    model="gpt-4o-mini",
)

# Suppose `file_cache` holds the raw contents of the files to analyze

file_cache = {"big_file.py": open("big_file.py").read()}
batches = analyzer.get_batches(list(file_cache), file_cache)

for batch in batches:
    print(f"{batch.file_label} → {len(batch.content)} chars")

To execute the analysis asynchronously:

import asyncio

async def analyze():
    results = await analyzer.arun_batches(batches, metadata_text="repo metadata")
    for batch, findings in results:
        print(batch.file_label, "found", len(findings), "issues")

asyncio.run(analyze())

Key Files in the Architecture

Understanding the codebase requires familiarity with these specific modules:

Summary

  • Token-aware chunking: SkillSpector calculates model-specific limits using get_max_input_tokens and estimates usage with a 4-character-per-token heuristic.
  • Overlapping line-based splits: The chunk_file_by_lines method creates chunks with 50 lines of overlap to maintain context across boundaries.
  • Independent batch processing: Each chunk becomes a separate LLM call via run_batches or arun_batches, preventing context window overflows.
  • Precise traceability: Line numbering with L<nn> prefixes ensures findings map accurately back to original source locations.
  • Scalable architecture: The system handles monorepos, multi-megabyte files, and tens of thousands of lines without truncation.

Frequently Asked Questions

How does SkillSpector avoid exceeding LLM token limits?

SkillSpector respects token limits by first querying the model's maximum input capacity via get_max_input_tokens in src/skillspector/model_info.py. It then uses the estimate_tokens function to check if files fit within the budget. Oversized files are automatically split into smaller chunks using chunk_file_by_lines, with each chunk processed as a separate LLM call.

What is the default chunk overlap size?

The default overlap is 50 lines, defined by the CHUNK_OVERLAP_LINES constant. This overlap ensures that code patterns or vulnerabilities spanning chunk boundaries retain sufficient context for accurate detection and reporting.

How does SkillSpector maintain line number accuracy across chunks?

The build_prompt method prefixes every line with its absolute line number in L<nn> format using the number_lines utility. When chunks are created, the Batch object stores the original start_line and end_line values, allowing the system to map LLM findings back to the exact locations in the original source file regardless of chunking.

Can SkillSpector handle monorepos with thousands of files?

Yes. The batching architecture treats each file independently, creating one or more batches per file depending on size. The arun_batches method processes these batches asynchronously, allowing SkillSpector to scale to repositories containing thousands of files or individual files with tens of thousands of lines without exceeding API rate limits or token constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →