# How SkillSpector Handles Large Codebases: Chunking and Token Management Strategy

> SkillSpector processes large codebases by chunking files and managing tokens effectively. Discover its strategy for handling extensive repositories without exceeding model limits.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: internals
- Published: 2026-07-10

---

**TLDR:** SkillSpector processes arbitrarily large repositories by splitting oversized files into overlapping line-based chunks and treating each chunk as an independent LLM call, ensuring no request exceeds the model's token limit while preserving precise line-level traceability.

Modern AI code analysis tools often fail when confronted with enterprise-scale repositories containing multi-megabyte source files or monorepos with tens of thousands of lines. NVIDIA's SkillSpector solves this scalability challenge through a sophisticated chunking architecture implemented in [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py). This system dynamically splits files into token-safe batches with overlapping context boundaries, allowing the tool to analyze codebases of any size without truncation or loss of semantic context.

## Token Budget Calculation and Estimation

Before chunking begins, SkillSpector determines the available context window for the selected model. The `get_max_input_tokens` function in [`src/skillspector/model_info.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/model_info.py) retrieves the model-specific token limit from the internal registry. This value sets the hard ceiling for each LLM request.

To avoid expensive tokenization calls during the splitting phase, SkillSpector uses a fast heuristic defined in the analyzer base. The system assumes approximately **4 characters per token** (`CHARS_PER_TOKEN = 4`), allowing the `estimate_tokens` helper to quickly convert string lengths into approximate token counts. This lightweight estimation enables rapid file segmentation without loading heavy tokenization libraries for every chunking decision.

## Line-Based Chunking with Overlap

The core scaling mechanism resides in `chunk_file_by_lines` within [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py). This method processes files line-by-line, accumulating token counts until adding the next line would exceed the budget determined by the model's maximum input tokens.

When a chunk reaches capacity, the system emits it alongside its start and end line numbers. To prevent findings from being lost at chunk boundaries, consecutive chunks share **50 lines of context** by default (`CHUNK_OVERLAP_LINES = 50`). This overlap ensures that code patterns spanning the boundary retain sufficient surrounding context for accurate analysis.

## Batch Construction and Prompt Engineering

The `LLMAnalyzerBase.get_batches` method constructs a `Batch` object for every file. If the file fits within the token budget, it becomes a single batch containing the complete source. Otherwise, the method splits the file using the line-based chunking strategy, creating multiple batches that each record the original file path, chunk text, and precise line range (`start_line`/`end_line`).

During prompt construction, `build_prompt` uses `number_lines` to prefix each line with its real line number in `L<nn>` format. This numbering allows the LLM to report findings with exact line references that map directly back to the original source file, maintaining traceability across the chunked analysis.

## Asynchronous Execution and Result Merging

The `run_batches` method (and its async counterpart `arun_batches`) sends each batch to the LLM independently. Because every batch respects the model's token limit, no request is truncated, even for files containing thousands of lines of code. After all batches return, the analyzer reassembles the findings, preserving the correct file and line numbers, and passes them downstream to the meta-analyzer or reporting nodes.

## Implementation Example

The following example demonstrates how to initialize an analyzer and process large files using the batching system:

```python
from skillspector.llm_analyzer_base import LLMAnalyzerBase

# Initialise the analyzer with a prompt and the target model

analyzer = LLMAnalyzerBase(
    base_prompt="Find security issues in the following code.",
    model="gpt-4o-mini",
)

# Suppose `file_cache` holds the raw contents of the files to analyze

file_cache = {"big_file.py": open("big_file.py").read()}
batches = analyzer.get_batches(list(file_cache), file_cache)

for batch in batches:
    print(f"{batch.file_label} → {len(batch.content)} chars")

```

To execute the analysis asynchronously:

```python
import asyncio

async def analyze():
    results = await analyzer.arun_batches(batches, metadata_text="repo metadata")
    for batch, findings in results:
        print(batch.file_label, "found", len(findings), "issues")

asyncio.run(analyze())

```

## Key Files in the Architecture

Understanding the codebase requires familiarity with these specific modules:

- **[`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py)**: Contains the core chunking logic (`chunk_file_by_lines`), the `Batch` dataclass, token budgeting, and batch execution methods.
- **[`src/skillspector/model_info.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/model_info.py)**: Provides `get_max_input_tokens` for retrieving model-specific context limits.
- **[`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py)**: Demonstrates how high-level nodes request batches and handle per-file results.
- **[`src/skillspector/llm_utils.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_utils.py)**: Supplies the `get_chat_model` factory used by `LLMAnalyzerBase`.
- **[`src/skillspector/nodes/build_context.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py)**: Part of the graph-building pipeline that feeds file contents into the analyzer.

## Summary

- **Token-aware chunking**: SkillSpector calculates model-specific limits using `get_max_input_tokens` and estimates usage with a 4-character-per-token heuristic.
- **Overlapping line-based splits**: The `chunk_file_by_lines` method creates chunks with 50 lines of overlap to maintain context across boundaries.
- **Independent batch processing**: Each chunk becomes a separate LLM call via `run_batches` or `arun_batches`, preventing context window overflows.
- **Precise traceability**: Line numbering with `L<nn>` prefixes ensures findings map accurately back to original source locations.
- **Scalable architecture**: The system handles monorepos, multi-megabyte files, and tens of thousands of lines without truncation.

## Frequently Asked Questions

### How does SkillSpector avoid exceeding LLM token limits?

SkillSpector respects token limits by first querying the model's maximum input capacity via `get_max_input_tokens` in [`src/skillspector/model_info.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/model_info.py). It then uses the `estimate_tokens` function to check if files fit within the budget. Oversized files are automatically split into smaller chunks using `chunk_file_by_lines`, with each chunk processed as a separate LLM call.

### What is the default chunk overlap size?

The default overlap is **50 lines**, defined by the `CHUNK_OVERLAP_LINES` constant. This overlap ensures that code patterns or vulnerabilities spanning chunk boundaries retain sufficient context for accurate detection and reporting.

### How does SkillSpector maintain line number accuracy across chunks?

The `build_prompt` method prefixes every line with its absolute line number in `L<nn>` format using the `number_lines` utility. When chunks are created, the `Batch` object stores the original `start_line` and `end_line` values, allowing the system to map LLM findings back to the exact locations in the original source file regardless of chunking.

### Can SkillSpector handle monorepos with thousands of files?

Yes. The batching architecture treats each file independently, creating one or more batches per file depending on size. The `arun_batches` method processes these batches asynchronously, allowing SkillSpector to scale to repositories containing thousands of files or individual files with tens of thousands of lines without exceeding API rate limits or token constraints.