How OpenClaude Builds a Repo Map for Codebase Awareness
OpenClaude constructs a repo map through a seven-stage pipeline that parses the repository with Tree-sitter, builds a weighted reference graph, ranks files using PageRank, and renders a token-budgeted summary of the most important definitions.
OpenClaude is an open-source coding assistant that injects repository context into LLM prompts without exceeding token limits. The system generates a repo map—a condensed structural summary—by analyzing source files, calculating symbol importance, and prioritizing definitions that matter most. This process is orchestrated by the buildRepoMap function in src/context/repoMap/index.ts and leverages graph algorithms to surface critical code relationships.
The Seven-Stage Repo Map Pipeline
OpenClaude's repo generation follows a deterministic pipeline designed for speed and repeatability. According to the source code in src/context/repoMap/index.ts (lines 94-78), the buildRepoMap function executes these stages sequentially, with aggressive caching to avoid re-parsing unchanged files.
1. File Discovery with Git Integration
The pipeline begins by inventorying the codebase. The getRepoFiles function in src/context/repoMap/gitFiles.ts (lines 24-33) first attempts to use git ls-files for rapid enumeration, falling back to a manual directory walk only when necessary. This ensures the system respects .gitignore rules while maintaining performance on large repositories.
2. Intelligent Cache Validation
Before parsing, OpenClaude computes a deterministic hash of the repository configuration—including the file list, token budget, and focus settings—using computeMapHash. If loadCache finds an existing rendered map for this hash in src/context/repoMap/index.ts (lines 15-27), the function returns immediately, eliminating redundant computation.
3. Symbol Extraction via Tree-sitter
For each uncached file, extractTags in src/context/repoMap/symbolExtractor.ts (lines 73-45) initializes a Tree-sitter parser and executes language-specific queries to identify definition (def) and reference (ref) tags. Results are cached per-file to accelerate subsequent builds, storing the structural metadata needed for graph construction.
4. Weighted Reference Graph Construction
OpenClaude models the codebase as a directed graph where nodes represent files and edges indicate cross-file references. The buildGraph function in src/context/repoMap/graph.ts (lines 24-90) creates an edge A → B when file A references a symbol defined in file B. Edge weights follow the formula refCount × idf(symbol), where IDF (Inverse Document Frequency) penalizes common symbols listed in COMMON_NAMES, ensuring ubiquitous utilities don't dominate the rankings.
5. PageRank-Based File Ranking
With the graph constructed, rankFiles in src/context/repoMap/pagerank.ts (lines 43-78) applies the PageRank algorithm from graphology-metrics to determine file importance. The system then applies focus boosts: if specific files or symbols are requested via the focusFiles or focusSymbols parameters, those nodes and their direct neighbors receive elevated scores to surface contextually relevant code.
6. Token-Budgeted Rendering
The final assembly happens in renderMap within src/context/repoMap/renderer.ts (lines 18-52). Files are processed in descending rank order, with renderFileSection emitting only definition tags for each. The countTokens helper verifies that adding each section stays within the maxTokens budget; if a file would exceed the limit, it is omitted entirely rather than truncated, preserving structural integrity.
7. Cache Persistence
After rendering, setRenderedCache and saveCache in src/context/repoMap/index.ts (lines 66-70) persist the map, token count, and metadata to disk. Future invocations with identical configurations retrieve this cache instantly, making iterative development workflows responsive even on multi-million line codebases.
Practical Usage Examples
The buildRepoMap API accepts configuration options to customize token budgets and focus areas.
Basic invocation using defaults:
import { buildRepoMap } from './src/context/repoMap/index.js';
// Uses current working directory with 2048-token default
const result = await buildRepoMap();
console.log(`Generated map from ${result.fileCount} files`);
console.log(result.map);
Focus-aware generation prioritizing specific symbols:
const result = await buildRepoMap({
maxTokens: 1500,
focusFiles: ['src/tools/RepoMapTool/RepoMapTool.ts'],
focusSymbols: ['buildRepoMap'],
});
Both calls return an object containing:
map: The rendered string suitable for LLM injectioncacheHit: Boolean indicating if a cached version was returnedbuildTimeMs: Duration of the operationfileCount: Number of files included in the final maptotalFileCount: Total files discovered in the repositorytokenCount: Actual tokens consumed by the map
Summary
- OpenClaude generates repo maps through a deterministic pipeline orchestrated by
buildRepoMapinsrc/context/repoMap/index.ts. - The system uses Tree-sitter parsers via
extractTagsto extract definitions and references with language-specific accuracy. - A weighted directed graph is constructed in
src/context/repoMap/graph.ts, where edge weights use IDF scoring to penalize common symbols. - PageRank in
src/context/repoMap/pagerank.tsidentifies important files, with optional boosts for focus files and their neighbors. - The token budget is strictly enforced in
src/context/repoMap/renderer.ts, rendering only complete file sections that fit withinmaxTokens. - Multi-layer caching at the symbol extraction and final render levels eliminates redundant parsing across repeated calls.
Frequently Asked Questions
How does OpenClaude handle repositories that are not Git repositories?
When git ls-files fails or returns no results, the getRepoFiles function in src/context/repoMap/gitFiles.ts automatically falls back to a manual directory traversal. This allows the repo map generator to function on any local directory structure, though it may include files that would normally be excluded by .gitignore rules.
Why does OpenClaude use PageRank instead of simple reference counting?
Simple reference counts would over-value utility files used everywhere. By implementing PageRank in src/context/repoMap/pagerank.ts (lines 43-78), OpenClaude treats the codebase as a link graph where a reference from a highly-ranked file carries more weight than one from an obscure module. This surfaces architectural core files rather than just frequently-imported utilities.
What determines whether a file is included in the final repo map?
Files are ranked by importance, then processed in descending order by renderMap in src/context/repoMap/renderer.ts. A file is included only if adding its definition tags would not exceed the maxTokens budget, as verified by countTokens. Lower-ranked files are dropped entirely once the budget is exhausted, ensuring the most critical definitions are always preserved.
How does the caching mechanism improve performance?
OpenClaude implements two caching layers: per-file symbol caches store parsed AST results in src/context/repoMap/symbolExtractor.ts, while rendered map caches store final output hashes in src/context/repoMap/index.ts. The system computes a configuration hash via computeMapHash (lines 15-27) to instantly return identical maps for unchanged repositories, reducing build times from seconds to milliseconds on subsequent runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →