# How OpenClaude Builds a Repo Map for Codebase Awareness

> Discover how OpenClaude builds a repo map for codebase awareness. Learn about its seven-stage pipeline, from parsing with Tree-sitter to rendering a summary of key definitions.

- Repository: [Gitlawb/openclaude](https://github.com/Gitlawb/openclaude)
- Tags: architecture
- Published: 2026-09-06

---

**OpenClaude constructs a repo map through a seven-stage pipeline that parses the repository with Tree-sitter, builds a weighted reference graph, ranks files using PageRank, and renders a token-budgeted summary of the most important definitions.**

OpenClaude is an open-source coding assistant that injects repository context into LLM prompts without exceeding token limits. The system generates a **repo map**—a condensed structural summary—by analyzing source files, calculating symbol importance, and prioritizing definitions that matter most. This process is orchestrated by the `buildRepoMap` function in [`src/context/repoMap/index.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/index.ts) and leverages graph algorithms to surface critical code relationships.

## The Seven-Stage Repo Map Pipeline

OpenClaude's repo generation follows a deterministic pipeline designed for speed and repeatability. According to the source code in [`src/context/repoMap/index.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/index.ts) (lines 94-78), the `buildRepoMap` function executes these stages sequentially, with aggressive caching to avoid re-parsing unchanged files.

### 1. File Discovery with Git Integration

The pipeline begins by inventorying the codebase. The `getRepoFiles` function in [`src/context/repoMap/gitFiles.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/gitFiles.ts) (lines 24-33) first attempts to use `git ls-files` for rapid enumeration, falling back to a manual directory walk only when necessary. This ensures the system respects `.gitignore` rules while maintaining performance on large repositories.

### 2. Intelligent Cache Validation

Before parsing, OpenClaude computes a deterministic hash of the repository configuration—including the file list, token budget, and focus settings—using `computeMapHash`. If `loadCache` finds an existing rendered map for this hash in [`src/context/repoMap/index.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/index.ts) (lines 15-27), the function returns immediately, eliminating redundant computation.

### 3. Symbol Extraction via Tree-sitter

For each uncached file, `extractTags` in [`src/context/repoMap/symbolExtractor.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/symbolExtractor.ts) (lines 73-45) initializes a Tree-sitter parser and executes language-specific queries to identify **definition** (`def`) and **reference** (`ref`) tags. Results are cached per-file to accelerate subsequent builds, storing the structural metadata needed for graph construction.

### 4. Weighted Reference Graph Construction

OpenClaude models the codebase as a directed graph where nodes represent files and edges indicate cross-file references. The `buildGraph` function in [`src/context/repoMap/graph.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/graph.ts) (lines 24-90) creates an edge *A → B* when file A references a symbol defined in file B. Edge weights follow the formula `refCount × idf(symbol)`, where **IDF** (Inverse Document Frequency) penalizes common symbols listed in `COMMON_NAMES`, ensuring ubiquitous utilities don't dominate the rankings.

### 5. PageRank-Based File Ranking

With the graph constructed, `rankFiles` in [`src/context/repoMap/pagerank.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/pagerank.ts) (lines 43-78) applies the PageRank algorithm from `graphology-metrics` to determine file importance. The system then applies **focus boosts**: if specific files or symbols are requested via the `focusFiles` or `focusSymbols` parameters, those nodes and their direct neighbors receive elevated scores to surface contextually relevant code.

### 6. Token-Budgeted Rendering

The final assembly happens in `renderMap` within [`src/context/repoMap/renderer.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/renderer.ts) (lines 18-52). Files are processed in descending rank order, with `renderFileSection` emitting only definition tags for each. The `countTokens` helper verifies that adding each section stays within the `maxTokens` budget; if a file would exceed the limit, it is omitted entirely rather than truncated, preserving structural integrity.

### 7. Cache Persistence

After rendering, `setRenderedCache` and `saveCache` in [`src/context/repoMap/index.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/index.ts) (lines 66-70) persist the map, token count, and metadata to disk. Future invocations with identical configurations retrieve this cache instantly, making iterative development workflows responsive even on multi-million line codebases.

## Practical Usage Examples

The `buildRepoMap` API accepts configuration options to customize token budgets and focus areas.

Basic invocation using defaults:

```typescript
import { buildRepoMap } from './src/context/repoMap/index.js';

// Uses current working directory with 2048-token default
const result = await buildRepoMap();
console.log(`Generated map from ${result.fileCount} files`);
console.log(result.map);

```

Focus-aware generation prioritizing specific symbols:

```typescript
const result = await buildRepoMap({
  maxTokens: 1500,
  focusFiles: ['src/tools/RepoMapTool/RepoMapTool.ts'],
  focusSymbols: ['buildRepoMap'],
});

```

Both calls return an object containing:

- `map`: The rendered string suitable for LLM injection
- `cacheHit`: Boolean indicating if a cached version was returned
- `buildTimeMs`: Duration of the operation
- `fileCount`: Number of files included in the final map
- `totalFileCount`: Total files discovered in the repository
- `tokenCount`: Actual tokens consumed by the map

## Summary

- **OpenClaude** generates repo maps through a deterministic pipeline orchestrated by `buildRepoMap` in [`src/context/repoMap/index.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/index.ts).
- The system uses **Tree-sitter** parsers via `extractTags` to extract definitions and references with language-specific accuracy.
- A **weighted directed graph** is constructed in [`src/context/repoMap/graph.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/graph.ts), where edge weights use IDF scoring to penalize common symbols.
- **PageRank** in [`src/context/repoMap/pagerank.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/pagerank.ts) identifies important files, with optional boosts for focus files and their neighbors.
- The **token budget** is strictly enforced in [`src/context/repoMap/renderer.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/renderer.ts), rendering only complete file sections that fit within `maxTokens`.
- **Multi-layer caching** at the symbol extraction and final render levels eliminates redundant parsing across repeated calls.

## Frequently Asked Questions

### How does OpenClaude handle repositories that are not Git repositories?

When `git ls-files` fails or returns no results, the `getRepoFiles` function in [`src/context/repoMap/gitFiles.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/gitFiles.ts) automatically falls back to a manual directory traversal. This allows the repo map generator to function on any local directory structure, though it may include files that would normally be excluded by `.gitignore` rules.

### Why does OpenClaude use PageRank instead of simple reference counting?

Simple reference counts would over-value utility files used everywhere. By implementing **PageRank** in [`src/context/repoMap/pagerank.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/pagerank.ts) (lines 43-78), OpenClaude treats the codebase as a link graph where a reference from a highly-ranked file carries more weight than one from an obscure module. This surfaces architectural core files rather than just frequently-imported utilities.

### What determines whether a file is included in the final repo map?

Files are ranked by importance, then processed in descending order by `renderMap` in [`src/context/repoMap/renderer.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/renderer.ts). A file is included only if adding its definition tags would not exceed the `maxTokens` budget, as verified by `countTokens`. Lower-ranked files are dropped entirely once the budget is exhausted, ensuring the most critical definitions are always preserved.

### How does the caching mechanism improve performance?

OpenClaude implements two caching layers: **per-file symbol caches** store parsed AST results in [`src/context/repoMap/symbolExtractor.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/symbolExtractor.ts), while **rendered map caches** store final output hashes in [`src/context/repoMap/index.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/context/repoMap/index.ts). The system computes a configuration hash via `computeMapHash` (lines 15-27) to instantly return identical maps for unchanged repositories, reducing build times from seconds to milliseconds on subsequent runs.