# How Batch Computation Achieves Semantic Grouping for File Analysis in Egonex-AI Understand Anything

> Discover how Egonex-AI Understand Anything uses batch computation and Louvain community detection to semantically group files, optimizing analysis throughput and preserving module boundaries.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: internals
- Published: 2026-06-13

---

**Batch computation achieves semantic grouping by applying Louvain community detection to import graphs, extracting export symbols via Tree-Sitter, and applying domain-specific rules to non-code files, producing batches that preserve logical module boundaries while optimizing for analysis throughput.**

The Egonex-AI/Understand-Anything repository orchestrates complex code analysis through a Phase 1.5 pipeline that transforms raw [`scan-result.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/scan-result.json) data into semantically coherent file batches. At the core of this transformation lies [`compute-batches.mjs`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs), which implements the clustering logic that enables downstream file analyzers to process related code together.

## The Semantic Batch Generation Pipeline

The pipeline begins by ingesting the scanner output and progresses through symbol extraction, graph analysis, and intelligent grouping.

### Ingesting the Import Graph

The [`main()`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L64-L73) function reads [`scan-result.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/scan-result.json), which contains the `importMap` listing every file's dependencies and the `fileCategory` classifications distinguishing code from non-code assets. This import graph serves as the foundation for all subsequent clustering decisions.

### Extracting Export Symbols with Tree-Sitter

Before clustering, the system identifies public APIs to enable cross-batch dependency tracking. The [`extractExports`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L51-L71) function parses each code file using the `TreeSitterPlugin` to collect top-level exported symbol names. These symbols attach to neighbour entries later, allowing the analyzer to understand not just that files import from other batches, but specifically which symbols they reference.

### Louvain Community Detection for Code Files

The primary semantic grouping mechanism resides in the [`runLouvain`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L30-L48) function. This implementation:

1. Constructs an undirected graph using `graphology` with self-loops disabled.
2. Executes the Louvain community detection algorithm to identify tightly coupled modules.
3. Returns a `Map<path, communityId>` representing the initial semantic batches.

Files that import each other heavily receive the same community ID, ensuring logical modules remain together during analysis.

## Handling Non-Code Files and Edge Cases

Not all project files participate in the import graph. The system handles these through specialized grouping logic and fallback mechanisms.

### Semantic Grouping for Configuration Files

The [`buildNonCodeBatches`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L23-L98) function implements domain-aware clustering for infrastructure and configuration files:

- **Group A**: Dockerfile clusters grouped by parent directory (marked `mergeable: false`).
- **Group B**: GitHub Actions workflow files (`.github/workflows/*`).
- **Group C**: CI configuration files ([`.gitlab-ci.yml`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.gitlab-ci.yml), `.circleci/*`).
- **Group D**: SQL migration files grouped by `migrations/` directory.
- **Group E**: Catch-all parent-directory groups with a maximum of 20 files per batch (marked `mergeable: true`).

These groups preserve important logical boundaries—such as keeping Dockerfiles with their service context—while preventing configuration sprawl from creating noise in the code analysis batches.

### Fallback Strategies for Small or Failed Communities

When the Louvain algorithm fails due to missing native bindings or other errors, the system degrades gracefully to [`countBasedAssignment`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L52-L61), which assigns files deterministically into batches of 12 files each. This ensures the pipeline remains operational even when graph analysis dependencies are unavailable.

For successful Louvain runs, communities exceeding `MAX_COMMUNITY_SIZE` (35 files) undergo alphabetical splitting ([lines 100-124](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L100-L124)), preventing oversized batches that would overwhelm the analyzer agents.

## Optimizing Batch Size and Cross-Batch Relationships

After initial grouping, the pipeline refines batch boundaries and constructs dependency maps that enable cross-batch reasoning.

### Merging Small Batches

The [`mergeSmallBatches`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L78-L98) function pools batches marked as `mergeable` that contain fewer than `MIN_BATCH_SIZE` (3 files). These tiny batches aggregate into "misc" batches containing up to `MAX_MERGE_TARGET` (25 files), reducing dispatch overhead while maintaining the semantic integrity of the larger, well-defined groups.

### Building the Neighbor Map

To enable cross-batch analysis, the system constructs a `neighborMap` for each file ([lines 84-124](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L84-L124)). This structure records:

- `batchImportData`: Direct imports remaining within the same batch.
- Cross-batch neighbours: Files imported from other batches, including the target batch index and exported symbols.

The neighbour list respects `MAX_NEIGHBORS` (50 entries), with excess connections ranked by degree (inbound + outbound connection count) to retain the most significant cross-module dependencies.

## Implementation Examples

```bash

# Run the batch computation on a checked-out project

node ./understand-anything-plugin/skills/understand/compute-batches.mjs /path/to/project

```

```javascript
// Minimal programmatic usage of the Louvain algorithm
import { runLouvain } from './compute-batches.mjs';

const communityMap = runLouvain(codeFiles, importMap);
// communityMap is a Map<string, string> mapping file paths to community IDs

```

```javascript
// Merging tiny batches after initial grouping
import { mergeSmallBatches } from './compute-batches.mjs';

const merged = mergeSmallBatches(bareBatches);
// Returns optimized batches with consolidated small groups

```

## Output Format and Consumption

The final output, written to [`batches.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/batches.json), contains:

- `algorithm`: Either `louvain` or `count-fallback` indicating the grouping method used.
- `exportsByPath`: The symbol tables extracted during preprocessing.
- `batches`: An array of objects with `batchIndex`, `files`, `batchImportData`, and `neighborMap`.

This format is consumed by the file-analyzer agents, which can now leverage semantic context rather than processing arbitrary file chunks. The companion script [[`merge-batch-graphs.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-batch-graphs.py)](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/merge-batch-graphs.py) further aggregates these per-batch graphs into global analysis structures.

## Summary

- **Louvain clustering** on import graphs creates initial semantic batches that keep coupled code together.
- **Tree-Sitter parsing** extracts export symbols to enable precise cross-batch dependency tracking.
- **Domain-specific rules** group non-code files (Dockerfiles, CI configs, SQL migrations) by logical function and location.
- **Size enforcement** splits oversized communities (>35 files) and merges tiny batches (<3 files) to optimize analyzer throughput.
- **Neighbor maps** cap cross-batch references at 50 entries, prioritized by connection degree, ensuring analyzers receive relevant context without noise.

## Frequently Asked Questions

### What algorithm does the batch computation use for semantic grouping?

The system primarily uses the **Louvain community detection algorithm** via the `graphology` library to identify clusters of files with dense import interconnections. When this fails, it falls back to a deterministic count-based assignment of 12 files per batch.

### How does the system handle non-code files like Dockerfiles and CI configurations?

Non-code files bypass the import graph and enter **predefined semantic groups** in `buildNonCodeBatches`. Dockerfiles cluster by directory, GitHub workflows group together, SQL migrations organize by folder, and remaining files collect into parent-directory buckets. Groups A through D are protected from merging, while Group E participates in the small-batch consolidation.

### What happens if the Louvain algorithm fails to run?

If the Louvain algorithm throws—typically due to missing native bindings—the `runLouvain` function catches the error and the pipeline invokes `countBasedAssignment`. This fallback creates fixed-size batches of 12 files using alphabetical ordering, ensuring the analysis pipeline continues without graph-theoretic clustering.

### How are cross-batch dependencies tracked in the semantic grouping?

The system constructs a `neighborMap` during the finalization phase. For each file importing from another batch, the map records the neighbour's path, batch index, and exported symbols consumed. This map is capped at `MAX_NEIGHBORS` (50) and sorted by degree centrality, providing downstream analyzers with the most relevant cross-module context.