How Batch Computation Achieves Semantic Grouping for File Analysis in Egonex-AI Understand Anything

Batch computation achieves semantic grouping by applying Louvain community detection to import graphs, extracting export symbols via Tree-Sitter, and applying domain-specific rules to non-code files, producing batches that preserve logical module boundaries while optimizing for analysis throughput.

The Egonex-AI/Understand-Anything repository orchestrates complex code analysis through a Phase 1.5 pipeline that transforms raw scan-result.json data into semantically coherent file batches. At the core of this transformation lies compute-batches.mjs, which implements the clustering logic that enables downstream file analyzers to process related code together.

The Semantic Batch Generation Pipeline

The pipeline begins by ingesting the scanner output and progresses through symbol extraction, graph analysis, and intelligent grouping.

Ingesting the Import Graph

The main() function reads scan-result.json, which contains the importMap listing every file's dependencies and the fileCategory classifications distinguishing code from non-code assets. This import graph serves as the foundation for all subsequent clustering decisions.

Extracting Export Symbols with Tree-Sitter

Before clustering, the system identifies public APIs to enable cross-batch dependency tracking. The extractExports function parses each code file using the TreeSitterPlugin to collect top-level exported symbol names. These symbols attach to neighbour entries later, allowing the analyzer to understand not just that files import from other batches, but specifically which symbols they reference.

Louvain Community Detection for Code Files

The primary semantic grouping mechanism resides in the runLouvain function. This implementation:

  1. Constructs an undirected graph using graphology with self-loops disabled.
  2. Executes the Louvain community detection algorithm to identify tightly coupled modules.
  3. Returns a Map<path, communityId> representing the initial semantic batches.

Files that import each other heavily receive the same community ID, ensuring logical modules remain together during analysis.

Handling Non-Code Files and Edge Cases

Not all project files participate in the import graph. The system handles these through specialized grouping logic and fallback mechanisms.

Semantic Grouping for Configuration Files

The buildNonCodeBatches function implements domain-aware clustering for infrastructure and configuration files:

  • Group A: Dockerfile clusters grouped by parent directory (marked mergeable: false).
  • Group B: GitHub Actions workflow files (.github/workflows/*).
  • Group C: CI configuration files (.gitlab-ci.yml, .circleci/*).
  • Group D: SQL migration files grouped by migrations/ directory.
  • Group E: Catch-all parent-directory groups with a maximum of 20 files per batch (marked mergeable: true).

These groups preserve important logical boundaries—such as keeping Dockerfiles with their service context—while preventing configuration sprawl from creating noise in the code analysis batches.

Fallback Strategies for Small or Failed Communities

When the Louvain algorithm fails due to missing native bindings or other errors, the system degrades gracefully to countBasedAssignment, which assigns files deterministically into batches of 12 files each. This ensures the pipeline remains operational even when graph analysis dependencies are unavailable.

For successful Louvain runs, communities exceeding MAX_COMMUNITY_SIZE (35 files) undergo alphabetical splitting (lines 100-124), preventing oversized batches that would overwhelm the analyzer agents.

Optimizing Batch Size and Cross-Batch Relationships

After initial grouping, the pipeline refines batch boundaries and constructs dependency maps that enable cross-batch reasoning.

Merging Small Batches

The mergeSmallBatches function pools batches marked as mergeable that contain fewer than MIN_BATCH_SIZE (3 files). These tiny batches aggregate into "misc" batches containing up to MAX_MERGE_TARGET (25 files), reducing dispatch overhead while maintaining the semantic integrity of the larger, well-defined groups.

Building the Neighbor Map

To enable cross-batch analysis, the system constructs a neighborMap for each file (lines 84-124). This structure records:

  • batchImportData: Direct imports remaining within the same batch.
  • Cross-batch neighbours: Files imported from other batches, including the target batch index and exported symbols.

The neighbour list respects MAX_NEIGHBORS (50 entries), with excess connections ranked by degree (inbound + outbound connection count) to retain the most significant cross-module dependencies.

Implementation Examples


# Run the batch computation on a checked-out project

node ./understand-anything-plugin/skills/understand/compute-batches.mjs /path/to/project
// Minimal programmatic usage of the Louvain algorithm
import { runLouvain } from './compute-batches.mjs';

const communityMap = runLouvain(codeFiles, importMap);
// communityMap is a Map<string, string> mapping file paths to community IDs
// Merging tiny batches after initial grouping
import { mergeSmallBatches } from './compute-batches.mjs';

const merged = mergeSmallBatches(bareBatches);
// Returns optimized batches with consolidated small groups

Output Format and Consumption

The final output, written to batches.json, contains:

  • algorithm: Either louvain or count-fallback indicating the grouping method used.
  • exportsByPath: The symbol tables extracted during preprocessing.
  • batches: An array of objects with batchIndex, files, batchImportData, and neighborMap.

This format is consumed by the file-analyzer agents, which can now leverage semantic context rather than processing arbitrary file chunks. The companion script [merge-batch-graphs.py](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/merge-batch-graphs.py) further aggregates these per-batch graphs into global analysis structures.

Summary

  • Louvain clustering on import graphs creates initial semantic batches that keep coupled code together.
  • Tree-Sitter parsing extracts export symbols to enable precise cross-batch dependency tracking.
  • Domain-specific rules group non-code files (Dockerfiles, CI configs, SQL migrations) by logical function and location.
  • Size enforcement splits oversized communities (>35 files) and merges tiny batches (<3 files) to optimize analyzer throughput.
  • Neighbor maps cap cross-batch references at 50 entries, prioritized by connection degree, ensuring analyzers receive relevant context without noise.

Frequently Asked Questions

What algorithm does the batch computation use for semantic grouping?

The system primarily uses the Louvain community detection algorithm via the graphology library to identify clusters of files with dense import interconnections. When this fails, it falls back to a deterministic count-based assignment of 12 files per batch.

How does the system handle non-code files like Dockerfiles and CI configurations?

Non-code files bypass the import graph and enter predefined semantic groups in buildNonCodeBatches. Dockerfiles cluster by directory, GitHub workflows group together, SQL migrations organize by folder, and remaining files collect into parent-directory buckets. Groups A through D are protected from merging, while Group E participates in the small-batch consolidation.

What happens if the Louvain algorithm fails to run?

If the Louvain algorithm throws—typically due to missing native bindings—the runLouvain function catches the error and the pipeline invokes countBasedAssignment. This fallback creates fixed-size batches of 12 files using alphabetical ordering, ensuring the analysis pipeline continues without graph-theoretic clustering.

How are cross-batch dependencies tracked in the semantic grouping?

The system constructs a neighborMap during the finalization phase. For each file importing from another batch, the map records the neighbour's path, batch index, and exported symbols consumed. This map is capped at MAX_NEIGHBORS (50) and sorted by degree centrality, providing downstream analyzers with the most relevant cross-module context.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →