How to Handle Large Graph Files with Git LFS in Understand-Anything

Enable Git LFS in Understand-Anything by running git lfs track ".understand-anything/*.json" to offload large knowledge graph JSON files (≥10 MB) to external storage while keeping lightweight pointers in your repository.

The Understand-Anything tool generates knowledge graphs that represent your codebase structure as JSON files in the .understand-anything/ directory. When these graphs grow beyond a few megabytes, they trigger Git's size limitations and create bloated repository histories that slow down clones. Using Git Large File Storage (Git LFS) allows you to store the actual graph contents outside the regular Git object database while maintaining seamless access to the data for the Understand-Anything analysis engine.

Why Knowledge Graphs Exceed Git's Limits

The .understand-anything/ directory stores knowledge-graph.json files that map your project's code structure. As repositories scale, these JSON files often exceed 10 MB—particularly when processing large codebases with thousands of nodes. According to the README.md "Large graphs" section, the project explicitly recommends Git LFS for these assets to prevent slow clones and excessive storage costs in standard Git history.

Setting Up Git LFS for Understand-Anything

Follow these steps to configure Git LFS for the Understand-Anything repository. This ensures that any future large graph files are automatically managed by LFS rather than stored directly in Git's object store.

Install Git LFS

Run the installation command once per machine to set up the Git LFS extension:

git lfs install

This command configures the necessary Git hooks and filters required for LFS operations on your system.

Configure Tracking Rules

Navigate to your repository root and tell Git LFS to track all JSON files within the hidden graph directory:

git lfs track ".understand-anything/*.json"

This command generates a .gitattributes file containing the pattern *.json filter=lfs diff=lfs merge=lfs -text, ensuring that every JSON file in .understand-anything/ is treated as a large file rather than a standard Git object.

Commit and Push

Stage the generated .gitattributes file along with your existing graph files:

git add .gitattributes .understand-anything/
git commit -m "Enable LFS for large knowledge-graph JSON files"
git push

When you push, the large files are uploaded to the LFS server while only lightweight pointer files (typically 130 bytes) remain in the Git repository, keeping your history slim.

Generating Test Data

To verify your LFS configuration or benchmark performance, the repository includes a helper script that generates synthetic graphs. The scripts/generate-large-graph.mjs script creates a mock knowledge graph with a specified number of nodes:

node scripts/generate-large-graph.mjs 3000

This generates approximately 10 MB of JSON data in .understand-anything/knowledge-graph.json. On your next commit, Git automatically recognizes this file as an LFS object based on your .gitattributes configuration, allowing you to test the workflow before committing production data.

Verifying LFS Integration

Confirm that your graph files are properly tracked by running:

git lfs ls-files

You should see entries for .understand-anything/knowledge-graph.json indicating they are stored in LFS. If you encounter issues, ensure the .gitattributes file is checked into the repository root and that git lfs install was run before the tracking commands were executed.

Summary

  • Store Understand-Anything knowledge graphs in Git LFS by tracking the .understand-anything/*.json pattern.
  • Run git lfs install once per machine, then git lfs track ".understand-anything/*.json" to generate the required .gitattributes configuration.
  • Commit .gitattributes alongside your graph files to ensure consistent LFS behavior across all team members' clones.
  • Use scripts/generate-large-graph.mjs with 3000 nodes (≈10 MB) to validate that large files are captured by LFS before pushing production data.

Frequently Asked Questions

What is the file size threshold for using Git LFS with Understand-Anything?

While Git technically allows files up to 100 MB, the Understand-Anything documentation recommends Git LFS for graphs around 10 MB or larger. At this size, JSON knowledge graphs begin impacting clone performance and repository efficiency, making LFS the optimal solution for teams collaborating on large codebases.

Where does Understand-Anything store its knowledge graph files?

The tool stores graph files in a hidden directory named .understand-anything/ at your repository root, specifically files like knowledge-graph.json. You must configure Git LFS to track this specific path pattern rather than all JSON files in your project to avoid inadvertently LFS-tracking configuration files or small metadata.

Can I use the generate-large-graph script to test my LFS setup?

Yes. Run node scripts/generate-large-graph.mjs 3000 from the repository root to create a synthetic graph with 3,000 nodes (approximately 10 MB). This generates a .understand-anything/knowledge-graph.json file that triggers LFS tracking rules, allowing you to verify that git lfs track properly captures large files before committing production data.

Do other contributors need to install Git LFS to clone the repository?

Yes. Anyone cloning the repository needs Git LFS installed (git lfs install) to download the actual graph file contents. Without it, they will only receive the 130-byte pointer files and will see placeholder text instead of the actual JSON data when opening the knowledge graph files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →