How Graphify's Semantic Caching Strategy Minimizes LLM API Calls

Graphify reduces LLM API usage by caching extracted graph nodes and edges keyed to file content hashes, checking this semantic cache before every LLM invocation, and only calling the API when source files change or caching is explicitly bypassed.

The Graphify open-source project (available at Graphify-Labs/graphify) implements an intelligent semantic caching strategy that eliminates redundant calls to large language model (LLM) APIs. By persisting the results of expensive semantic extraction operations—specifically the graph nodes and edges derived from source code—Graphify can reuse previous LLM outputs when processing unchanged files. This approach significantly reduces API costs, network latency, and rate-limit issues during iterative development workflows.

The Semantic Cache Architecture

Graphify's caching mechanism operates on a simple but effective principle: if a file's content has not changed, the semantic graph derived from it remains identical. The system implements this through several coordinated components.

Hash-Based Cache Keys

For every source file processed, Graphify computes a stable hash of the file's contents combined with its path. This hash serves as the primary key for the semantic cache. By using content-derived hashes rather than timestamps, Graphify ensures that even files moved or renamed without modification still match their cached entries. The cache data is stored in a JSON file (.graphify_cached.json) located in the specified output directory.

Pre-Flight Cache Lookup

Before invoking any LLM API, Graphify calls check_semantic_cache (defined in [graphify/cache.py at line 531](https://github.com/Graphify-Labs/graphify/blob/v8/graphify/cache.py#L531)). This function checks whether the computed hash exists in the cache and returns the previously extracted nodes, edges, and metadata. If a match is found, the LLM call is skipped entirely, and the cached graph fragments are loaded directly into the processing pipeline.

Implementation Deep Dive

The caching layer is tightly integrated into Graphify's extraction pipeline, ensuring minimal overhead and maximum reusability.

Core Cache Functions in cache.py

The [graphify/cache.py](https://github.com/Graphify-Labs/graphify/blob/v8/graphify/cache.py) module provides three primary functions that manage the cache lifecycle:

  • check_semantic_cache (line 531): Validates cache hits by comparing file hashes against stored entries. Returns cached nodes, edges, and provenance data when available.
  • save_semantic_cache (line 560): Persists newly extracted graph data to the cache file after a successful LLM invocation. This function serializes the nodes and edges alongside their content hashes.
  • prune_semantic_cache (line 493): Removes cache entries for files that no longer exist in the current corpus, preventing unbounded cache growth.

LLM Integration Layer

The LLM backend in [graphify/llm.py](https://github.com/Graphify-Labs/graphify/blob/v8/graphify/llm.py) integrates directly with the cache utilities. Around line 1916, the code checks the semantic cache before constructing LLM prompts. When check_semantic_cache returns a hit, the LLM code path short-circuits, returning the cached result rather than invoking the external API. This integration ensures that expensive LLM calls only occur when the cache misses or when the user explicitly forces a refresh.

Scoped Writes to Prevent Contamination

To avoid polluting the cache with stale data from unrelated files, Graphify implements scoped cache writes in [tools/skillgen/gen.py](https://github.com/Graphify-Labs/graphify/blob/v8/tools/skillgen/gen.py) (lines 905-917). This logic ensures that save_semantic_cache only writes entries for files that were actually dispatched for extraction during the current run. By scoping writes to specific file sets, Graphify prevents scenarios where partial runs or targeted extractions might overwrite valid cache entries for other unchanged files.

Cache Lifecycle Management

Maintaining cache integrity requires periodic cleanup of obsolete entries.

Pruning Obsolete Entries

After processing, Graphify calls prune_semantic_cache (line 493 in cache.py) to remove any cached items whose hashes no longer correspond to files in the current working directory. This garbage collection keeps the cache size bounded and ensures that only relevant, up-to-date semantic data persists across sessions. The pruning operation is idempotent and safe to run multiple times without affecting valid cache entries.

Practical Usage Examples

Graphify's semantic caching operates transparently during normal extraction workflows, but developers can also interact with it programmatically.

Automatic Cache Hits

When running Graphify repeatedly on the same codebase, the semantic cache automatically prevents redundant LLM calls:

from graphify import extract
from pathlib import Path

# First run: LLM processes each file and caches results

graph = extract(["src/*.py"], cache_root=Path("output"))

# Second run: Instant cache retrieval, zero LLM calls

graph = extract(["src/*.py"], cache_root=Path("output"))

# Behind the scenes: check_semantic_cache returns cached nodes for unchanged files

Bypassing the Cache

When source files change or you need fresh extractions, Graphify supports cache bypassing:


# Force re-extraction regardless of cache status

graph = extract(["src/*.py"], cache_root=Path("output"), force=True)

# Or manually check cache status before extraction

from graphify.cache import check_semantic_cache

cached_nodes, edges, metadata, hit = check_semantic_cache(
    ["src/main.py"], root=Path("output")
)
if hit:
    print("Cache hit: Loading from semantic cache")
else:
    print("Cache miss: Will invoke LLM API")

Summary

  • Graphify's semantic caching strategy uses content hashing to cache LLM extraction results, eliminating redundant API calls for unchanged files.
  • The check_semantic_cache function in graphify/cache.py (line 531) provides pre-flight lookups that short-circuit LLM invocations when valid cached data exists.
  • save_semantic_cache (line 560) persists new extractions immediately after LLM calls, while prune_semantic_cache (line 493) garbage-collects stale entries.
  • Scoped cache writes in tools/skillgen/gen.py (lines 905-917) prevent cross-contamination between extraction runs.
  • The integration layer in graphify/llm.py (around line 1916) ensures seamless cache checking before every LLM operation.

Frequently Asked Questions

How does Graphify determine if a file's semantic data is still valid?

Graphify computes a cryptographic hash of each file's content combined with its path. This hash is stored alongside the extracted graph data in the semantic cache. When processing a file, Graphify recalculates the hash and compares it against the cache. If the hashes match, the cached semantic data is considered valid and returned immediately from check_semantic_cache without invoking the LLM.

What happens when I modify a source file and run Graphify again?

When a file's content changes, its hash changes, causing a cache miss in check_semantic_cache. Graphify then invokes the LLM API to extract fresh nodes and edges, immediately persisting the new results via save_semantic_cache (line 560). The next run with the modified file will return these updated cached values, ensuring the semantic graph always reflects the current source code while still benefiting from caching for unchanged files.

Can I disable semantic caching when running Graphify?

Yes. The extraction API accepts a force parameter that bypasses the cache lookup entirely. When force=True, Graphify skips check_semantic_cache and proceeds directly to LLM invocation, then overwrites any existing cache entries for the processed files with the new results. This is useful when updating LLM models or when you suspect cached data may be corrupted.

Where does Graphify store the semantic cache data?

By default, Graphify writes semantic cache data to a file named .graphify_cached.json located in the cache_root directory (typically your specified output folder). This JSON file contains serialized nodes, edges, content hashes, and metadata. The prune_semantic_cache function manages this file's size by removing entries for deleted or modified files, ensuring the cache remains compact and relevant.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →