How to Detect Duplicate Code with Code-Graph-RAG: A Complete Technical Guide
Use cgr duplicates to find exact and similar code clones by analyzing structural AST fingerprints stored in Memgraph, with configurable similarity thresholds and filtering options.
Code-Graph-RAG (cgr) provides a purpose-built duplicate code detection engine that operates on the code's structural graph representation. This guide explains how the detection algorithm works, how to configure it, and how to run it from both the CLI and Python according to the vitali87/code-graph-rag source code.
How Duplicate Detection Works in Code-Graph-RAG
The duplicate detection pipeline runs in six distinct stages when you execute cgr duplicates. Understanding these stages helps you tune the tool for your codebase.
Stage 1: Fetching Symbol Fingerprints from Memgraph
The engine begins by querying the graph database for every symbol's structural signature. In codebase_rag/duplicates.py, the GraphQueryClient.fetch_all method executes CYPHER_DUPLICATE_FINGERPRINTS to retrieve:
- The whole-AST fingerprint (complete structural hash)
- Branch-level fingerprints (subtree hashes for similarity comparison)
- Node count (for size-based filtering)
# From codebase_rag/duplicates.py lines 78-84
# The Cypher query aggregates fingerprints per symbol
rows = ingestor.gq.fetch_all(
CYPHER_DUPLICATE_FINGERPRINTS,
project=project_name
)
This query targets the Memgraph instance you've populated with cgr ingest, so detection quality depends on complete and recent ingestion.
Stage 2: Building Normalized Entries
Raw database rows transform into _Entry objects in codebase_rag/duplicates.py lines 18-33. A critical deduplication step occurs here: source spans are deduplicated so that the same source location appearing multiple times (common in C++ with declarations and definitions) counts only once.
Each _Entry represents one distinct whole-skeleton fingerprint, carrying:
- The fingerprint hash
- Member symbols sharing that fingerprint
- Source location metadata
Stage 3: Filtering by Size and Pattern
Before analysis, the engine applies two filters defined in lines 37-43:
| Filter | Purpose | Configuration |
|---|---|---|
min_nodes |
Exclude trivial code fragments | --min-size CLI flag |
exclude_patterns |
Skip generated/test files | --exclude CLI flag (glob patterns) |
Symbols lacking fingerprints (unsupported languages or parse failures) are discarded at this stage.
Stage 4: Exact Duplicate Detection
When multiple members share an identical whole-AST fingerprint, the engine emits an exact duplicate group (KIND_EXACT with similarity = 1.0). This is the fastest path—pure hash comparison with no approximate calculation.
# Pseudocode from codebase_rag/duplicates.py lines 73-81
if len(entry.members) > 1:
groups.append(DuplicateGroup(
kind=KIND_EXACT,
similarity=1.0,
members=entry.members
))
Stage 5: Similar Duplicate Detection (Approximate Clones)
For non-exact matches, Code-Graph-RAG uses Jaccard similarity over branch fingerprints with three sophisticated sub-stages:
Candidate Generation: A prefix-filtering All-Pairs/PPJoin algorithm (lines 86-100) dramatically reduces comparison count while guaranteeing no qualifying pairs are missed. This scales to large codebases where brute-force O(n²) comparison would be prohibitive.
Edge Qualification: Each candidate pair undergoes two checks (lines 83-92):
- Jaccard score meets the
thresholdparameter - Pair is not a "nested-only" relationship (eliminates false positives where one function simply calls another)
Clique Enumeration: The similarity graph's maximal cliques are found using a pivoted Bron-Kerbosch algorithm (lines 116-124). A configurable max_similar_groups cap prevents exponential runtime on pathological inputs.
Stage 6: Result Packaging
The final DuplicatesReport (lines 98-103) contains:
groups: List ofDuplicateGroupobjects (exact and similar)skipped_symbols: Count of symbols ignored (unsupported language or bodiless)truncated: Boolean flag if similar-group enumeration hit capsanalyzed_symbols: Total symbols examined
CLI Usage: Running Duplicate Detection
The duplicates subcommand in codebase_rag/cli.py (lines 55-71 and 155-165) exposes the full engine through a convenient interface.
Basic Detection
# Detect duplicates in current directory (after cgr ingest)
cgr duplicates
Common Configuration Options
# Require minimum 25 AST nodes (filters out small utilities)
cgr duplicates --min-size 25
# Exclude generated files and tests
cgr duplicates --exclude "*gen_*" --exclude "*_test.go"
# Lower similarity threshold for near-miss detection (default typically 0.8)
cgr duplicates --threshold 0.6
# JSON output for CI integration
cgr duplicates --format json
# Fail exit code if any duplicates found (quality gate)
cgr duplicates --fail-on-found
# Limit similar group enumeration (performance safety)
cgr duplicates --max-similar-groups 1000
Output formats include Rich tables (human-readable) and JSON (machine-parseable), selected via --format.
Programmatic Usage in Python
For custom integrations, import directly from codebase_rag/duplicates.py.
from codebase_rag.duplicates import (
default_duplicates_config,
collect_duplicates,
)
from codebase_rag.services.graph_service import MemgraphIngestor
# Initialize connection to your Memgraph instance
ingestor = MemgraphIngestor()
# Build configuration: 30+ nodes, Jaccard ≥ 0.6
config = default_duplicates_config(
threshold=0.6,
min_nodes=30,
exclude_patterns=["*gen_*", "*vendor/*"]
)
# Execute detection for project "my_project"
duplicate_groups = collect_duplicates(ingestor, "my_project", config)
# Process results
for group in duplicate_groups:
print(f"\nGroup type: {group['kind']}, similarity: {group['similarity']:.2f}")
for member in group["members"]:
loc = f"{member['path']}:{member['start_line']}-{member['end_line']}"
print(f" → {loc} {member['qualified_name']}")
The default_duplicates_config factory centralizes default values defined in codebase_rag/constants.py, ensuring consistent behavior between CLI and programmatic use.
Key Configuration Parameters
| Parameter | Default | Effect |
|---|---|---|
threshold |
~0.8 | Minimum Jaccard similarity for "similar" duplicates |
min_nodes |
10 | Minimum AST node count to analyze |
max_similar_groups |
10000 | Hard cap on clique enumeration (prevents runaway) |
exclude_patterns |
[] |
Glob patterns for paths to ignore |
Adjust threshold downward to catch heavily modified clones; increase min_nodes to focus on substantial duplication rather than boilerplate.
Architecture and Source Files
The duplicate detection system spans these components in the vitali87/code-graph-rag repository:
codebase_rag/duplicates.py— Core engine with_Entry,collect_duplicates_with_coverage, exact/similar group logic, and Bron-Kerbosch implementationcodebase_rag/cli.py— CLI argument parsing,DuplicatesConfigassembly, and output formatting (Rich/JSON)codebase_rag/cypher_queries.py—CYPHER_DUPLICATE_FINGERPRINTSandCYPHER_DUPLICATE_SKIPPED_COUNTfor Memgraph accesscodebase_rag/constants.py— Default thresholds and limit constantscodebase_rag/tools/duplicate_detection.py— Async wrapper for MCP toolchain integration
Summary
- Run
cgr duplicatesaftercgr ingestto detect structural code clones stored in Memgraph - Exact duplicates match whole-AST fingerprints identically; similar duplicates use Jaccard similarity over branch fingerprints with PPJoin optimization
- Configure filtering via
--min-size,--exclude, and--thresholdto balance precision and recall - Access programmatically through
collect_duplicates()withdefault_duplicates_config()for custom pipelines - Monitor truncation via the
truncatedflag in results whenmax_similar_groupslimits are hit
Frequently Asked Questions
What does Code-Graph-RAG consider a "duplicate"?
Code-Graph-RAG identifies two clone types: exact duplicates (identical AST structure, similarity = 1.0) and similar duplicates (structurally related code with Jaccard similarity above your threshold). The tool uses tree-sitter generated AST fingerprints rather than text comparison, so renaming variables or reformatting won't hide clones.
Can I use duplicate detection without the full RAG pipeline?
No—you must first run cgr ingest to populate Memgraph with structural fingerprints. The cgr duplicates command queries these precomputed fingerprints; it does not parse source files directly. This design enables fast repeated detection runs on large codebases.
Why are some symbols skipped during detection?
Symbols are skipped in three cases: the language lacks tree-sitter grammar support, parsing failed (bodiless or malformed), or the symbol falls below min_nodes. Check skipped_symbols in your DuplicatesReport and analyzed_symbols for coverage metrics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →