How to Detect Duplicate Code with Code-Graph-RAG: A Complete Technical Guide

Use cgr duplicates to find exact and similar code clones by analyzing structural AST fingerprints stored in Memgraph, with configurable similarity thresholds and filtering options.

Code-Graph-RAG (cgr) provides a purpose-built duplicate code detection engine that operates on the code's structural graph representation. This guide explains how the detection algorithm works, how to configure it, and how to run it from both the CLI and Python according to the vitali87/code-graph-rag source code.


How Duplicate Detection Works in Code-Graph-RAG

The duplicate detection pipeline runs in six distinct stages when you execute cgr duplicates. Understanding these stages helps you tune the tool for your codebase.

Stage 1: Fetching Symbol Fingerprints from Memgraph

The engine begins by querying the graph database for every symbol's structural signature. In codebase_rag/duplicates.py, the GraphQueryClient.fetch_all method executes CYPHER_DUPLICATE_FINGERPRINTS to retrieve:

  • The whole-AST fingerprint (complete structural hash)
  • Branch-level fingerprints (subtree hashes for similarity comparison)
  • Node count (for size-based filtering)

# From codebase_rag/duplicates.py lines 78-84

# The Cypher query aggregates fingerprints per symbol

rows = ingestor.gq.fetch_all(
    CYPHER_DUPLICATE_FINGERPRINTS,
    project=project_name
)

This query targets the Memgraph instance you've populated with cgr ingest, so detection quality depends on complete and recent ingestion.

Stage 2: Building Normalized Entries

Raw database rows transform into _Entry objects in codebase_rag/duplicates.py lines 18-33. A critical deduplication step occurs here: source spans are deduplicated so that the same source location appearing multiple times (common in C++ with declarations and definitions) counts only once.

Each _Entry represents one distinct whole-skeleton fingerprint, carrying:

  • The fingerprint hash
  • Member symbols sharing that fingerprint
  • Source location metadata

Stage 3: Filtering by Size and Pattern

Before analysis, the engine applies two filters defined in lines 37-43:

Filter Purpose Configuration
min_nodes Exclude trivial code fragments --min-size CLI flag
exclude_patterns Skip generated/test files --exclude CLI flag (glob patterns)

Symbols lacking fingerprints (unsupported languages or parse failures) are discarded at this stage.

Stage 4: Exact Duplicate Detection

When multiple members share an identical whole-AST fingerprint, the engine emits an exact duplicate group (KIND_EXACT with similarity = 1.0). This is the fastest path—pure hash comparison with no approximate calculation.


# Pseudocode from codebase_rag/duplicates.py lines 73-81

if len(entry.members) > 1:
    groups.append(DuplicateGroup(
        kind=KIND_EXACT,
        similarity=1.0,
        members=entry.members
    ))

Stage 5: Similar Duplicate Detection (Approximate Clones)

For non-exact matches, Code-Graph-RAG uses Jaccard similarity over branch fingerprints with three sophisticated sub-stages:

Candidate Generation: A prefix-filtering All-Pairs/PPJoin algorithm (lines 86-100) dramatically reduces comparison count while guaranteeing no qualifying pairs are missed. This scales to large codebases where brute-force O(n²) comparison would be prohibitive.

Edge Qualification: Each candidate pair undergoes two checks (lines 83-92):

  • Jaccard score meets the threshold parameter
  • Pair is not a "nested-only" relationship (eliminates false positives where one function simply calls another)

Clique Enumeration: The similarity graph's maximal cliques are found using a pivoted Bron-Kerbosch algorithm (lines 116-124). A configurable max_similar_groups cap prevents exponential runtime on pathological inputs.

Stage 6: Result Packaging

The final DuplicatesReport (lines 98-103) contains:

  • groups: List of DuplicateGroup objects (exact and similar)
  • skipped_symbols: Count of symbols ignored (unsupported language or bodiless)
  • truncated: Boolean flag if similar-group enumeration hit caps
  • analyzed_symbols: Total symbols examined

CLI Usage: Running Duplicate Detection

The duplicates subcommand in codebase_rag/cli.py (lines 55-71 and 155-165) exposes the full engine through a convenient interface.

Basic Detection


# Detect duplicates in current directory (after cgr ingest)

cgr duplicates

Common Configuration Options


# Require minimum 25 AST nodes (filters out small utilities)

cgr duplicates --min-size 25

# Exclude generated files and tests

cgr duplicates --exclude "*gen_*" --exclude "*_test.go"

# Lower similarity threshold for near-miss detection (default typically 0.8)

cgr duplicates --threshold 0.6

# JSON output for CI integration

cgr duplicates --format json

# Fail exit code if any duplicates found (quality gate)

cgr duplicates --fail-on-found

# Limit similar group enumeration (performance safety)

cgr duplicates --max-similar-groups 1000

Output formats include Rich tables (human-readable) and JSON (machine-parseable), selected via --format.


Programmatic Usage in Python

For custom integrations, import directly from codebase_rag/duplicates.py.

from codebase_rag.duplicates import (
    default_duplicates_config,
    collect_duplicates,
)
from codebase_rag.services.graph_service import MemgraphIngestor

# Initialize connection to your Memgraph instance

ingestor = MemgraphIngestor()

# Build configuration: 30+ nodes, Jaccard ≥ 0.6

config = default_duplicates_config(
    threshold=0.6,
    min_nodes=30,
    exclude_patterns=["*gen_*", "*vendor/*"]
)

# Execute detection for project "my_project"

duplicate_groups = collect_duplicates(ingestor, "my_project", config)

# Process results

for group in duplicate_groups:
    print(f"\nGroup type: {group['kind']}, similarity: {group['similarity']:.2f}")
    for member in group["members"]:
        loc = f"{member['path']}:{member['start_line']}-{member['end_line']}"
        print(f"  → {loc}  {member['qualified_name']}")

The default_duplicates_config factory centralizes default values defined in codebase_rag/constants.py, ensuring consistent behavior between CLI and programmatic use.


Key Configuration Parameters

Parameter Default Effect
threshold ~0.8 Minimum Jaccard similarity for "similar" duplicates
min_nodes 10 Minimum AST node count to analyze
max_similar_groups 10000 Hard cap on clique enumeration (prevents runaway)
exclude_patterns [] Glob patterns for paths to ignore

Adjust threshold downward to catch heavily modified clones; increase min_nodes to focus on substantial duplication rather than boilerplate.


Architecture and Source Files

The duplicate detection system spans these components in the vitali87/code-graph-rag repository:


Summary

  • Run cgr duplicates after cgr ingest to detect structural code clones stored in Memgraph
  • Exact duplicates match whole-AST fingerprints identically; similar duplicates use Jaccard similarity over branch fingerprints with PPJoin optimization
  • Configure filtering via --min-size, --exclude, and --threshold to balance precision and recall
  • Access programmatically through collect_duplicates() with default_duplicates_config() for custom pipelines
  • Monitor truncation via the truncated flag in results when max_similar_groups limits are hit

Frequently Asked Questions

What does Code-Graph-RAG consider a "duplicate"?

Code-Graph-RAG identifies two clone types: exact duplicates (identical AST structure, similarity = 1.0) and similar duplicates (structurally related code with Jaccard similarity above your threshold). The tool uses tree-sitter generated AST fingerprints rather than text comparison, so renaming variables or reformatting won't hide clones.

Can I use duplicate detection without the full RAG pipeline?

No—you must first run cgr ingest to populate Memgraph with structural fingerprints. The cgr duplicates command queries these precomputed fingerprints; it does not parse source files directly. This design enables fast repeated detection runs on large codebases.

Why are some symbols skipped during detection?

Symbols are skipped in three cases: the language lacks tree-sitter grammar support, parsing failed (bodiless or malformed), or the symbol falls below min_nodes. Check skipped_symbols in your DuplicatesReport and analyzed_symbols for coverage metrics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →