# How to Detect Duplicate Code with Code-Graph-RAG: A Complete Technical Guide

> Detect duplicate code with Code-Graph-RAG using cgr duplicates. Analyze AST fingerprints in Memgraph to find exact and similar code clones with custom thresholds.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Use `cgr duplicates` to find exact and similar code clones by analyzing structural AST fingerprints stored in Memgraph, with configurable similarity thresholds and filtering options.**

Code-Graph-RAG (cgr) provides a purpose-built duplicate code detection engine that operates on the code's structural graph representation. This guide explains how the detection algorithm works, how to configure it, and how to run it from both the CLI and Python according to the vitali87/code-graph-rag source code.

---

## How Duplicate Detection Works in Code-Graph-RAG

The duplicate detection pipeline runs in six distinct stages when you execute `cgr duplicates`. Understanding these stages helps you tune the tool for your codebase.

### Stage 1: Fetching Symbol Fingerprints from Memgraph

The engine begins by querying the graph database for every symbol's structural signature. In [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py), the `GraphQueryClient.fetch_all` method executes `CYPHER_DUPLICATE_FINGERPRINTS` to retrieve:

- The **whole-AST fingerprint** (complete structural hash)
- **Branch-level fingerprints** (subtree hashes for similarity comparison)
- **Node count** (for size-based filtering)

```python

# From codebase_rag/duplicates.py lines 78-84

# The Cypher query aggregates fingerprints per symbol

rows = ingestor.gq.fetch_all(
    CYPHER_DUPLICATE_FINGERPRINTS,
    project=project_name
)

```

This query targets the Memgraph instance you've populated with `cgr ingest`, so detection quality depends on complete and recent ingestion.

### Stage 2: Building Normalized Entries

Raw database rows transform into `_Entry` objects in [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) lines 18-33. A critical deduplication step occurs here: **source spans are deduplicated** so that the same source location appearing multiple times (common in C++ with declarations and definitions) counts only once.

Each `_Entry` represents one distinct whole-skeleton fingerprint, carrying:
- The fingerprint hash
- Member symbols sharing that fingerprint
- Source location metadata

### Stage 3: Filtering by Size and Pattern

Before analysis, the engine applies two filters defined in lines 37-43:

| Filter | Purpose | Configuration |
|--------|---------|---------------|
| `min_nodes` | Exclude trivial code fragments | `--min-size` CLI flag |
| `exclude_patterns` | Skip generated/test files | `--exclude` CLI flag (glob patterns) |

Symbols lacking fingerprints (unsupported languages or parse failures) are discarded at this stage.

### Stage 4: Exact Duplicate Detection

When multiple members share an identical whole-AST fingerprint, the engine emits an **exact duplicate group** (`KIND_EXACT` with similarity = 1.0). This is the fastest path—pure hash comparison with no approximate calculation.

```python

# Pseudocode from codebase_rag/duplicates.py lines 73-81

if len(entry.members) > 1:
    groups.append(DuplicateGroup(
        kind=KIND_EXACT,
        similarity=1.0,
        members=entry.members
    ))

```

### Stage 5: Similar Duplicate Detection (Approximate Clones)

For non-exact matches, Code-Graph-RAG uses **Jaccard similarity over branch fingerprints** with three sophisticated sub-stages:

**Candidate Generation:** A prefix-filtering All-Pairs/PPJoin algorithm (lines 86-100) dramatically reduces comparison count while guaranteeing no qualifying pairs are missed. This scales to large codebases where brute-force O(n²) comparison would be prohibitive.

**Edge Qualification:** Each candidate pair undergoes two checks (lines 83-92):
- Jaccard score meets the `threshold` parameter
- Pair is not a "nested-only" relationship (eliminates false positives where one function simply calls another)

**Clique Enumeration:** The similarity graph's maximal cliques are found using a **pivoted Bron-Kerbosch algorithm** (lines 116-124). A configurable `max_similar_groups` cap prevents exponential runtime on pathological inputs.

### Stage 6: Result Packaging

The final `DuplicatesReport` (lines 98-103) contains:

- `groups`: List of `DuplicateGroup` objects (exact and similar)
- `skipped_symbols`: Count of symbols ignored (unsupported language or bodiless)
- `truncated`: Boolean flag if similar-group enumeration hit caps
- `analyzed_symbols`: Total symbols examined

---

## CLI Usage: Running Duplicate Detection

The `duplicates` subcommand in [`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) (lines 55-71 and 155-165) exposes the full engine through a convenient interface.

### Basic Detection

```bash

# Detect duplicates in current directory (after cgr ingest)

cgr duplicates

```

### Common Configuration Options

```bash

# Require minimum 25 AST nodes (filters out small utilities)

cgr duplicates --min-size 25

# Exclude generated files and tests

cgr duplicates --exclude "*gen_*" --exclude "*_test.go"

# Lower similarity threshold for near-miss detection (default typically 0.8)

cgr duplicates --threshold 0.6

# JSON output for CI integration

cgr duplicates --format json

# Fail exit code if any duplicates found (quality gate)

cgr duplicates --fail-on-found

# Limit similar group enumeration (performance safety)

cgr duplicates --max-similar-groups 1000

```

Output formats include **Rich tables** (human-readable) and **JSON** (machine-parseable), selected via `--format`.

---

## Programmatic Usage in Python

For custom integrations, import directly from [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py).

```python
from codebase_rag.duplicates import (
    default_duplicates_config,
    collect_duplicates,
)
from codebase_rag.services.graph_service import MemgraphIngestor

# Initialize connection to your Memgraph instance

ingestor = MemgraphIngestor()

# Build configuration: 30+ nodes, Jaccard ≥ 0.6

config = default_duplicates_config(
    threshold=0.6,
    min_nodes=30,
    exclude_patterns=["*gen_*", "*vendor/*"]
)

# Execute detection for project "my_project"

duplicate_groups = collect_duplicates(ingestor, "my_project", config)

# Process results

for group in duplicate_groups:
    print(f"\nGroup type: {group['kind']}, similarity: {group['similarity']:.2f}")
    for member in group["members"]:
        loc = f"{member['path']}:{member['start_line']}-{member['end_line']}"
        print(f"  → {loc}  {member['qualified_name']}")

```

The `default_duplicates_config` factory centralizes default values defined in [`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py), ensuring consistent behavior between CLI and programmatic use.

---

## Key Configuration Parameters

| Parameter | Default | Effect |
|-----------|---------|--------|
| `threshold` | ~0.8 | Minimum Jaccard similarity for "similar" duplicates |
| `min_nodes` | 10 | Minimum AST node count to analyze |
| `max_similar_groups` | 10000 | Hard cap on clique enumeration (prevents runaway) |
| `exclude_patterns` | `[]` | Glob patterns for paths to ignore |

Adjust `threshold` downward to catch heavily modified clones; increase `min_nodes` to focus on substantial duplication rather than boilerplate.

---

## Architecture and Source Files

The duplicate detection system spans these components in the vitali87/code-graph-rag repository:

- **[`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py)** — Core engine with `_Entry`, `collect_duplicates_with_coverage`, exact/similar group logic, and Bron-Kerbosch implementation
- **[`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)** — CLI argument parsing, `DuplicatesConfig` assembly, and output formatting (Rich/JSON)
- **[`codebase_rag/cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py)** — `CYPHER_DUPLICATE_FINGERPRINTS` and `CYPHER_DUPLICATE_SKIPPED_COUNT` for Memgraph access
- **[`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py)** — Default thresholds and limit constants
- **[`codebase_rag/tools/duplicate_detection.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py)** — Async wrapper for MCP toolchain integration

---

## Summary

- **Run `cgr duplicates`** after `cgr ingest` to detect structural code clones stored in Memgraph
- **Exact duplicates** match whole-AST fingerprints identically; **similar duplicates** use Jaccard similarity over branch fingerprints with PPJoin optimization
- **Configure filtering** via `--min-size`, `--exclude`, and `--threshold` to balance precision and recall
- **Access programmatically** through `collect_duplicates()` with `default_duplicates_config()` for custom pipelines
- **Monitor truncation** via the `truncated` flag in results when `max_similar_groups` limits are hit

---

## Frequently Asked Questions

### What does Code-Graph-RAG consider a "duplicate"?

Code-Graph-RAG identifies two clone types: **exact duplicates** (identical AST structure, similarity = 1.0) and **similar duplicates** (structurally related code with Jaccard similarity above your threshold). The tool uses tree-sitter generated AST fingerprints rather than text comparison, so renaming variables or reformatting won't hide clones.

### Can I use duplicate detection without the full RAG pipeline?

No—you must first run `cgr ingest` to populate Memgraph with structural fingerprints. The `cgr duplicates` command queries these precomputed fingerprints; it does not parse source files directly. This design enables fast repeated detection runs on large codebases.

### Why are some symbols skipped during detection?

Symbols are skipped in three cases: the language lacks tree-sitter grammar support, parsing failed (bodiless or malformed), or the symbol falls below `min_nodes`. Check `skipped_symbols` in your `DuplicatesReport` and `analyzed_symbols` for coverage metrics.