# How to Detect Code Duplication with Code-Graph-RAG: A Complete Guide to Finding Exact and Similar Clones

> Detect code duplication with Code-Graph-RAG. Learn to find exact and similar code clones using AST and branch fingerprint analysis. A complete guide for developers.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-04

---

**Code-Graph-RAG detects code duplication by building a fingerprint graph during ingestion and running a two-stage engine that identifies exact copies via AST fingerprints and similar clones through Jaccard similarity analysis on branch fingerprints.**

Code-Graph-RAG is a multilanguage codebase analysis tool that uses Tree-sitter parsing to detect structural code clones. The system creates **AST fingerprints** during the ingestion phase, then analyzes these fingerprints to find both verbatim copies and edited variants across your entire repository.

## Understanding the Two-Stage Duplicate Detection Engine

The duplicate detection system in [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) operates through two complementary stages. Each stage targets a different type of code duplication, from identical function copies to logic that has been modified but remains structurally similar.

### Stage 1: Detecting Exact and Renamed Copies

The first stage identifies functions with identical *whole-skeleton* fingerprints. During ingestion, every function node receives an `AST_FINGERPRINT` value generated by the Tree-sitter parser.

In [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py), the `_entries_from_rows` function collects functions sharing the same fingerprint into a single `_Entry`. The `_exact_groups` method then emits these collections as exact duplicate groups. This catches direct copy-paste operations and renamed variables, since the underlying AST structure remains unchanged.

This stage runs quickly because it relies on exact hash matching rather than pairwise comparison.

### Stage 2: Detecting Similar and Edited Copies

The second stage finds code that has been edited but retains high structural overlap. This process uses **branch fingerprints** (statement-level hashes) and involves several sophisticated algorithms:

1. **Prefix Indexing (PPJoin)**: The `_prefix_index` method builds an index only on the rarest branch fingerprints needed to satisfy the Jaccard threshold, dramatically reducing the candidate search space.

2. **Candidate Generation**: The `_candidate_pairs` function generates potential duplicate pairs from this prefix index.

3. **Threshold Filtering**: The `_threshold_adjacency` method filters pairs by Jaccard similarity and excludes pure nesting relationships.

4. **Clique Enumeration**: The `_maximal_cliques` function implements the **Bron-Kerbosch** algorithm to find maximal cliques in the similarity graph, where each clique represents a group of mutually similar functions.

The `_similar_groups` function orchestrates this entire pipeline, emitting groups with similarity scores and node counts.

## Configuration and Deduplication Logic

### Handling Nested Functions and Containers

To avoid reporting closures as duplicates of their containing functions, Code-Graph-RAG implements sophisticated containment detection. The system uses line-range checks via `_span_contains` and qualified-name hierarchy analysis through `_qn_within`.

The `_drop_contained_members` function removes nested members from duplicate groups, ensuring that inner functions are not flagged as duplicates of their outer scope. This prevents false positives when analyzing languages with nested function support like JavaScript or Python.

### Configuring Detection Parameters

Detection behavior is controlled by the `DuplicatesConfig` dataclass defined in [`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py) (lines 735-749). Default values include:

- **Similarity threshold**: Minimum Jaccard similarity for "similar" duplicates
- **Minimum node count**: Minimum AST node size to consider a function
- **Exclusion patterns**: Glob patterns for files to ignore (e.g., generated code)
- **Maximum groups**: Cap on reported duplicate clusters

You can override these defaults via CLI flags or programmatic configuration when using the detection API.

## Running Duplicate Detection via CLI

The Typer-based CLI in [`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) exposes duplicate detection through the `cgr duplicates` command. The implementation uses `_duplicates_group_cell` and `_duplicates_location_cell` for formatted output.

Basic usage prints all duplicate groups for the current project:

```bash
cgr duplicates

```

Tune detection sensitivity with flags:

```bash

# Report only substantial functions (≥30 nodes) with 80% similarity

cgr duplicates --min-size 30 --threshold 0.8

```

Exclude generated or vendor files:

```bash
cgr duplicates --exclude-patterns "*gen_*" --exclude-patterns "*/vendor/*"

```

## Programmatic and Agentic Access

Beyond the CLI, the detection engine is exposed as an MCP (Model Context Protocol) tool for AI-driven analysis. The [`codebase_rag/tools/duplicate_detection.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py) module provides `create_find_duplicates_tool`, which wraps the engine for async usage.

Use the tool programmatically:

```python
from codebase_rag.tools.duplicate_detection import create_find_duplicates_tool

tool = create_find_duplicates_tool(ingestor)  # ingestor is a GraphQueryClient

result = await tool.function(
    project="myproj",
    threshold=0.75,
    min_size=20,
    limit=10
)
print(result)  # Human-readable duplicate report

```

This enables AI agents to call `find_duplicate_code` during automated refactoring workflows, receiving structured data about duplication clusters across the codebase.

## Summary

- **Two-stage detection**: Code-Graph-RAG first finds exact copies via `AST_FINGERPRINT` matching, then discovers similar clones using Jaccard similarity on branch fingerprints.
- **Algorithmic sophistication**: The system uses PPJoin indexing for performance and Bron-Kerbosch clique enumeration for accurate grouping of similar functions.
- **Smart filtering**: Nested functions are automatically excluded from parent function duplicate groups via `_span_contains` and `_qn_within` checks.
- **Multiple interfaces**: Access detection via `cgr duplicates` CLI, direct Python API, or MCP agent tools in [`codebase_rag/tools/duplicate_detection.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py).
- **Configurable thresholds**: Adjust sensitivity through `DuplicatesConfig` in [`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py) or CLI flags like `--threshold` and `--min-size`.

## Frequently Asked Questions

### How does Code-Graph-RAG differentiate between exact copies and similar clones?

Exact copies share identical `AST_FINGERPRINT` values generated by Tree-sitter parsing, detected through hash lookup in `_exact_groups`. Similar clones are detected via Jaccard similarity of branch fingerprints in `_similar_groups`, requiring configurable threshold matching (default typically 0.7-0.8) and Bron-Kerbosch clique analysis to find related function clusters.

### What algorithms does Code-Graph-RAG use for scalable duplicate detection?

The system implements **PPJoin (Position-Prefix Join)** indexing in `_prefix_index` to reduce candidate pairs, followed by **Bron-Kerbosch** maximal clique enumeration in `_maximal_cliques` to identify duplicate groups. This combination avoids the O(n²) complexity of naive pairwise comparison, making it feasible for large codebases.

### Can Code-Graph-RAG detect code duplication across different programming languages?

Yes, because the detection relies on **AST fingerprints** rather than text comparison. Since Tree-sitter parses each language into a normalized AST structure, functions in Python and JavaScript could potentially be flagged as similar if their control flow structures match, though the tool primarily focuses on finding duplicates within the same language family.

### How do I prevent test files or generated code from appearing in duplicate detection results?

Use the `--exclude-patterns` CLI flag or configure exclusion globs in `DuplicatesConfig`. For example, run `cgr duplicates --exclude-patterns "*_test.py" --exclude-patterns "*/generated/*"` to skip test suites and auto-generated files. These patterns are evaluated during the candidate generation phase in [`duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/duplicates.py).