# How to Find Structurally Duplicated Code with Code-Graph-RAG

> Learn how to find structurally duplicated code using Code-Graph-RAG. This tool employs a two-stage pipeline for precise detection and offers both CLI and Python async interfaces for seamless integration.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Code-Graph-RAG detects structural code duplication through a two-stage pipeline—first grouping exact copies via whole-skeleton fingerprints, then detecting similar functions using Jaccard similarity of branch fingerprints and maximal clique extraction—exposed through both a command-line interface and a Python async tool.**

Finding structurally duplicated code is essential for maintaining clean, refactorable repositories. The **Code-Graph-RAG** project (`vitali87/code-graph-rag`) provides a sophisticated engine that analyzes Abstract Syntax Tree (AST) skeletons to detect both exact copies and near-duplicate functions. This article explains how to leverage the duplicate detection capabilities using the CLI and programmatic APIs.

## How Structural Duplicate Detection Works

The detection engine resides in [[`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) and operates in two distinct stages to balance precision and performance.

### Stage 1: Exact Copy Grouping

The engine first identifies functions that are identical except for identifier names by grouping entries sharing the same **whole-skeleton fingerprint**. This fingerprint is computed at ingest time and stored for every function definition.

The [`_exact_groups`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py#L73-L84) function builds these groups in linear time by aggregating members with identical fingerprints:

```python

# _exact_groups – creates a DuplicateGroup for each fingerprint with >1 members

# (see lines 73-84 of duplicates.py)

def _exact_groups(entries: dict[str, _Entry]) -> list[DuplicateGroup]:
    return [
        DuplicateGroup(
            kind=cs.KIND_EXACT,
            similarity=1.0,
            node_count=entry.node_count,
            members=_sorted_members(entry.members),
        )
        for entry in entries.values()
        if len(entry.members) > 1
    ]

```

This stage runs in **O(n)** time because it relies on a simple dictionary lookup of pre-computed hashes.

### Stage 2: Similar Copy Detection

When `exact_only` is disabled, the engine searches for **similar** functions that may have undergone minor edits. This stage uses a **prefix-filtered AllPairs/PPJoin** algorithm to avoid an expensive O(n²) comparison.

1. **Candidate Generation**: The [`_candidate_pairs`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py#L86-L100) function indexes only the rarest branch fingerprints that must appear in any qualifying pair, guaranteeing no valid matches are missed while pruning the search space.

2. **Edge Qualification**: For each candidate pair, the [`_edge_qualifies`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py#L94) function calculates **Jaccard similarity** between the sets of branch-level fingerprints. It also filters pairs that represent merely nested definitions using [`_only_nested_members`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py).

3. **Clique Extraction**: The engine builds a similarity graph where nodes are functions and edges represent qualified pairs. It then extracts **maximal cliques** using a pivoted Bron-Kerbosch algorithm in [`_maximal_cliques`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py#L76-L90). Each clique represents a group of functions where every member meets the similarity threshold with every other member.

The default similarity threshold is **0.8** and the minimum node count is **5**, both configurable via [`default_duplicates_config`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py).

## Finding Duplicates via the CLI

The [`cgr duplicates`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py#L1938-L1945) command in [[`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) provides immediate access to the detection engine:

```bash

# Scan the default project with default settings (threshold 0.8, min-size 5)

cgr duplicates

# Scan a specific project with custom parameters

cgr duplicates --project my_project --threshold 0.9 --min-size 10 --limit 5

# Output as JSON for further processing

cgr duplicates --project my_project --format json

```

The CLI invokes `collect_duplicates_with_coverage`, formats the resulting `DuplicatesReport`, and prints either a human-readable table or structured JSON depending on the `--format` flag.

## Detecting Duplicates Programmatically

For integration into async applications or agentic workflows, use the tool wrapper in [[`codebase_rag/tools/duplicate_detection.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py):

```python
from codebase_rag.services import GraphQueryClient
from codebase_rag.tools.duplicate_detection import create_find_duplicates_tool

# Initialize your GraphQL query client (must implement QueryProtocol)

client = GraphQueryClient()

# Create the tool instance

tool = create_find_duplicates_tool(client)

# Execute detection asynchronously

report = await tool.find_duplicate_code(
    project="my_project",
    threshold=0.85,
    min_size=8,
    limit=10,
)

print(report)  # Human-readable summary of duplicate groups

```

The tool validates arguments, resolves the target project, executes [`collect_duplicates`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) in a background thread to prevent blocking, and returns a formatted text report.

For synchronous, low-level access, call the engine directly:

```python
from codebase_rag.duplicates import collect_duplicates, default_duplicates_config

config = default_duplicates_config(threshold=0.75, min_nodes=6)
groups = collect_duplicates(
    ingestor=client, 
    project_name="my_project", 
    config=config
)

for group in groups:
    print(f"Group ({group['kind']}) – similarity {group['similarity']}")
    for member in group["members"]:
        print(f"  {member['qualified_name']} ({member['path']}:{member['start_line']}-{member['end_line']})")

```

## Key Source Files and Architecture

- **[[`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py)** – Core engine implementing exact grouping (lines 73-84), candidate pair generation (lines 86-100), Jaccard evaluation, and maximal clique extraction (lines 76-90).
- **[[`codebase_rag/tools/duplicate_detection.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py)** – Async tool wrapper that validates inputs and formats reports for agentic consumption.
- **[[`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)** – Command-line entry point defining the `duplicates` subcommand around line 1938.
- **[[`codebase_rag/constants/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/duplicates.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/duplicates.py)** – Default thresholds, message templates, and symbolic constants used throughout the detection pipeline.

## Summary

- Code-Graph-RAG detects structural duplication using **AST skeleton fingerprints**, not just text comparison.
- **Stage 1** finds exact copies in linear time via whole-skeleton fingerprint grouping.
- **Stage 2** finds similar code using **Jaccard similarity** of branch fingerprints, prefix-filtered candidate generation, and **maximal clique extraction**.
- Access the functionality via the **`cgr duplicates`** CLI command or the **`create_find_duplicates_tool`** async Python API.
- Default settings require a **0.8 similarity threshold** and a minimum of **5 AST nodes** per function, both configurable per scan.

## Frequently Asked Questions

### What is the difference between exact and similar duplicate detection in Code-Graph-RAG?

**Exact detection** identifies functions with identical AST skeletons where only identifier names differ, grouped by matching whole-skeleton fingerprints. **Similar detection** uses Jaccard overlap of branch-level fingerprints to find functions that have been edited or refactored but retain structural resemblance, extracted via maximal cliques in a similarity graph.

### How does the candidate pair generation optimize performance?

The [`_candidate_pairs`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py#L86-L100) function implements a **prefix-filtered AllPairs/PPJoin** strategy. Instead of comparing every function against every other function (O(n²)), it indexes only the rarest branch fingerprints required for a pair to meet the Jaccard threshold. This guarantees that no qualifying pair is missed while dramatically reducing the number of comparisons.

### Can I adjust the similarity threshold for duplicate detection?

Yes. Both the CLI (`--threshold` flag) and the programmatic API (`threshold` parameter in `default_duplicates_config` or `find_duplicate_code`) accept values between 0.0 and 1.0. The default is **0.8**, meaning two functions must share at least 80% of their branch fingerprints to qualify as similar duplicates.

### Is the duplicate detection process run synchronously or asynchronously?

The core engine in [`duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/duplicates.py) is synchronous, but the [`create_find_duplicates_tool`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py) wrapper exposes an **async interface** (`find_duplicate_code`) that runs the blocking detection logic in a background thread using `asyncio.to_thread`. This prevents event-loop blocking in async applications while leveraging the CPU-bound similarity calculations.