# How to Find Duplicate Code with code-graph-rag: Complete Detection Guide

> Find duplicate code with code-graph-rag using its two-stage detection pipeline. Detect exact clones and structurally similar functions via the CLI or Python API for complete code analysis.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-05

---

**Find duplicate code in your repository using code-graph-rag's two-stage detection pipeline via the `cgr duplicates` CLI or Python API, which identifies both exact clones and structurally similar functions through graph-based fingerprinting.**

The `vitali87/code-graph-rag` repository provides a sophisticated duplicate detection system that analyzes your codebase's structure to identify redundant implementations. By combining exact byte-for-byte matching with similarity analysis of control-flow branches, this tool helps you eliminate technical debt through the `codebase_rag.duplicates` module. You can invoke detection via command line, Python scripts, or LangChain-compatible agents.

## How Duplicate Detection Works

The detection engine lives in [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) and processes your codebase through a sequential pipeline. It first identifies perfect clones, then evaluates structural similarity for near-duplicates, while filtering out nested function false positives.

### Stage 1: Exact-Only Detection

The engine begins by grouping functions that are byte-for-byte identical, accounting for renamed copies. It groups rows by the **whole-skeleton fingerprint** stored at ingest time. When a fingerprint appears on more than one member, the system emits an `Exact` group. This stage executes during lines 44-85 of [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py).

### Stage 2: Similar Code Detection

For edited copies sharing high proportions of control-flow logic, the system implements a four-step process:

1. **Branch Fingerprinting**: Each function's control-flow branches receive unique fingerprints
2. **Candidate Generation**: An **All-Pairs/PPJoin** prefix filter generates candidate pairs while guaranteeing no qualifying pair is missed
3. **Similarity Scoring**: Jaccard similarity measures the overlap between branch sets, keeping pairs above the user-specified `threshold`
4. **Clique Enumeration**: The Bron-Kerbosch algorithm with pivoting reduces the remaining graph to **maximal cliques**, producing `Similar` groups

This logic resides in lines 86-124 of [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py).

### Nested Function Handling

To prevent a container function from being reported as a duplicate of its own closure, the engine applies a series of pruning checks. The functions `_member_nested_in`, `_only_nested_members`, and `_drop_contained_members` execute before similarity evaluation to remove nested members from consideration, as implemented in lines 40-80 of [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py).

## CLI Usage and Configuration

The `cgr duplicates` command wraps the detection engine and provides user-level filtering. Located in [`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) (lines 1515-1858), it resolves the project name, applies filters, and formats output as either a rich table or JSON.

### Basic Usage

```bash

# Detect duplicates in the current repository (project name inferred)

cgr duplicates

# Find only exact copies, ignore small functions, output JSON

cgr duplicates --exact-only --min-size 20 --output dup_report.json --format json

# CI integration: fail build if duplicates found with strict threshold

cgr duplicates --threshold 0.9 --fail-on-found

```

### Configuration Options

All parameters pass through `DuplicatesConfig` to `collect_duplicates_with_coverage` (lines 1629-1663 in [`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)):

| Option | Default | Purpose |
|--------|---------|---------|
| `--threshold` | `0.8` | Minimum Jaccard similarity for clone pairs (`cs.DUPLICATES_DEFAULT_THRESHOLD`) |
| `--min-size` | `10` | Minimum AST node count to include a function (`cs.DUPLICATES_DEFAULT_MIN_NODES`) |
| `--exact-only` | `False` | Skip similarity stage, report only exact copies |
| `--exclude` | `[]` | Glob patterns to filter paths (e.g., `*/generated/*`) |
| `--format` | `table` | Output format: `table` or `json` |
| `--fail-on-found` | `False` | Exit non-zero if duplicates detected (useful for CI pipelines) |

When the project root is known, table output includes clickable **OSC 8** hyperlinks that open source files in your configured editor.

## Python API Integration

For embedding detection in scripts or automation, import `collect_duplicates_with_coverage` and `default_duplicates_config` from `codebase_rag.duplicates`.

```python
from pathlib import Path
from codebase_rag.duplicates import (
    collect_duplicates_with_coverage,
    default_duplicates_config,
)
from codebase_rag.graph_service import MemgraphIngestor
from codebase_rag.utils.path_utils import derive_project_name

# Initialize your ingested graph connection

ingestor = MemgraphIngestor()  # Existing connected instance

project_name = derive_project_name(Path("/path/to/repo"))

config = default_duplicates_config(
    threshold=0.85,
    min_nodes=15,
    exclude_patterns=("*/generated/*",),
)

report = collect_duplicates_with_coverage(ingestor, project_name, config)

# Process results

for group in report.groups:
    print(f"Group ({group['kind']}) similarity {group['similarity']:.2%}")
    for member in group["members"]:
        print(f"  - {member['path']}:{member['start_line']}-{member['end_line']}")

```

## LLM Agent Integration

The repository ships a LangChain-compatible tool for autonomous agent workflows. Located in [`codebase_rag/tools/duplicate_detection.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py), this wrapper exposes the same engine through an async interface.

```python
from codebase_rag.tools.duplicate_detection import create_find_duplicates_tool
from codebase_rag.graph_service import MemgraphIngestor

tool = create_find_duplicates_tool(MemgraphIngestor())

# Async invocation with parameters

result = await tool.ainvoke({
    "project": "myproj",
    "threshold": 0.9,
    "min_size": 25,
    "limit": 10
})

```

## Summary

- **Two-stage pipeline**: The system first detects exact clones via skeleton fingerprints, then finds similar code using branch-based Jaccard similarity and maximal clique enumeration.
- **Configurable thresholds**: Adjust similarity cutoff (`--threshold`), minimum function size (`--min-size`), and exclusion patterns to reduce noise.
- **Multiple interfaces**: Access detection via the `cgr duplicates` CLI, direct Python API (`collect_duplicates_with_coverage`), or LangChain agent tools.
- **CI-ready**: Use `--fail-on-found` to enforce code quality gates in continuous integration pipelines.
- **Smart filtering**: Automatic nested function detection prevents false positives where closures are incorrectly matched against their parent containers.

## Frequently Asked Questions

### How does code-graph-rag distinguish between exact and similar duplicates?

Exact duplicates are identified by matching **whole-skeleton fingerprints** stored during ingestion, which capture the complete AST structure of a function. Similar duplicates are detected by comparing sets of **branch fingerprints** using Jaccard similarity, then grouping results into maximal cliques. Exact matches require byte-for-byte equivalence (excluding metadata), while similar matches require a configurable threshold of structural overlap (default 80%).

### Can I integrate duplicate detection into my CI pipeline?

Yes. Pass the `--fail-on-found` flag to the `cgr duplicates` command to return a non-zero exit code whenever duplicate groups are detected. Combine this with `--threshold 0.9` or `--exact-only` to enforce strict quality gates, and use `--exclude` patterns to skip generated code or third-party libraries that would otherwise trigger false positives.

### What prevents the tool from reporting a function as a duplicate of its inner closure?

The detection engine runs a preprocessing step that identifies nested function relationships. Using the helper functions `_member_nested_in`, `_only_nested_members`, and `_drop_contained_members` in [`codebase_rag/duplicates.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) (lines 40-80), the system prunes nested members before similarity evaluation. This ensures that a parent function containing helper definitions is never compared against those internal helpers.

### How do I exclude generated code or specific directories from analysis?

Use the `--exclude` parameter with glob patterns when running the CLI command: `cgr duplicates --exclude "*/generated/*" --exclude "*_test.go"`. In the Python API, pass a tuple of patterns to the `exclude_patterns` parameter of `default_duplicates_config()`. These filters apply during the candidate generation phase to reduce computational overhead and eliminate false positives from boilerplate or auto-generated files.