How to Find Duplicate Code with code-graph-rag: Complete Detection Guide
Find duplicate code in your repository using code-graph-rag's two-stage detection pipeline via the cgr duplicates CLI or Python API, which identifies both exact clones and structurally similar functions through graph-based fingerprinting.
The vitali87/code-graph-rag repository provides a sophisticated duplicate detection system that analyzes your codebase's structure to identify redundant implementations. By combining exact byte-for-byte matching with similarity analysis of control-flow branches, this tool helps you eliminate technical debt through the codebase_rag.duplicates module. You can invoke detection via command line, Python scripts, or LangChain-compatible agents.
How Duplicate Detection Works
The detection engine lives in codebase_rag/duplicates.py and processes your codebase through a sequential pipeline. It first identifies perfect clones, then evaluates structural similarity for near-duplicates, while filtering out nested function false positives.
Stage 1: Exact-Only Detection
The engine begins by grouping functions that are byte-for-byte identical, accounting for renamed copies. It groups rows by the whole-skeleton fingerprint stored at ingest time. When a fingerprint appears on more than one member, the system emits an Exact group. This stage executes during lines 44-85 of codebase_rag/duplicates.py.
Stage 2: Similar Code Detection
For edited copies sharing high proportions of control-flow logic, the system implements a four-step process:
- Branch Fingerprinting: Each function's control-flow branches receive unique fingerprints
- Candidate Generation: An All-Pairs/PPJoin prefix filter generates candidate pairs while guaranteeing no qualifying pair is missed
- Similarity Scoring: Jaccard similarity measures the overlap between branch sets, keeping pairs above the user-specified
threshold - Clique Enumeration: The Bron-Kerbosch algorithm with pivoting reduces the remaining graph to maximal cliques, producing
Similargroups
This logic resides in lines 86-124 of codebase_rag/duplicates.py.
Nested Function Handling
To prevent a container function from being reported as a duplicate of its own closure, the engine applies a series of pruning checks. The functions _member_nested_in, _only_nested_members, and _drop_contained_members execute before similarity evaluation to remove nested members from consideration, as implemented in lines 40-80 of codebase_rag/duplicates.py.
CLI Usage and Configuration
The cgr duplicates command wraps the detection engine and provides user-level filtering. Located in codebase_rag/cli.py (lines 1515-1858), it resolves the project name, applies filters, and formats output as either a rich table or JSON.
Basic Usage
# Detect duplicates in the current repository (project name inferred)
cgr duplicates
# Find only exact copies, ignore small functions, output JSON
cgr duplicates --exact-only --min-size 20 --output dup_report.json --format json
# CI integration: fail build if duplicates found with strict threshold
cgr duplicates --threshold 0.9 --fail-on-found
Configuration Options
All parameters pass through DuplicatesConfig to collect_duplicates_with_coverage (lines 1629-1663 in codebase_rag/cli.py):
| Option | Default | Purpose |
|---|---|---|
--threshold |
0.8 |
Minimum Jaccard similarity for clone pairs (cs.DUPLICATES_DEFAULT_THRESHOLD) |
--min-size |
10 |
Minimum AST node count to include a function (cs.DUPLICATES_DEFAULT_MIN_NODES) |
--exact-only |
False |
Skip similarity stage, report only exact copies |
--exclude |
[] |
Glob patterns to filter paths (e.g., */generated/*) |
--format |
table |
Output format: table or json |
--fail-on-found |
False |
Exit non-zero if duplicates detected (useful for CI pipelines) |
When the project root is known, table output includes clickable OSC 8 hyperlinks that open source files in your configured editor.
Python API Integration
For embedding detection in scripts or automation, import collect_duplicates_with_coverage and default_duplicates_config from codebase_rag.duplicates.
from pathlib import Path
from codebase_rag.duplicates import (
collect_duplicates_with_coverage,
default_duplicates_config,
)
from codebase_rag.graph_service import MemgraphIngestor
from codebase_rag.utils.path_utils import derive_project_name
# Initialize your ingested graph connection
ingestor = MemgraphIngestor() # Existing connected instance
project_name = derive_project_name(Path("/path/to/repo"))
config = default_duplicates_config(
threshold=0.85,
min_nodes=15,
exclude_patterns=("*/generated/*",),
)
report = collect_duplicates_with_coverage(ingestor, project_name, config)
# Process results
for group in report.groups:
print(f"Group ({group['kind']}) similarity {group['similarity']:.2%}")
for member in group["members"]:
print(f" - {member['path']}:{member['start_line']}-{member['end_line']}")
LLM Agent Integration
The repository ships a LangChain-compatible tool for autonomous agent workflows. Located in codebase_rag/tools/duplicate_detection.py, this wrapper exposes the same engine through an async interface.
from codebase_rag.tools.duplicate_detection import create_find_duplicates_tool
from codebase_rag.graph_service import MemgraphIngestor
tool = create_find_duplicates_tool(MemgraphIngestor())
# Async invocation with parameters
result = await tool.ainvoke({
"project": "myproj",
"threshold": 0.9,
"min_size": 25,
"limit": 10
})
Summary
- Two-stage pipeline: The system first detects exact clones via skeleton fingerprints, then finds similar code using branch-based Jaccard similarity and maximal clique enumeration.
- Configurable thresholds: Adjust similarity cutoff (
--threshold), minimum function size (--min-size), and exclusion patterns to reduce noise. - Multiple interfaces: Access detection via the
cgr duplicatesCLI, direct Python API (collect_duplicates_with_coverage), or LangChain agent tools. - CI-ready: Use
--fail-on-foundto enforce code quality gates in continuous integration pipelines. - Smart filtering: Automatic nested function detection prevents false positives where closures are incorrectly matched against their parent containers.
Frequently Asked Questions
How does code-graph-rag distinguish between exact and similar duplicates?
Exact duplicates are identified by matching whole-skeleton fingerprints stored during ingestion, which capture the complete AST structure of a function. Similar duplicates are detected by comparing sets of branch fingerprints using Jaccard similarity, then grouping results into maximal cliques. Exact matches require byte-for-byte equivalence (excluding metadata), while similar matches require a configurable threshold of structural overlap (default 80%).
Can I integrate duplicate detection into my CI pipeline?
Yes. Pass the --fail-on-found flag to the cgr duplicates command to return a non-zero exit code whenever duplicate groups are detected. Combine this with --threshold 0.9 or --exact-only to enforce strict quality gates, and use --exclude patterns to skip generated code or third-party libraries that would otherwise trigger false positives.
What prevents the tool from reporting a function as a duplicate of its inner closure?
The detection engine runs a preprocessing step that identifies nested function relationships. Using the helper functions _member_nested_in, _only_nested_members, and _drop_contained_members in codebase_rag/duplicates.py (lines 40-80), the system prunes nested members before similarity evaluation. This ensures that a parent function containing helper definitions is never compared against those internal helpers.
How do I exclude generated code or specific directories from analysis?
Use the --exclude parameter with glob patterns when running the CLI command: cgr duplicates --exclude "*/generated/*" --exclude "*_test.go". In the Python API, pass a tuple of patterns to the exclude_patterns parameter of default_duplicates_config(). These filters apply during the candidate generation phase to reduce computational overhead and eliminate false positives from boilerplate or auto-generated files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →