How to Find Structurally Duplicated Code with Code-Graph-RAG
Code-Graph-RAG detects structural code duplication through a two-stage pipeline—first grouping exact copies via whole-skeleton fingerprints, then detecting similar functions using Jaccard similarity of branch fingerprints and maximal clique extraction—exposed through both a command-line interface and a Python async tool.
Finding structurally duplicated code is essential for maintaining clean, refactorable repositories. The Code-Graph-RAG project (vitali87/code-graph-rag) provides a sophisticated engine that analyzes Abstract Syntax Tree (AST) skeletons to detect both exact copies and near-duplicate functions. This article explains how to leverage the duplicate detection capabilities using the CLI and programmatic APIs.
How Structural Duplicate Detection Works
The detection engine resides in [codebase_rag/duplicates.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) and operates in two distinct stages to balance precision and performance.
Stage 1: Exact Copy Grouping
The engine first identifies functions that are identical except for identifier names by grouping entries sharing the same whole-skeleton fingerprint. This fingerprint is computed at ingest time and stored for every function definition.
The _exact_groups function builds these groups in linear time by aggregating members with identical fingerprints:
# _exact_groups – creates a DuplicateGroup for each fingerprint with >1 members
# (see lines 73-84 of duplicates.py)
def _exact_groups(entries: dict[str, _Entry]) -> list[DuplicateGroup]:
return [
DuplicateGroup(
kind=cs.KIND_EXACT,
similarity=1.0,
node_count=entry.node_count,
members=_sorted_members(entry.members),
)
for entry in entries.values()
if len(entry.members) > 1
]
This stage runs in O(n) time because it relies on a simple dictionary lookup of pre-computed hashes.
Stage 2: Similar Copy Detection
When exact_only is disabled, the engine searches for similar functions that may have undergone minor edits. This stage uses a prefix-filtered AllPairs/PPJoin algorithm to avoid an expensive O(n²) comparison.
-
Candidate Generation: The
_candidate_pairsfunction indexes only the rarest branch fingerprints that must appear in any qualifying pair, guaranteeing no valid matches are missed while pruning the search space. -
Edge Qualification: For each candidate pair, the
_edge_qualifiesfunction calculates Jaccard similarity between the sets of branch-level fingerprints. It also filters pairs that represent merely nested definitions using_only_nested_members. -
Clique Extraction: The engine builds a similarity graph where nodes are functions and edges represent qualified pairs. It then extracts maximal cliques using a pivoted Bron-Kerbosch algorithm in
_maximal_cliques. Each clique represents a group of functions where every member meets the similarity threshold with every other member.
The default similarity threshold is 0.8 and the minimum node count is 5, both configurable via default_duplicates_config.
Finding Duplicates via the CLI
The cgr duplicates command in [codebase_rag/cli.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) provides immediate access to the detection engine:
# Scan the default project with default settings (threshold 0.8, min-size 5)
cgr duplicates
# Scan a specific project with custom parameters
cgr duplicates --project my_project --threshold 0.9 --min-size 10 --limit 5
# Output as JSON for further processing
cgr duplicates --project my_project --format json
The CLI invokes collect_duplicates_with_coverage, formats the resulting DuplicatesReport, and prints either a human-readable table or structured JSON depending on the --format flag.
Detecting Duplicates Programmatically
For integration into async applications or agentic workflows, use the tool wrapper in [codebase_rag/tools/duplicate_detection.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py):
from codebase_rag.services import GraphQueryClient
from codebase_rag.tools.duplicate_detection import create_find_duplicates_tool
# Initialize your GraphQL query client (must implement QueryProtocol)
client = GraphQueryClient()
# Create the tool instance
tool = create_find_duplicates_tool(client)
# Execute detection asynchronously
report = await tool.find_duplicate_code(
project="my_project",
threshold=0.85,
min_size=8,
limit=10,
)
print(report) # Human-readable summary of duplicate groups
The tool validates arguments, resolves the target project, executes collect_duplicates in a background thread to prevent blocking, and returns a formatted text report.
For synchronous, low-level access, call the engine directly:
from codebase_rag.duplicates import collect_duplicates, default_duplicates_config
config = default_duplicates_config(threshold=0.75, min_nodes=6)
groups = collect_duplicates(
ingestor=client,
project_name="my_project",
config=config
)
for group in groups:
print(f"Group ({group['kind']}) – similarity {group['similarity']}")
for member in group["members"]:
print(f" {member['qualified_name']} ({member['path']}:{member['start_line']}-{member['end_line']})")
Key Source Files and Architecture
- [
codebase_rag/duplicates.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/duplicates.py) – Core engine implementing exact grouping (lines 73-84), candidate pair generation (lines 86-100), Jaccard evaluation, and maximal clique extraction (lines 76-90). - [
codebase_rag/tools/duplicate_detection.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/duplicate_detection.py) – Async tool wrapper that validates inputs and formats reports for agentic consumption. - [
codebase_rag/cli.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) – Command-line entry point defining theduplicatessubcommand around line 1938. - [
codebase_rag/constants/duplicates.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/duplicates.py) – Default thresholds, message templates, and symbolic constants used throughout the detection pipeline.
Summary
- Code-Graph-RAG detects structural duplication using AST skeleton fingerprints, not just text comparison.
- Stage 1 finds exact copies in linear time via whole-skeleton fingerprint grouping.
- Stage 2 finds similar code using Jaccard similarity of branch fingerprints, prefix-filtered candidate generation, and maximal clique extraction.
- Access the functionality via the
cgr duplicatesCLI command or thecreate_find_duplicates_toolasync Python API. - Default settings require a 0.8 similarity threshold and a minimum of 5 AST nodes per function, both configurable per scan.
Frequently Asked Questions
What is the difference between exact and similar duplicate detection in Code-Graph-RAG?
Exact detection identifies functions with identical AST skeletons where only identifier names differ, grouped by matching whole-skeleton fingerprints. Similar detection uses Jaccard overlap of branch-level fingerprints to find functions that have been edited or refactored but retain structural resemblance, extracted via maximal cliques in a similarity graph.
How does the candidate pair generation optimize performance?
The _candidate_pairs function implements a prefix-filtered AllPairs/PPJoin strategy. Instead of comparing every function against every other function (O(n²)), it indexes only the rarest branch fingerprints required for a pair to meet the Jaccard threshold. This guarantees that no qualifying pair is missed while dramatically reducing the number of comparisons.
Can I adjust the similarity threshold for duplicate detection?
Yes. Both the CLI (--threshold flag) and the programmatic API (threshold parameter in default_duplicates_config or find_duplicate_code) accept values between 0.0 and 1.0. The default is 0.8, meaning two functions must share at least 80% of their branch fingerprints to qualify as similar duplicates.
Is the duplicate detection process run synchronously or asynchronously?
The core engine in duplicates.py is synchronous, but the create_find_duplicates_tool wrapper exposes an async interface (find_duplicate_code) that runs the blocking detection logic in a background thread using asyncio.to_thread. This prevents event-loop blocking in async applications while leveraging the CPU-bound similarity calculations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →