How to Detect Code Duplication with Code-Graph-RAG: A Complete Guide to Finding Exact and Similar Clones

Code-Graph-RAG detects code duplication by building a fingerprint graph during ingestion and running a two-stage engine that identifies exact copies via AST fingerprints and similar clones through Jaccard similarity analysis on branch fingerprints.

Code-Graph-RAG is a multilanguage codebase analysis tool that uses Tree-sitter parsing to detect structural code clones. The system creates AST fingerprints during the ingestion phase, then analyzes these fingerprints to find both verbatim copies and edited variants across your entire repository.

Understanding the Two-Stage Duplicate Detection Engine

The duplicate detection system in codebase_rag/duplicates.py operates through two complementary stages. Each stage targets a different type of code duplication, from identical function copies to logic that has been modified but remains structurally similar.

Stage 1: Detecting Exact and Renamed Copies

The first stage identifies functions with identical whole-skeleton fingerprints. During ingestion, every function node receives an AST_FINGERPRINT value generated by the Tree-sitter parser.

In codebase_rag/duplicates.py, the _entries_from_rows function collects functions sharing the same fingerprint into a single _Entry. The _exact_groups method then emits these collections as exact duplicate groups. This catches direct copy-paste operations and renamed variables, since the underlying AST structure remains unchanged.

This stage runs quickly because it relies on exact hash matching rather than pairwise comparison.

Stage 2: Detecting Similar and Edited Copies

The second stage finds code that has been edited but retains high structural overlap. This process uses branch fingerprints (statement-level hashes) and involves several sophisticated algorithms:

  1. Prefix Indexing (PPJoin): The _prefix_index method builds an index only on the rarest branch fingerprints needed to satisfy the Jaccard threshold, dramatically reducing the candidate search space.

  2. Candidate Generation: The _candidate_pairs function generates potential duplicate pairs from this prefix index.

  3. Threshold Filtering: The _threshold_adjacency method filters pairs by Jaccard similarity and excludes pure nesting relationships.

  4. Clique Enumeration: The _maximal_cliques function implements the Bron-Kerbosch algorithm to find maximal cliques in the similarity graph, where each clique represents a group of mutually similar functions.

The _similar_groups function orchestrates this entire pipeline, emitting groups with similarity scores and node counts.

Configuration and Deduplication Logic

Handling Nested Functions and Containers

To avoid reporting closures as duplicates of their containing functions, Code-Graph-RAG implements sophisticated containment detection. The system uses line-range checks via _span_contains and qualified-name hierarchy analysis through _qn_within.

The _drop_contained_members function removes nested members from duplicate groups, ensuring that inner functions are not flagged as duplicates of their outer scope. This prevents false positives when analyzing languages with nested function support like JavaScript or Python.

Configuring Detection Parameters

Detection behavior is controlled by the DuplicatesConfig dataclass defined in codebase_rag/constants.py (lines 735-749). Default values include:

  • Similarity threshold: Minimum Jaccard similarity for "similar" duplicates
  • Minimum node count: Minimum AST node size to consider a function
  • Exclusion patterns: Glob patterns for files to ignore (e.g., generated code)
  • Maximum groups: Cap on reported duplicate clusters

You can override these defaults via CLI flags or programmatic configuration when using the detection API.

Running Duplicate Detection via CLI

The Typer-based CLI in codebase_rag/cli.py exposes duplicate detection through the cgr duplicates command. The implementation uses _duplicates_group_cell and _duplicates_location_cell for formatted output.

Basic usage prints all duplicate groups for the current project:

cgr duplicates

Tune detection sensitivity with flags:


# Report only substantial functions (≥30 nodes) with 80% similarity

cgr duplicates --min-size 30 --threshold 0.8

Exclude generated or vendor files:

cgr duplicates --exclude-patterns "*gen_*" --exclude-patterns "*/vendor/*"

Programmatic and Agentic Access

Beyond the CLI, the detection engine is exposed as an MCP (Model Context Protocol) tool for AI-driven analysis. The codebase_rag/tools/duplicate_detection.py module provides create_find_duplicates_tool, which wraps the engine for async usage.

Use the tool programmatically:

from codebase_rag.tools.duplicate_detection import create_find_duplicates_tool

tool = create_find_duplicates_tool(ingestor)  # ingestor is a GraphQueryClient

result = await tool.function(
    project="myproj",
    threshold=0.75,
    min_size=20,
    limit=10
)
print(result)  # Human-readable duplicate report

This enables AI agents to call find_duplicate_code during automated refactoring workflows, receiving structured data about duplication clusters across the codebase.

Summary

  • Two-stage detection: Code-Graph-RAG first finds exact copies via AST_FINGERPRINT matching, then discovers similar clones using Jaccard similarity on branch fingerprints.
  • Algorithmic sophistication: The system uses PPJoin indexing for performance and Bron-Kerbosch clique enumeration for accurate grouping of similar functions.
  • Smart filtering: Nested functions are automatically excluded from parent function duplicate groups via _span_contains and _qn_within checks.
  • Multiple interfaces: Access detection via cgr duplicates CLI, direct Python API, or MCP agent tools in codebase_rag/tools/duplicate_detection.py.
  • Configurable thresholds: Adjust sensitivity through DuplicatesConfig in codebase_rag/constants.py or CLI flags like --threshold and --min-size.

Frequently Asked Questions

How does Code-Graph-RAG differentiate between exact copies and similar clones?

Exact copies share identical AST_FINGERPRINT values generated by Tree-sitter parsing, detected through hash lookup in _exact_groups. Similar clones are detected via Jaccard similarity of branch fingerprints in _similar_groups, requiring configurable threshold matching (default typically 0.7-0.8) and Bron-Kerbosch clique analysis to find related function clusters.

What algorithms does Code-Graph-RAG use for scalable duplicate detection?

The system implements PPJoin (Position-Prefix Join) indexing in _prefix_index to reduce candidate pairs, followed by Bron-Kerbosch maximal clique enumeration in _maximal_cliques to identify duplicate groups. This combination avoids the O(n²) complexity of naive pairwise comparison, making it feasible for large codebases.

Can Code-Graph-RAG detect code duplication across different programming languages?

Yes, because the detection relies on AST fingerprints rather than text comparison. Since Tree-sitter parses each language into a normalized AST structure, functions in Python and JavaScript could potentially be flagged as similar if their control flow structures match, though the tool primarily focuses on finding duplicates within the same language family.

How do I prevent test files or generated code from appearing in duplicate detection results?

Use the --exclude-patterns CLI flag or configure exclusion globs in DuplicatesConfig. For example, run cgr duplicates --exclude-patterns "*_test.py" --exclude-patterns "*/generated/*" to skip test suites and auto-generated files. These patterns are evaluated during the candidate generation phase in duplicates.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →