Can AI Detect Duplicate Issues in a GitHub Repository? A Practical Implementation Guide

Yes, AI can detect duplicate GitHub issues by leveraging GitHub Actions workflows that feed issue data to large language models via GitHub Models, analyzing titles and descriptions for semantic similarity to flag potential duplicates automatically.

The awesome-continuous-ai repository demonstrates this capability through an automated workflow designed for content deduplication. While the repository implements this pattern specifically for README tool entries, the underlying architecture applies directly to GitHub issue deduplication by modifying the input source and system prompt.

How AI Detects Duplicate GitHub Issues

The GitHub Actions and GitHub Models Architecture

The detection capability centers on the integration of GitHub Actions with GitHub Models (managed LLM inference). According to the source code in githubnext/awesome-continuous-ai, the workflow .github/workflows/detect-duplicate-tools.yml orchestrates this process by sending repository content to an AI model and interpreting structured responses.

The workflow triggers on every push or pull request that modifies README.md and executes four critical operations:

  1. Checkout repository using actions/checkout@v4
  2. Run AI inference via actions/ai-inference@main with the openai/gpt-4o model and a maximum token limit of 14,000
  3. Parse model output for specific signals—either DUPLICATES_FOUND followed by a structured list, or NO_DUPLICATES_FOUND
  4. Create GitHub issues automatically when duplicates are detected, applying labels genai and documentation for triage

Prompt-Driven Detection Logic

Rather than implementing hard-coded comparison algorithms, the approach encodes detection rules within a system prompt (defined in lines 31-73 of the workflow file). The prompt instructs the model to identify duplicates based on three specific criteria: exact URL matches, similar tool names (case-insensitive with punctuation normalization), and identical products appearing in different sections. The prompt explicitly excludes duplicates within the "Programming Frameworks" subsection, demonstrating how to configure domain-specific exceptions.

This prompt-centric architecture provides extensibility—changing the input file from README.md to an issues export and adjusting the prompt wording adapts the workflow for detecting duplicate GitHub issues instead of documentation entries.

Implementing Duplicate Issue Detection

Adapting the Workflow for Issues

To detect duplicate issues rather than README tools, modify the workflow to fetch open issues as JSON and adjust the system prompt to analyze issue metadata. The repository suggests fetching data via the GitHub CLI and processing this through the same inference action.


# .github/workflows/detect-duplicate-issues.yml

name: Detect Duplicate Issues
on:
  schedule:
    - cron: '0 2 * * *'   # run daily

  workflow_dispatch:

jobs:
  detect-duplicates:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Fetch open issues as JSON
        run: |
          gh api repos/${{ github.repository }}/issues?state=open > issues.json
        env:
          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
      - uses: actions/ai-inference@main
        id: detect-issues
        with:
          token: ${{ secrets.GITHUB_TOKEN }}
          model: openai/gpt-4o
          max-tokens: 12000
          prompt-file: 'issues.json'
          system-prompt: |
            You are an AI assistant tasked with finding duplicate GitHub issues.
            Consider two issues duplicates if they have:
            - Identical or near-identical titles (case-insensitive, ignore punctuation)
            - Highly similar descriptions or body text
            - Overlapping labels or tags
            Respond with DUPLICATES_FOUND and list the issue numbers, or NO_DUPLICATES_FOUND.
      - name: Create duplicate-issue report
        if: contains(steps.detect-issues.outputs.response, 'DUPLICATES_FOUND')
        run: |
          gh issue create --title "Duplicate Issues Detected" \
            --body-file ${{ steps.detect-issues.outputs.response-file }} \
            --label "genai,triage"
        env:
          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}

Key modifications from the original tool-detection workflow include:

  • Scheduled execution using cron: '0 2 * * *' for daily scanning rather than file-triggered runs
  • API endpoint targeting repos/${{ github.repository }}/issues?state=open instead of README content
  • Prompt focus on issue titles, bodies, and labels rather than tool URLs and descriptions

Using the Reusable Action

The repository references pelikhan/action-genai-issue-dedup, a reusable GitHub Action specifically designed for issue deduplication. This action encapsulates the inference logic and can trigger in real-time when issues are opened or edited:


# .github/workflows/issue-dedup.yml

name: Issue Deduplication
on:
  issues:
    types: [opened, edited]

jobs:
  dedup:
    runs-on: ubuntu-latest
    steps:
      - uses: pelikhan/action-genai-issue-dedup@v1
        with:
          github-token: ${{ secrets.GITHUB_TOKEN }}
          model: openai/gpt-4o
          max-tokens: 8000

This implementation provides immediate duplicate detection as issues are created, eliminating the latency of scheduled scans while automatically flagging potential duplicates for maintainers.

Technical Benefits of the AI Approach

The architecture implemented in githubnext/awesome-continuous-ai offers distinct advantages for duplicate detection:

  • Semantic Understanding: The openai/gpt-4o model identifies conceptual duplicates beyond exact string matches, catching paraphrased issues or those with similar intent but different wording that traditional regex or TF-IDF algorithms would miss.
  • Scalability: With a max-tokens configuration of up to 14,000 tokens, the workflow handles large backlogs without custom chunking logic or embedding pipeline maintenance.
  • Low Maintenance: No custom fuzzy-matching code, similarity thresholds, or vector databases are required; the LLM handles comparison logic through the system prompt.
  • Automated Triage: The workflow uses gh issue create to generate structured reports with consistent labeling, ensuring human reviewers receive formatted duplicate lists without manual screening.

Summary

  • AI detects duplicate GitHub issues by analyzing textual similarity through LLM inference within GitHub Actions workflows, as demonstrated in githubnext/awesome-continuous-ai.
  • The reference implementation in .github/workflows/detect-duplicate-tools.yml uses actions/ai-inference@main with a custom system prompt to encode detection criteria and structured output formats.
  • Adapting the pattern for issues requires fetching open issues via gh api, adjusting the prompt to analyze titles and bodies, and maintaining the DUPLICATES_FOUND response parsing logic.
  • For production deployment, either customize the generic workflow pattern for batch scanning or implement action-genai-issue-dedup for real-time detection on issue creation.

Frequently Asked Questions

How accurate is AI at detecting duplicate GitHub issues?

Accuracy depends on the LLM model quality and prompt specificity. The openai/gpt-4o implementation in the awesome-continuous-ai workflow provides high semantic accuracy, identifying paraphrased duplicates and conceptually similar issues that keyword matching would miss. However, results may vary based on issue complexity and require prompt tuning for domain-specific terminology.

What are the costs associated with AI-powered duplicate detection?

GitHub Models provides managed inference with rate limits suitable for typical repository volumes. The workflow configuration uses max-tokens: 14000 for processing large inputs, but costs scale with token consumption. For high-volume repositories, implement the scheduled cron approach (daily scans) rather than real-time triggers to manage API quotas effectively.

Can this workflow detect duplicates across multiple repositories?

Yes, by extending the gh api call to fetch issues from multiple repositories and aggregating them into the JSON payload sent to the model. The system prompt would require adjustment to include repository identifiers in the duplicate comparison logic, but the actions/ai-inference@main action processes arbitrary text inputs regardless of source origin.

How does AI detection compare to traditional duplicate detection algorithms?

Traditional approaches rely on Jaccard similarity, TF-IDF vector comparison, or cosine similarity on embeddings with manually tuned thresholds. The AI approach in detect-duplicate-tools.yml eliminates embedding pipeline maintenance and threshold calibration, leveraging the LLM's built-in semantic understanding instead. This reduces engineering overhead but introduces dependency on model availability and potential variability in output formatting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →