Can AI Detect Duplicate Issues in a GitHub Repository? A Practical Implementation Guide
Yes, AI can detect duplicate GitHub issues by leveraging GitHub Actions workflows that feed issue data to large language models via GitHub Models, analyzing titles and descriptions for semantic similarity to flag potential duplicates automatically.
The awesome-continuous-ai repository demonstrates this capability through an automated workflow designed for content deduplication. While the repository implements this pattern specifically for README tool entries, the underlying architecture applies directly to GitHub issue deduplication by modifying the input source and system prompt.
How AI Detects Duplicate GitHub Issues
The GitHub Actions and GitHub Models Architecture
The detection capability centers on the integration of GitHub Actions with GitHub Models (managed LLM inference). According to the source code in githubnext/awesome-continuous-ai, the workflow .github/workflows/detect-duplicate-tools.yml orchestrates this process by sending repository content to an AI model and interpreting structured responses.
The workflow triggers on every push or pull request that modifies README.md and executes four critical operations:
- Checkout repository using
actions/checkout@v4 - Run AI inference via
actions/ai-inference@mainwith theopenai/gpt-4omodel and a maximum token limit of 14,000 - Parse model output for specific signals—either
DUPLICATES_FOUNDfollowed by a structured list, orNO_DUPLICATES_FOUND - Create GitHub issues automatically when duplicates are detected, applying labels
genaianddocumentationfor triage
Prompt-Driven Detection Logic
Rather than implementing hard-coded comparison algorithms, the approach encodes detection rules within a system prompt (defined in lines 31-73 of the workflow file). The prompt instructs the model to identify duplicates based on three specific criteria: exact URL matches, similar tool names (case-insensitive with punctuation normalization), and identical products appearing in different sections. The prompt explicitly excludes duplicates within the "Programming Frameworks" subsection, demonstrating how to configure domain-specific exceptions.
This prompt-centric architecture provides extensibility—changing the input file from README.md to an issues export and adjusting the prompt wording adapts the workflow for detecting duplicate GitHub issues instead of documentation entries.
Implementing Duplicate Issue Detection
Adapting the Workflow for Issues
To detect duplicate issues rather than README tools, modify the workflow to fetch open issues as JSON and adjust the system prompt to analyze issue metadata. The repository suggests fetching data via the GitHub CLI and processing this through the same inference action.
# .github/workflows/detect-duplicate-issues.yml
name: Detect Duplicate Issues
on:
schedule:
- cron: '0 2 * * *' # run daily
workflow_dispatch:
jobs:
detect-duplicates:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Fetch open issues as JSON
run: |
gh api repos/${{ github.repository }}/issues?state=open > issues.json
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
- uses: actions/ai-inference@main
id: detect-issues
with:
token: ${{ secrets.GITHUB_TOKEN }}
model: openai/gpt-4o
max-tokens: 12000
prompt-file: 'issues.json'
system-prompt: |
You are an AI assistant tasked with finding duplicate GitHub issues.
Consider two issues duplicates if they have:
- Identical or near-identical titles (case-insensitive, ignore punctuation)
- Highly similar descriptions or body text
- Overlapping labels or tags
Respond with DUPLICATES_FOUND and list the issue numbers, or NO_DUPLICATES_FOUND.
- name: Create duplicate-issue report
if: contains(steps.detect-issues.outputs.response, 'DUPLICATES_FOUND')
run: |
gh issue create --title "Duplicate Issues Detected" \
--body-file ${{ steps.detect-issues.outputs.response-file }} \
--label "genai,triage"
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
Key modifications from the original tool-detection workflow include:
- Scheduled execution using
cron: '0 2 * * *'for daily scanning rather than file-triggered runs - API endpoint targeting
repos/${{ github.repository }}/issues?state=openinstead of README content - Prompt focus on issue titles, bodies, and labels rather than tool URLs and descriptions
Using the Reusable Action
The repository references pelikhan/action-genai-issue-dedup, a reusable GitHub Action specifically designed for issue deduplication. This action encapsulates the inference logic and can trigger in real-time when issues are opened or edited:
# .github/workflows/issue-dedup.yml
name: Issue Deduplication
on:
issues:
types: [opened, edited]
jobs:
dedup:
runs-on: ubuntu-latest
steps:
- uses: pelikhan/action-genai-issue-dedup@v1
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: openai/gpt-4o
max-tokens: 8000
This implementation provides immediate duplicate detection as issues are created, eliminating the latency of scheduled scans while automatically flagging potential duplicates for maintainers.
Technical Benefits of the AI Approach
The architecture implemented in githubnext/awesome-continuous-ai offers distinct advantages for duplicate detection:
- Semantic Understanding: The
openai/gpt-4omodel identifies conceptual duplicates beyond exact string matches, catching paraphrased issues or those with similar intent but different wording that traditional regex or TF-IDF algorithms would miss. - Scalability: With a
max-tokensconfiguration of up to 14,000 tokens, the workflow handles large backlogs without custom chunking logic or embedding pipeline maintenance. - Low Maintenance: No custom fuzzy-matching code, similarity thresholds, or vector databases are required; the LLM handles comparison logic through the system prompt.
- Automated Triage: The workflow uses
gh issue createto generate structured reports with consistent labeling, ensuring human reviewers receive formatted duplicate lists without manual screening.
Summary
- AI detects duplicate GitHub issues by analyzing textual similarity through LLM inference within GitHub Actions workflows, as demonstrated in
githubnext/awesome-continuous-ai. - The reference implementation in
.github/workflows/detect-duplicate-tools.ymlusesactions/ai-inference@mainwith a custom system prompt to encode detection criteria and structured output formats. - Adapting the pattern for issues requires fetching open issues via
gh api, adjusting the prompt to analyze titles and bodies, and maintaining theDUPLICATES_FOUNDresponse parsing logic. - For production deployment, either customize the generic workflow pattern for batch scanning or implement
action-genai-issue-dedupfor real-time detection on issue creation.
Frequently Asked Questions
How accurate is AI at detecting duplicate GitHub issues?
Accuracy depends on the LLM model quality and prompt specificity. The openai/gpt-4o implementation in the awesome-continuous-ai workflow provides high semantic accuracy, identifying paraphrased duplicates and conceptually similar issues that keyword matching would miss. However, results may vary based on issue complexity and require prompt tuning for domain-specific terminology.
What are the costs associated with AI-powered duplicate detection?
GitHub Models provides managed inference with rate limits suitable for typical repository volumes. The workflow configuration uses max-tokens: 14000 for processing large inputs, but costs scale with token consumption. For high-volume repositories, implement the scheduled cron approach (daily scans) rather than real-time triggers to manage API quotas effectively.
Can this workflow detect duplicates across multiple repositories?
Yes, by extending the gh api call to fetch issues from multiple repositories and aggregating them into the JSON payload sent to the model. The system prompt would require adjustment to include repository identifiers in the duplicate comparison logic, but the actions/ai-inference@main action processes arbitrary text inputs regardless of source origin.
How does AI detection compare to traditional duplicate detection algorithms?
Traditional approaches rely on Jaccard similarity, TF-IDF vector comparison, or cosine similarity on embeddings with manually tuned thresholds. The AI approach in detect-duplicate-tools.yml eliminates embedding pipeline maintenance and threshold calibration, leveraging the LLM's built-in semantic understanding instead. This reduces engineering overhead but introduces dependency on model availability and potential variability in output formatting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →