How grep-mcp Groups and Sorts Search Results by Repository

grep-mcp aggregates raw grep.app API hits into repository-specific buckets using a dictionary keyed by repository name, then sorts the final output in descending order by match count to prioritize the most relevant repositories.

The grep-mcp server transforms scattered code search results into structured, repository-centric datasets that developers can navigate efficiently. When querying the grep.app API across thousands of public repositories, the tool organizes individual file hits into coherent groups based on their source origin. This article examines the implementation in src/grep_mcp/server.py to explain exactly how grep-mcp groups and sorts search results by repository.

The Response Processing Pipeline

After receiving JSON from the grep.app API, the _format_grep_response function in src/grep_mcp/server.py processes the payload synchronously to ensure deterministic ordering. The function extracts the nested hits array from the response body, then iterates through results to build a grouped structure before returning data to the MCP client.

Repository Grouping Implementation

The grouping logic relies on three core operations: extracting repository identifiers, building temporary storage buckets, and aggregating match metadata.

Extracting Repository Identifiers

The code first retrieves the hits array from the API response around line 992:

hits = data.get("hits", {}).get("hits", [])

For each hit in the collection (limited to 10 results), the function extracts the repository name from the nested repo.raw field:

repo = hit.get("repo", {}).get("raw", "Unknown")

This value typically follows the format owner/repository and serves as the dictionary key for grouping results.

Building Repository Buckets

The function initializes an empty dictionary to hold repository-specific results:

repo_groups = {}

As the code iterates through hits, it checks for existing keys and initializes new lists when encountering unseen repositories. This logic appears around lines 1028–1031:

if repo not in repo_groups:
    repo_groups[repo] = []
repo_groups[repo].append(result)

This bucket approach ensures all matches from the same repository cluster together regardless of their original order in the API response.

Constructing Result Objects

Each hit gets transformed into a structured result object containing the file path, branch name, line numbers, language hint, and formatted code snippet. These objects append to their respective repository lists, maintaining the relationship between match metadata and source repository.

Sorting by Match Count

After populating the dictionary, the function converts repo_groups into a list of repository objects. Each object tracks a matches_count field representing the total hits for that repository. The final sorting operation arranges repositories by relevance around lines 1042–1044:

results_by_repo.sort(key=lambda x: x["matches_count"], reverse=True)

This descending sort ensures repositories with the most matches appear first in the results_by_repository array returned to the client.

Consuming the Grouped Output

Developers interact with the grouped results through the grep_query tool. The following example demonstrates a search for asyncio patterns:

import asyncio
from grep_mcp.server import grep_query

async def demo():
    result_json = await grep_query(
        query="asyncio.run",
        language="python"
    )
    print(result_json)

asyncio.run(demo())

The returned JSON contains a top-level results_by_repository array already sorted by relevance. To inspect the grouped structure programmatically:

import json

data = json.loads(result_json)

for repo in data["results_by_repository"]:
    print(f"Repo: {repo['repository']} (total matches: {repo['matches_count']})")
    for file in repo["files"]:
        print(f"  • {file['file_path']} → lines {file['line_numbers']}")

Summary

  • grep-mcp processes grep.app API responses in src/grep_mcp/server.py using the _format_grep_response helper function.
  • Results group by repository name extracted from hit["repo"]["raw"] and store in a dictionary structure keyed by repository identifier.
  • The implementation limits processing to 10 hits per query for performance optimization.
  • Final output sorts repositories by matches_count in descending order to surface the most relevant sources first.
  • Synchronous processing guarantees deterministic ordering for downstream MCP clients.

Frequently Asked Questions

How does grep-mcp handle repositories with no matches?

Repositories with no matches do not appear in the output. The grouping logic only creates entries for repositories present in the grep.app API response hits array, so repositories without matching files are implicitly excluded from the results.

What is the maximum number of repositories returned in one query?

The tool processes a hard limit of 10 hits from the API response. Since multiple hits can belong to the same repository, the actual number of unique repositories returned depends on how the matches distribute across sources, but will not exceed 10 distinct repositories.

Can I customize the sorting order to alphabetical instead of by match count?

The current implementation in _format_grep_response hardcodes descending sort by matches_count. To change the ordering, you would need to modify the sort key lambda in src/grep_mcp/server.py or implement client-side reordering of the results_by_repository array after receiving the response.

Where does the repository name parsing occur in the source code?

Repository name extraction happens within the _format_grep_response function in src/grep_mcp/server.py, specifically using hit.get("repo", {}).get("raw", "Unknown") to handle missing metadata gracefully and default to "Unknown" when repository information is unavailable.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →