How grep-mcp Groups and Sorts Search Results by Repository
grep-mcp aggregates raw grep.app API hits into repository-specific buckets using a dictionary keyed by repository name, then sorts the final output in descending order by match count to prioritize the most relevant repositories.
The grep-mcp server transforms scattered code search results into structured, repository-centric datasets that developers can navigate efficiently. When querying the grep.app API across thousands of public repositories, the tool organizes individual file hits into coherent groups based on their source origin. This article examines the implementation in src/grep_mcp/server.py to explain exactly how grep-mcp groups and sorts search results by repository.
The Response Processing Pipeline
After receiving JSON from the grep.app API, the _format_grep_response function in src/grep_mcp/server.py processes the payload synchronously to ensure deterministic ordering. The function extracts the nested hits array from the response body, then iterates through results to build a grouped structure before returning data to the MCP client.
Repository Grouping Implementation
The grouping logic relies on three core operations: extracting repository identifiers, building temporary storage buckets, and aggregating match metadata.
Extracting Repository Identifiers
The code first retrieves the hits array from the API response around line 992:
hits = data.get("hits", {}).get("hits", [])
For each hit in the collection (limited to 10 results), the function extracts the repository name from the nested repo.raw field:
repo = hit.get("repo", {}).get("raw", "Unknown")
This value typically follows the format owner/repository and serves as the dictionary key for grouping results.
Building Repository Buckets
The function initializes an empty dictionary to hold repository-specific results:
repo_groups = {}
As the code iterates through hits, it checks for existing keys and initializes new lists when encountering unseen repositories. This logic appears around lines 1028–1031:
if repo not in repo_groups:
repo_groups[repo] = []
repo_groups[repo].append(result)
This bucket approach ensures all matches from the same repository cluster together regardless of their original order in the API response.
Constructing Result Objects
Each hit gets transformed into a structured result object containing the file path, branch name, line numbers, language hint, and formatted code snippet. These objects append to their respective repository lists, maintaining the relationship between match metadata and source repository.
Sorting by Match Count
After populating the dictionary, the function converts repo_groups into a list of repository objects. Each object tracks a matches_count field representing the total hits for that repository. The final sorting operation arranges repositories by relevance around lines 1042–1044:
results_by_repo.sort(key=lambda x: x["matches_count"], reverse=True)
This descending sort ensures repositories with the most matches appear first in the results_by_repository array returned to the client.
Consuming the Grouped Output
Developers interact with the grouped results through the grep_query tool. The following example demonstrates a search for asyncio patterns:
import asyncio
from grep_mcp.server import grep_query
async def demo():
result_json = await grep_query(
query="asyncio.run",
language="python"
)
print(result_json)
asyncio.run(demo())
The returned JSON contains a top-level results_by_repository array already sorted by relevance. To inspect the grouped structure programmatically:
import json
data = json.loads(result_json)
for repo in data["results_by_repository"]:
print(f"Repo: {repo['repository']} (total matches: {repo['matches_count']})")
for file in repo["files"]:
print(f" • {file['file_path']} → lines {file['line_numbers']}")
Summary
- grep-mcp processes grep.app API responses in
src/grep_mcp/server.pyusing the_format_grep_responsehelper function. - Results group by repository name extracted from
hit["repo"]["raw"]and store in a dictionary structure keyed by repository identifier. - The implementation limits processing to 10 hits per query for performance optimization.
- Final output sorts repositories by
matches_countin descending order to surface the most relevant sources first. - Synchronous processing guarantees deterministic ordering for downstream MCP clients.
Frequently Asked Questions
How does grep-mcp handle repositories with no matches?
Repositories with no matches do not appear in the output. The grouping logic only creates entries for repositories present in the grep.app API response hits array, so repositories without matching files are implicitly excluded from the results.
What is the maximum number of repositories returned in one query?
The tool processes a hard limit of 10 hits from the API response. Since multiple hits can belong to the same repository, the actual number of unique repositories returned depends on how the matches distribute across sources, but will not exceed 10 distinct repositories.
Can I customize the sorting order to alphabetical instead of by match count?
The current implementation in _format_grep_response hardcodes descending sort by matches_count. To change the ordering, you would need to modify the sort key lambda in src/grep_mcp/server.py or implement client-side reordering of the results_by_repository array after receiving the response.
Where does the repository name parsing occur in the source code?
Repository name extraction happens within the _format_grep_response function in src/grep_mcp/server.py, specifically using hit.get("repo", {}).get("raw", "Unknown") to handle missing metadata gracefully and default to "Unknown" when repository information is unavailable.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →