# How grep-mcp Groups and Sorts Search Results by Repository

> Learn how grep-mcp groups and sorts search results by repository. It aggregates API hits into buckets keyed by repo name and sorts by match count for prioritized results.

- Repository: [gal peretz/grep-mcp](https://github.com/galprz/grep-mcp)
- Tags: how-to-guide
- Published: 2026-02-16

---

**grep-mcp aggregates raw grep.app API hits into repository-specific buckets using a dictionary keyed by repository name, then sorts the final output in descending order by match count to prioritize the most relevant repositories.**

The `grep-mcp` server transforms scattered code search results into structured, repository-centric datasets that developers can navigate efficiently. When querying the grep.app API across thousands of public repositories, the tool organizes individual file hits into coherent groups based on their source origin. This article examines the implementation in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) to explain exactly how grep-mcp groups and sorts search results by repository.

## The Response Processing Pipeline

After receiving JSON from the grep.app API, the `_format_grep_response` function in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) processes the payload synchronously to ensure deterministic ordering. The function extracts the nested `hits` array from the response body, then iterates through results to build a grouped structure before returning data to the MCP client.

## Repository Grouping Implementation

The grouping logic relies on three core operations: extracting repository identifiers, building temporary storage buckets, and aggregating match metadata.

### Extracting Repository Identifiers

The code first retrieves the hits array from the API response around line 992:

```python
hits = data.get("hits", {}).get("hits", [])

```

For each hit in the collection (limited to 10 results), the function extracts the repository name from the nested `repo.raw` field:

```python
repo = hit.get("repo", {}).get("raw", "Unknown")

```

This value typically follows the format `owner/repository` and serves as the dictionary key for grouping results.

### Building Repository Buckets

The function initializes an empty dictionary to hold repository-specific results:

```python
repo_groups = {}

```

As the code iterates through hits, it checks for existing keys and initializes new lists when encountering unseen repositories. This logic appears around lines 1028–1031:

```python
if repo not in repo_groups:
    repo_groups[repo] = []
repo_groups[repo].append(result)

```

This bucket approach ensures all matches from the same repository cluster together regardless of their original order in the API response.

### Constructing Result Objects

Each hit gets transformed into a structured result object containing the **file path**, **branch name**, **line numbers**, **language hint**, and **formatted code snippet**. These objects append to their respective repository lists, maintaining the relationship between match metadata and source repository.

## Sorting by Match Count

After populating the dictionary, the function converts `repo_groups` into a list of repository objects. Each object tracks a `matches_count` field representing the total hits for that repository. The final sorting operation arranges repositories by relevance around lines 1042–1044:

```python
results_by_repo.sort(key=lambda x: x["matches_count"], reverse=True)

```

This descending sort ensures repositories with the most matches appear first in the `results_by_repository` array returned to the client.

## Consuming the Grouped Output

Developers interact with the grouped results through the `grep_query` tool. The following example demonstrates a search for asyncio patterns:

```python
import asyncio
from grep_mcp.server import grep_query

async def demo():
    result_json = await grep_query(
        query="asyncio.run",
        language="python"
    )
    print(result_json)

asyncio.run(demo())

```

The returned JSON contains a top-level `results_by_repository` array already sorted by relevance. To inspect the grouped structure programmatically:

```python
import json

data = json.loads(result_json)

for repo in data["results_by_repository"]:
    print(f"Repo: {repo['repository']} (total matches: {repo['matches_count']})")
    for file in repo["files"]:
        print(f"  • {file['file_path']} → lines {file['line_numbers']}")

```

## Summary

- **grep-mcp** processes grep.app API responses in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) using the `_format_grep_response` helper function.
- Results group by repository name extracted from `hit["repo"]["raw"]` and store in a dictionary structure keyed by repository identifier.
- The implementation limits processing to 10 hits per query for performance optimization.
- Final output sorts repositories by `matches_count` in descending order to surface the most relevant sources first.
- Synchronous processing guarantees deterministic ordering for downstream MCP clients.

## Frequently Asked Questions

### How does grep-mcp handle repositories with no matches?

Repositories with no matches do not appear in the output. The grouping logic only creates entries for repositories present in the grep.app API response hits array, so repositories without matching files are implicitly excluded from the results.

### What is the maximum number of repositories returned in one query?

The tool processes a hard limit of 10 hits from the API response. Since multiple hits can belong to the same repository, the actual number of unique repositories returned depends on how the matches distribute across sources, but will not exceed 10 distinct repositories.

### Can I customize the sorting order to alphabetical instead of by match count?

The current implementation in `_format_grep_response` hardcodes descending sort by `matches_count`. To change the ordering, you would need to modify the sort key lambda in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) or implement client-side reordering of the `results_by_repository` array after receiving the response.

### Where does the repository name parsing occur in the source code?

Repository name extraction happens within the `_format_grep_response` function in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py), specifically using `hit.get("repo", {}).get("raw", "Unknown")` to handle missing metadata gracefully and default to "Unknown" when repository information is unavailable.