# How grep-mcp Extracts and Displays Programming Language and Repository Statistics

> Learn how grep-mcp extracts and displays programming language and repository statistics by parsing the grep.app API and structuring JSON summaries.

- Repository: [gal peretz/grep-mcp](https://github.com/galprz/grep-mcp)
- Tags: how-to-guide
- Published: 2026-02-16

---

**grep-mcp parses the facets object from the grep.app API response to extract top programming languages and repositories, then aggregates individual code hits into a structured JSON summary using the `_format_grep_response` helper in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py).**

The `galprz/grep-mcp` repository provides a Model Context Protocol (MCP) server that interfaces with the grep.app search API. When you query for code patterns, the tool extracts and displays statistics for programming languages and repositories by transforming raw API data into a concise, human-readable format. The implementation centers on parsing facet buckets and grouping results by repository to deliver actionable insights.

## Understanding the grep.app API Response Structure

The grep.app API returns a JSON payload containing a `"facets"` object that pre-aggregates search metadata. This object includes grouped counts for programming languages (`"lang"`) and repositories (`"repo"`), along with an overall hit count.

```python
facets = data.get("facets", {})
total_count = facets.get("count", 0)
lang_buckets = facets.get("lang", {}).get("buckets", [])
repo_buckets = facets.get("repo", {}).get("buckets", [])

```

Each bucket in `lang_buckets` and `repo_buckets` contains a `"val"` field (the language or repository name) and a `"count"` field (the number of matches). The `_format_grep_response` function in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) (lines 66-90) processes these buckets to build the final statistics summary.

## Extracting Top Programming Language Statistics

### Parsing Language Facets from the API

The tool extracts language statistics by iterating over the `lang_buckets` list returned by the API. Each bucket represents a distinct programming language found in the search results, ordered by frequency.

```python
languages = []
for bucket in lang_buckets[:5]:  # Keep only the five most common languages

    languages.append({
        "language": bucket.get("val", "Unknown"),
        "count": bucket.get("count", 0)
    })

```

### Building the Top 5 Language List

The slicing operation `[:5]` ensures only the top five languages by match count are included in the final output. This prevents information overload while highlighting the most relevant technologies. The resulting list is stored under `summary["top_languages"]` in the final JSON response.

## Extracting Top Repository Statistics

The repository extraction logic mirrors the language processing pattern. The `_format_grep_response` function processes `repo_buckets` to identify where code matches occur most frequently.

```python
repositories = []
for bucket in repo_buckets[:5]:  # Keep only the five most common repos

    repositories.append({
        "repository": bucket.get("val", "Unknown"),
        "count": bucket.get("count", 0)
    })

```

These entries populate `summary["top_repositories"]`, providing users with immediate insight into which GitHub repositories contain the most occurrences of their search pattern.

## Aggregating Per-Repository Search Results

### Processing Individual Code Hits

Beyond high-level statistics, grep-mcp aggregates individual search results by repository. The function processes the first **10** hits (`result_limit = 10`) to avoid overwhelming output while maintaining relevance.

For each hit, the tool extracts:
- **Repository name** and **file path**
- **Branch** information
- **Total matches** within the file
- **Line numbers** where matches occur
- **Code snippet** with syntax highlighting

The language for syntax highlighting is determined by `_get_language_from_extension` (lines 55-63 in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py)), which maps file extensions to language identifiers.

### Grouping Results by Repository

Results are organized using a `repo_groups` dictionary that accumulates files under their respective repository keys. Each repository entry tracks the cumulative `matches_count` and maintains a list of file objects containing the detailed match information.

After processing all hits, the groups are converted to a list and sorted descending by `matches_count` to prioritize repositories with the highest match density. This sorted list becomes `results_by_repository` in the final output.

## The Final Statistics Response Format

The `_format_grep_response` function assembles a structured JSON object that combines summary statistics with detailed results:

```json
{
  "query": "<original query>",
  "summary": {
    "total_results": <total hits>,
    "results_shown": <number of files returned>,
    "repositories_found": <number of distinct repos>,
    "top_languages": [
      {"language": "Python", "count": 42},
      {"language": "JavaScript", "count": 15}
    ],
    "top_repositories": [
      {"repository": "owner/repo", "count": 25}
    ]
  },
  "results_by_repository": [...]
}

```

The JSON is serialized with `json.dumps(..., indent=2)` and returned to the MCP tool caller, providing both high-level statistical insights and granular code matches.

## Implementation in src/grep_mcp/server.py

The core logic resides in the `_format_grep_response` helper function located at lines 66-90 of [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py). This function orchestrates the extraction of language and repository statistics from the API facets and manages the grouping of individual search results.

```python
def _format_grep_response(data: dict, query: str) -> str:
    """
    Transform raw grep.app API response into structured statistics.
    """
    facets = data.get("facets", {})
    total_count = facets.get("count", 0)
    
    # Extract top 5 languages

    lang_buckets = facets.get("lang", {}).get("buckets", [])
    top_languages = [
        {"language": b.get("val", "Unknown"), "count": b.get("count", 0)}
        for b in lang_buckets[:5]
    ]
    
    # Extract top 5 repositories

    repo_buckets = facets.get("repo", {}).get("buckets", [])
    top_repositories = [
        {"repository": b.get("val", "Unknown"), "count": b.get("count", 0)}
        for b in repo_buckets[:5]
    ]
    
    # ... aggregation logic for results_by_repository ...

    
    summary = {
        "total_results": total_count,
        "top_languages": top_languages,
        "top_repositories": top_repositories,
        # ... other fields ...

    }
    
    return json.dumps({"query": query, "summary": summary, ...}, indent=2)

```

The `grep_query` function (lines 51-52) serves as the public MCP tool entry point, validating parameters and invoking the API before delegating to `_format_grep_response` for statistical processing.

## Summary

- **grep-mcp** extracts statistics by parsing the `"facets"` object returned by the grep.app API, which pre-aggregates matches by language and repository.
- The `_format_grep_response` function in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) (lines 66-90) processes these facets to generate **top 5** lists for both programming languages and repositories.
- Individual search results are aggregated by repository using a grouping dictionary, sorted by match count, and limited to the first 10 hits to balance detail with readability.
- The final output combines high-level summary statistics with granular, syntax-highlighted code snippets in a structured JSON format suitable for MCP tool consumers.

## Frequently Asked Questions

### How does grep-mcp determine the top 5 programming languages?

The tool extracts the `"lang"` buckets from the API's `"facets"` object, which contains pre-sorted language counts. It slices the first five entries using `lang_buckets[:5]` and maps each bucket's `"val"` (language name) and `"count"` (match count) into a structured list stored under `summary["top_languages"]`.

### What is the difference between top_repositories and results_by_repository?

`top_repositories` provides a high-level statistical summary of the five repositories with the most matches across the entire search, derived from the API's `"repo"` facets. In contrast, `results_by_repository` contains the detailed, granular search results—including file paths, line numbers, and code snippets—grouped by repository and sorted by total match count within each repository.

### How does grep-mcp identify the language for syntax highlighting in code snippets?

The tool uses the `_get_language_from_extension` helper function located at lines 55-63 in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py). This function maps file extensions extracted from the file path to language identifiers (e.g., `.py` to `python`, `.js` to `javascript`), which are then used to format the code snippet with appropriate markdown code fences.

### Where is the statistics formatting logic implemented?

The core statistics extraction and formatting logic resides in the `_format_grep_response` function within [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) (lines 66-90). This private helper method orchestrates the parsing of language and repository facets, limits results to the top five entries for each category, aggregates individual hits by repository, and assembles the final structured JSON response.