# How `_format_grep_response` Transforms Raw grep.app API Data in galprz/grep-mcp

> Discover how _format_grep_response transforms raw grep.app API data into structured, syntax-highlighted JSON. Learn about code snippet normalization and grouping by repository.

- Repository: [gal peretz/grep-mcp](https://github.com/galprz/grep-mcp)
- Tags: how-to-guide
- Published: 2026-02-16

---

**The `_format_grep_response` function in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) normalizes raw HTML-rich grep.app search results into a structured JSON payload with syntax-highlighted code snippets, limited to 10 hits and grouped by repository.**

When building an MCP (Model Context Protocol) server that interfaces with the grep.app search engine, raw API responses contain messy HTML snippets and scattered metadata. The `_format_grep_response` function serves as the critical transformation layer that converts this unstructured data into a clean, consumer-friendly format that downstream AI tools can render directly.

## What `_format_grep_response` Does

Located at **[`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py)** (lines 66-127), this function acts as a data pipeline that performs eight distinct operations on the raw grep.app JSON:

1. **Summarizes search metadata** (total hits, top languages, top repositories)
2. **Limits result volume** to the first 10 hits to prevent token overflow
3. **Sanitizes HTML content** from code snippets
4. **Extracts precise line numbers** from `data-line` attributes
5. **Detects programming languages** from file extensions
6. **Formats code blocks** with proper fencing and truncation
7. **Groups results by repository** for organized display
8. **Sorts repositories** by match count (descending)

## Step-by-Step Transformation Process

### Extracting Search Summary Statistics

The function first pulls aggregate data from the API's `facets` object to give users immediate context about their search scope.

```python
facets = data.get("facets", {})
total_count = facets.get("count", 0)

# Extract top 5 languages

languages = [
    {"name": bucket["key"], "count": bucket["count"]}
    for bucket in facets.get("lang", {}).get("buckets", [])[:5]
]

```

This creates a `summary` block containing `total_results`, `languages_found`, and `repositories_found`.

### Limiting and Processing Individual Hits

To maintain concise output suitable for LLM context windows, the function enforces a hard limit of 10 results:

```python
result_limit = 10
hits = data.get("hits", {}).get("hits", [])
limited_hits = hits[:result_limit]

```

Each hit contains repository metadata (`repo`, `path`, `branch`) and raw HTML content that requires cleaning.

### Cleaning HTML and Extracting Line Numbers

Raw grep.app snippets contain HTML markup with syntax highlighting. The function delegates cleaning to `_extract_text_from_html` (which strips tags and decodes entities) and `_extract_line_numbers` (which parses `data-line` attributes):

```python
raw_html = content.get("snippet", "")
clean_code = _extract_text_from_html(raw_html)
line_numbers = _extract_line_numbers(raw_html)

```

This produces plain text code with an array of corresponding line numbers (e.g., `[42, 45, 78]`).

### Language Detection and Code Formatting

The function determines the programming language by analyzing the file extension through `_get_language_from_extension`, then passes the cleaned code to `_format_code_snippet`:

```python
language = _get_language_from_extension(path)
formatted_snippet = _format_code_snippet(clean_code, language)

```

`_format_code_snippet` performs three critical operations:
1. **Truncates** snippets exceeding 1000 characters
2. **Removes** leading/trailing empty lines
3. **Limits** display to 8 lines maximum
4. **Wraps** the result in markdown code fences (e.g., ` ```python `)

### Grouping and Sorting by Repository

Finally, the function organizes hits into repository buckets and sorts them by relevance:

```python

# Group by repository

repo_groups = defaultdict(list)
for hit in processed_hits:
    repo_groups[hit["repository"]].append(hit)

# Build sorted results

results_by_repo = []
for repo_name, files in repo_groups.items():
    total_matches = sum(f["total_matches"] for f in files)
    results_by_repo.append({
        "repository": repo_name,
        "matches_count": total_matches,
        "files": files
    })

# Sort by match count descending

results_by_repo.sort(key=lambda x: x["matches_count"], reverse=True)

```

The final payload combines the query, summary statistics, and sorted repository groups into a single structured JSON object.

## Code Example: Using the Formatted Response

When calling the MCP tool, the formatted response arrives as a pretty-printed JSON string ready for immediate use:

```python
import json

# Example response from grep_query tool

response = await grep_query(
    query="asyncio.run",
    language="python",
    repo="python/cpython"
)

# Parse the structured payload

payload = json.loads(response)

# Access summary statistics

print(f"Total matches: {payload['summary']['total_results']}")
print(f"Top language: {payload['summary']['languages_found'][0]['name']}")

# Iterate through repository groups

for repo in payload['results_by_repository']:
    print(f"\nRepository: {repo['repository']} ({repo['matches_count']} matches)")
    
    # Access formatted code snippets

    for file in repo['files']:
        print(f"  File: {file['file_path']}")
        print(f"  Lines: {file['line_numbers']}")
        # The code_snippet is already markdown-formatted

        print(file['code_snippet'])

```

The `code_snippet` field arrives pre-wrapped in markdown code fences (e.g., ` ```python `), making it immediately renderable in any markdown-compatible interface.

## Summary

- **`_format_grep_response`** in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) is the central transformation engine that converts raw grep.app HTML into structured JSON.
- The function **limits results to 10 hits** to prevent context window overflow in LLM applications.
- It **sanitizes HTML** content, **extracts line numbers** from `data-line` attributes, and **detects languages** from file extensions.
- Code snippets are **formatted with markdown fencing**, truncated to 8 lines maximum, and wrapped in language-specific code blocks.
- Results are **grouped by repository** and **sorted by match count** (descending) to prioritize the most relevant sources.

## Frequently Asked Questions

### Where is the `_format_grep_response` function defined?

The `_format_grep_response` function is defined in **[`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py)** within the `galprz/grep-mcp` repository. It operates on the raw JSON response returned by the grep.app search API and is called internally by the `grep_query` MCP tool implementation.

### How many search results does `_format_grep_response` return?

The function enforces a hard limit of **10 results** (`result_limit = 10`) to ensure the output remains concise and suitable for LLM context windows. This limit is applied immediately after fetching the raw hits from the API response, before any HTML processing or formatting occurs.

### How does `_format_grep_response` handle HTML in code snippets?

Raw grep.app responses contain HTML-wrapped code snippets with syntax highlighting. The function delegates HTML cleaning to `_extract_text_from_html`, which strips all HTML tags and decodes HTML entities to produce plain text code. Simultaneously, `_extract_line_numbers` parses `data-line` attributes from the original HTML to preserve accurate line number references.

### What determines the language detection in `_format_grep_response`?

Language detection is performed by `_get_language_from_extension`, which analyzes the file extension from the `path` field in each hit. The function maps common extensions (e.g., `.py`, `.js`, `.go`) to Prism-compatible language identifiers. This detected language is then passed to `_format_code_snippet`, which wraps the cleaned code in the appropriate markdown code fence (e.g., ` ```python `).