How `_format_grep_response` Transforms Raw grep.app API Data in galprz/grep-mcp
The _format_grep_response function in src/grep_mcp/server.py normalizes raw HTML-rich grep.app search results into a structured JSON payload with syntax-highlighted code snippets, limited to 10 hits and grouped by repository.
When building an MCP (Model Context Protocol) server that interfaces with the grep.app search engine, raw API responses contain messy HTML snippets and scattered metadata. The _format_grep_response function serves as the critical transformation layer that converts this unstructured data into a clean, consumer-friendly format that downstream AI tools can render directly.
What _format_grep_response Does
Located at src/grep_mcp/server.py (lines 66-127), this function acts as a data pipeline that performs eight distinct operations on the raw grep.app JSON:
- Summarizes search metadata (total hits, top languages, top repositories)
- Limits result volume to the first 10 hits to prevent token overflow
- Sanitizes HTML content from code snippets
- Extracts precise line numbers from
data-lineattributes - Detects programming languages from file extensions
- Formats code blocks with proper fencing and truncation
- Groups results by repository for organized display
- Sorts repositories by match count (descending)
Step-by-Step Transformation Process
Extracting Search Summary Statistics
The function first pulls aggregate data from the API's facets object to give users immediate context about their search scope.
facets = data.get("facets", {})
total_count = facets.get("count", 0)
# Extract top 5 languages
languages = [
{"name": bucket["key"], "count": bucket["count"]}
for bucket in facets.get("lang", {}).get("buckets", [])[:5]
]
This creates a summary block containing total_results, languages_found, and repositories_found.
Limiting and Processing Individual Hits
To maintain concise output suitable for LLM context windows, the function enforces a hard limit of 10 results:
result_limit = 10
hits = data.get("hits", {}).get("hits", [])
limited_hits = hits[:result_limit]
Each hit contains repository metadata (repo, path, branch) and raw HTML content that requires cleaning.
Cleaning HTML and Extracting Line Numbers
Raw grep.app snippets contain HTML markup with syntax highlighting. The function delegates cleaning to _extract_text_from_html (which strips tags and decodes entities) and _extract_line_numbers (which parses data-line attributes):
raw_html = content.get("snippet", "")
clean_code = _extract_text_from_html(raw_html)
line_numbers = _extract_line_numbers(raw_html)
This produces plain text code with an array of corresponding line numbers (e.g., [42, 45, 78]).
Language Detection and Code Formatting
The function determines the programming language by analyzing the file extension through _get_language_from_extension, then passes the cleaned code to _format_code_snippet:
language = _get_language_from_extension(path)
formatted_snippet = _format_code_snippet(clean_code, language)
_format_code_snippet performs three critical operations:
- Truncates snippets exceeding 1000 characters
- Removes leading/trailing empty lines
- Limits display to 8 lines maximum
- Wraps the result in markdown code fences (e.g.,
```python)
Grouping and Sorting by Repository
Finally, the function organizes hits into repository buckets and sorts them by relevance:
# Group by repository
repo_groups = defaultdict(list)
for hit in processed_hits:
repo_groups[hit["repository"]].append(hit)
# Build sorted results
results_by_repo = []
for repo_name, files in repo_groups.items():
total_matches = sum(f["total_matches"] for f in files)
results_by_repo.append({
"repository": repo_name,
"matches_count": total_matches,
"files": files
})
# Sort by match count descending
results_by_repo.sort(key=lambda x: x["matches_count"], reverse=True)
The final payload combines the query, summary statistics, and sorted repository groups into a single structured JSON object.
Code Example: Using the Formatted Response
When calling the MCP tool, the formatted response arrives as a pretty-printed JSON string ready for immediate use:
import json
# Example response from grep_query tool
response = await grep_query(
query="asyncio.run",
language="python",
repo="python/cpython"
)
# Parse the structured payload
payload = json.loads(response)
# Access summary statistics
print(f"Total matches: {payload['summary']['total_results']}")
print(f"Top language: {payload['summary']['languages_found'][0]['name']}")
# Iterate through repository groups
for repo in payload['results_by_repository']:
print(f"\nRepository: {repo['repository']} ({repo['matches_count']} matches)")
# Access formatted code snippets
for file in repo['files']:
print(f" File: {file['file_path']}")
print(f" Lines: {file['line_numbers']}")
# The code_snippet is already markdown-formatted
print(file['code_snippet'])
The code_snippet field arrives pre-wrapped in markdown code fences (e.g., ```python), making it immediately renderable in any markdown-compatible interface.
Summary
_format_grep_responseinsrc/grep_mcp/server.pyis the central transformation engine that converts raw grep.app HTML into structured JSON.- The function limits results to 10 hits to prevent context window overflow in LLM applications.
- It sanitizes HTML content, extracts line numbers from
data-lineattributes, and detects languages from file extensions. - Code snippets are formatted with markdown fencing, truncated to 8 lines maximum, and wrapped in language-specific code blocks.
- Results are grouped by repository and sorted by match count (descending) to prioritize the most relevant sources.
Frequently Asked Questions
Where is the _format_grep_response function defined?
The _format_grep_response function is defined in src/grep_mcp/server.py within the galprz/grep-mcp repository. It operates on the raw JSON response returned by the grep.app search API and is called internally by the grep_query MCP tool implementation.
How many search results does _format_grep_response return?
The function enforces a hard limit of 10 results (result_limit = 10) to ensure the output remains concise and suitable for LLM context windows. This limit is applied immediately after fetching the raw hits from the API response, before any HTML processing or formatting occurs.
How does _format_grep_response handle HTML in code snippets?
Raw grep.app responses contain HTML-wrapped code snippets with syntax highlighting. The function delegates HTML cleaning to _extract_text_from_html, which strips all HTML tags and decodes HTML entities to produce plain text code. Simultaneously, _extract_line_numbers parses data-line attributes from the original HTML to preserve accurate line number references.
What determines the language detection in _format_grep_response?
Language detection is performed by _get_language_from_extension, which analyzes the file extension from the path field in each hit. The function maps common extensions (e.g., .py, .js, .go) to Prism-compatible language identifiers. This detected language is then passed to _format_code_snippet, which wraps the cleaned code in the appropriate markdown code fence (e.g., ```python).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →