How grep-mcp Extracts Line Numbers from HTML Search Result Snippets: A Deep Dive into the data-line Attribute
TLDR: grep-mcp extracts line numbers from HTML search result snippets by parsing the data-line attribute embedded in <span> elements returned by the grep.app API, using a targeted regular expression in the _extract_line_numbers helper function.
The galprz/grep-mcp repository implements a Model Context Protocol (MCP) server that interfaces with the grep.app API to search code across millions of repositories. A critical challenge when processing these search results is reliably extracting original source line numbers from the HTML snippets that the API returns. This article examines how the codebase solves this problem through deterministic attribute parsing rather than fragile text extraction.
Understanding the grep.app API HTML Structure
When the grep.app API returns search results, it provides matches as HTML snippets where each line of code is wrapped in semantic markup. Rather than displaying raw text, the API encloses line content in <span> elements (or similar tags) that carry a data-line attribute containing the integer line number from the original source file.
This approach decouples the visual presentation from metadata, allowing syntax highlighters to apply CSS classes while preserving precise location data. For example, a snippet might contain <span class="hljs-keyword" data-line="42">def</span>, where 42 represents the actual line number in the repository file.
The Extraction Logic in server.py
The core extraction functionality resides in src/grep_mcp/server.py, which implements a two-stage pipeline: first parsing individual HTML snippets to isolate line numbers, then aggregating these into the final JSON response structure.
The _extract_line_numbers Helper Function
Located at lines 45-52, the _extract_line_numbers function performs the actual parsing using a regular expression designed to match the data-line attribute pattern:
def _extract_line_numbers(html_snippet: str) -> list[int]:
"""
Extract line numbers from data-line attributes in HTML snippets.
"""
pattern = r'data-line="(\d+)"'
matches = re.findall(pattern, html_snippet)
return [int(match) for match in matches]
This implementation uses re.findall to capture all digit sequences within data-line="..." constructs, converting each match to an integer. The regex is intentionally simple—focusing solely on the attribute name and value—to ensure robustness against variations in surrounding HTML, such as additional attributes or changing CSS class names.
Integration into the Response Pipeline
At lines 308-312, the server invokes this helper when formatting API responses for the grep_query tool. As the code iterates through search results, it processes each file's html_snippet field:
# Within the response formatting logic
line_numbers = _extract_line_numbers(html_snippet)
file_entry = {
"file_path": file_path,
"line_numbers": line_numbers,
# ... other fields
}
This integration ensures that every file entry in the final JSON payload contains a clean, deduplicated list of line numbers corresponding to the matches found in that file.
Practical Implementation Examples
You can leverage this extraction logic directly or through the public API interface.
Direct Use of the Extractor
For custom processing of raw grep.app HTML responses:
from grep_mcp.server import _extract_line_numbers
html_snippet = (
'<span class="hljs-keyword" data-line="42">def</span> '
'<span class="hljs-title function_" data-line="42">my_func</span>('
'<span data-line="42">)</span>:'
)
line_numbers = _extract_line_numbers(html_snippet)
print(line_numbers) # → [42]
Fetching Line Numbers via grep_query
The standard method for retrieving line numbers through the MCP tool:
import asyncio
from grep_mcp.server import grep_query
async def demo():
response_json = await grep_query("def hello_world", language="Python")
for repo in response_json["results_by_repository"]:
print(f"Repo: {repo['repository']}")
for file in repo["files"]:
print(f" {file['file_path']} → lines {file['line_numbers']}")
asyncio.run(demo())
Why This Approach Is Reliable
The data-line attribute extraction strategy provides several advantages over text-based parsing:
- Deterministic Metadata: The
data-lineattribute is explicitly designed to carry line number metadata, eliminating ambiguity from syntax highlighting markup or visible text formatting. - Regex Resilience: The pattern
data-line="(\d+)"targets only the specific attribute, ignoring surrounding CSS classes, HTML structure changes, or snippet truncation boundaries. - Performance: Using
re.findallwith a simple capture group is computationally efficient, processing even large HTML snippets with minimal overhead. - Language Agnostic: Since the extraction operates on HTML markup rather than code syntax, it works uniformly across all programming languages supported by grep.app.
Summary
grep-mcpextracts line numbers by parsing thedata-lineattribute from HTML snippets returned by the grep.app API.- The
_extract_line_numbershelper insrc/grep_mcp/server.py(lines 45-52) uses a targeted regex to capture these values. - This approach is integrated into the response pipeline at lines 308-312, ensuring every file entry includes accurate line number arrays.
- The method is robust against HTML structure changes and works across all programming languages.
Frequently Asked Questions
How does grep-mcp handle multiple matches on the same line?
When the grep.app API returns HTML snippets containing multiple matches on a single source line, the data-line attribute appears multiple times with the same value. The _extract_line_numbers function captures all instances, and the downstream logic in src/grep_mcp/server.py deduplicates these into a unique set of line numbers for the final response.
What happens if the HTML structure changes in future API updates?
The extraction regex data-line="(\d+)" specifically targets the attribute name and value pattern, which is part of grep.app's public HTML contract for semantic markup. While CSS classes and surrounding elements may change, the data-line attribute is designed as stable metadata. If the API were to remove this attribute, the current implementation would return an empty list rather than incorrect data, making failures detectable.
Can I use the line extraction logic for other HTML-based code search APIs?
Yes, the _extract_line_numbers function is generic and can process any HTML string containing data-line attributes with numeric values. To adapt it for other APIs, ensure the target service embeds line numbers in a similar attribute format. If the attribute name differs (e.g., data-linenumber), you would modify the regex pattern in the function to match the specific attribute name used by that API.
Does grep-mcp validate that extracted line numbers are positive integers?
The current implementation in src/grep_mcp/server.py relies on the regex pattern \d+ to match only digit sequences, which inherently filters out negative signs, decimals, and non-numeric characters. The int() conversion then produces positive integers (or zero) from these matches. While there is no explicit bounds checking for extremely large values, the grep.app API guarantees that data-line attributes contain valid source line numbers, making additional validation unnecessary for typical use cases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →