# How grep-mcp Extracts Line Numbers from HTML Search Result Snippets: A Deep Dive into the data-line Attribute

> Discover how grep-mcp extracts line numbers from HTML search results by parsing the data-line attribute. Understand the regex used for reliable extraction. Learn more!

- Repository: [gal peretz/grep-mcp](https://github.com/galprz/grep-mcp)
- Tags: deep-dive
- Published: 2026-02-16

---

**TLDR:** `grep-mcp` extracts line numbers from HTML search result snippets by parsing the `data-line` attribute embedded in `<span>` elements returned by the grep.app API, using a targeted regular expression in the `_extract_line_numbers` helper function.

The `galprz/grep-mcp` repository implements a Model Context Protocol (MCP) server that interfaces with the grep.app API to search code across millions of repositories. A critical challenge when processing these search results is reliably extracting original source line numbers from the HTML snippets that the API returns. This article examines how the codebase solves this problem through deterministic attribute parsing rather than fragile text extraction.

## Understanding the grep.app API HTML Structure

When the grep.app API returns search results, it provides matches as HTML snippets where each line of code is wrapped in semantic markup. Rather than displaying raw text, the API encloses line content in `<span>` elements (or similar tags) that carry a **`data-line`** attribute containing the integer line number from the original source file.

This approach decouples the visual presentation from metadata, allowing syntax highlighters to apply CSS classes while preserving precise location data. For example, a snippet might contain `<span class="hljs-keyword" data-line="42">def</span>`, where `42` represents the actual line number in the repository file.

## The Extraction Logic in server.py

The core extraction functionality resides in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py), which implements a two-stage pipeline: first parsing individual HTML snippets to isolate line numbers, then aggregating these into the final JSON response structure.

### The _extract_line_numbers Helper Function

Located at lines 45-52, the `_extract_line_numbers` function performs the actual parsing using a regular expression designed to match the `data-line` attribute pattern:

```python
def _extract_line_numbers(html_snippet: str) -> list[int]:
    """
    Extract line numbers from data-line attributes in HTML snippets.
    """
    pattern = r'data-line="(\d+)"'
    matches = re.findall(pattern, html_snippet)
    return [int(match) for match in matches]

```

This implementation uses `re.findall` to capture all digit sequences within `data-line="..."` constructs, converting each match to an integer. The regex is intentionally simple—focusing solely on the attribute name and value—to ensure robustness against variations in surrounding HTML, such as additional attributes or changing CSS class names.

### Integration into the Response Pipeline

At lines 308-312, the server invokes this helper when formatting API responses for the `grep_query` tool. As the code iterates through search results, it processes each file's `html_snippet` field:

```python

# Within the response formatting logic

line_numbers = _extract_line_numbers(html_snippet)
file_entry = {
    "file_path": file_path,
    "line_numbers": line_numbers,
    # ... other fields

}

```

This integration ensures that every file entry in the final JSON payload contains a clean, deduplicated list of line numbers corresponding to the matches found in that file.

## Practical Implementation Examples

You can leverage this extraction logic directly or through the public API interface.

### Direct Use of the Extractor

For custom processing of raw grep.app HTML responses:

```python
from grep_mcp.server import _extract_line_numbers

html_snippet = (
    '<span class="hljs-keyword" data-line="42">def</span> '
    '<span class="hljs-title function_" data-line="42">my_func</span>('
    '<span data-line="42">)</span>:'
)

line_numbers = _extract_line_numbers(html_snippet)
print(line_numbers)  # → [42]

```

### Fetching Line Numbers via grep_query

The standard method for retrieving line numbers through the MCP tool:

```python
import asyncio
from grep_mcp.server import grep_query

async def demo():
    response_json = await grep_query("def hello_world", language="Python")
    
    for repo in response_json["results_by_repository"]:
        print(f"Repo: {repo['repository']}")
        for file in repo["files"]:
            print(f"  {file['file_path']} → lines {file['line_numbers']}")

asyncio.run(demo())

```

## Why This Approach Is Reliable

The `data-line` attribute extraction strategy provides several advantages over text-based parsing:

- **Deterministic Metadata**: The `data-line` attribute is explicitly designed to carry line number metadata, eliminating ambiguity from syntax highlighting markup or visible text formatting.
- **Regex Resilience**: The pattern `data-line="(\d+)"` targets only the specific attribute, ignoring surrounding CSS classes, HTML structure changes, or snippet truncation boundaries.
- **Performance**: Using `re.findall` with a simple capture group is computationally efficient, processing even large HTML snippets with minimal overhead.
- **Language Agnostic**: Since the extraction operates on HTML markup rather than code syntax, it works uniformly across all programming languages supported by grep.app.

## Summary

- `grep-mcp` extracts line numbers by parsing the `data-line` attribute from HTML snippets returned by the grep.app API.
- The `_extract_line_numbers` helper in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) (lines 45-52) uses a targeted regex to capture these values.
- This approach is integrated into the response pipeline at lines 308-312, ensuring every file entry includes accurate line number arrays.
- The method is robust against HTML structure changes and works across all programming languages.

## Frequently Asked Questions

### How does grep-mcp handle multiple matches on the same line?

When the grep.app API returns HTML snippets containing multiple matches on a single source line, the `data-line` attribute appears multiple times with the same value. The `_extract_line_numbers` function captures all instances, and the downstream logic in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) deduplicates these into a unique set of line numbers for the final response.

### What happens if the HTML structure changes in future API updates?

The extraction regex `data-line="(\d+)"` specifically targets the attribute name and value pattern, which is part of grep.app's public HTML contract for semantic markup. While CSS classes and surrounding elements may change, the `data-line` attribute is designed as stable metadata. If the API were to remove this attribute, the current implementation would return an empty list rather than incorrect data, making failures detectable.

### Can I use the line extraction logic for other HTML-based code search APIs?

Yes, the `_extract_line_numbers` function is generic and can process any HTML string containing `data-line` attributes with numeric values. To adapt it for other APIs, ensure the target service embeds line numbers in a similar attribute format. If the attribute name differs (e.g., `data-linenumber`), you would modify the regex pattern in the function to match the specific attribute name used by that API.

### Does grep-mcp validate that extracted line numbers are positive integers?

The current implementation in [`src/grep_mcp/server.py`](https://github.com/galprz/grep-mcp/blob/main/src/grep_mcp/server.py) relies on the regex pattern `\d+` to match only digit sequences, which inherently filters out negative signs, decimals, and non-numeric characters. The `int()` conversion then produces positive integers (or zero) from these matches. While there is no explicit bounds checking for extremely large values, the grep.app API guarantees that `data-line` attributes contain valid source line numbers, making additional validation unnecessary for typical use cases.