How grep-mcp Extracts and Displays Programming Language and Repository Statistics

grep-mcp parses the facets object from the grep.app API response to extract top programming languages and repositories, then aggregates individual code hits into a structured JSON summary using the _format_grep_response helper in src/grep_mcp/server.py.

The galprz/grep-mcp repository provides a Model Context Protocol (MCP) server that interfaces with the grep.app search API. When you query for code patterns, the tool extracts and displays statistics for programming languages and repositories by transforming raw API data into a concise, human-readable format. The implementation centers on parsing facet buckets and grouping results by repository to deliver actionable insights.

Understanding the grep.app API Response Structure

The grep.app API returns a JSON payload containing a "facets" object that pre-aggregates search metadata. This object includes grouped counts for programming languages ("lang") and repositories ("repo"), along with an overall hit count.

facets = data.get("facets", {})
total_count = facets.get("count", 0)
lang_buckets = facets.get("lang", {}).get("buckets", [])
repo_buckets = facets.get("repo", {}).get("buckets", [])

Each bucket in lang_buckets and repo_buckets contains a "val" field (the language or repository name) and a "count" field (the number of matches). The _format_grep_response function in src/grep_mcp/server.py (lines 66-90) processes these buckets to build the final statistics summary.

Extracting Top Programming Language Statistics

Parsing Language Facets from the API

The tool extracts language statistics by iterating over the lang_buckets list returned by the API. Each bucket represents a distinct programming language found in the search results, ordered by frequency.

languages = []
for bucket in lang_buckets[:5]:  # Keep only the five most common languages

    languages.append({
        "language": bucket.get("val", "Unknown"),
        "count": bucket.get("count", 0)
    })

Building the Top 5 Language List

The slicing operation [:5] ensures only the top five languages by match count are included in the final output. This prevents information overload while highlighting the most relevant technologies. The resulting list is stored under summary["top_languages"] in the final JSON response.

Extracting Top Repository Statistics

The repository extraction logic mirrors the language processing pattern. The _format_grep_response function processes repo_buckets to identify where code matches occur most frequently.

repositories = []
for bucket in repo_buckets[:5]:  # Keep only the five most common repos

    repositories.append({
        "repository": bucket.get("val", "Unknown"),
        "count": bucket.get("count", 0)
    })

These entries populate summary["top_repositories"], providing users with immediate insight into which GitHub repositories contain the most occurrences of their search pattern.

Aggregating Per-Repository Search Results

Processing Individual Code Hits

Beyond high-level statistics, grep-mcp aggregates individual search results by repository. The function processes the first 10 hits (result_limit = 10) to avoid overwhelming output while maintaining relevance.

For each hit, the tool extracts:

  • Repository name and file path
  • Branch information
  • Total matches within the file
  • Line numbers where matches occur
  • Code snippet with syntax highlighting

The language for syntax highlighting is determined by _get_language_from_extension (lines 55-63 in src/grep_mcp/server.py), which maps file extensions to language identifiers.

Grouping Results by Repository

Results are organized using a repo_groups dictionary that accumulates files under their respective repository keys. Each repository entry tracks the cumulative matches_count and maintains a list of file objects containing the detailed match information.

After processing all hits, the groups are converted to a list and sorted descending by matches_count to prioritize repositories with the highest match density. This sorted list becomes results_by_repository in the final output.

The Final Statistics Response Format

The _format_grep_response function assembles a structured JSON object that combines summary statistics with detailed results:

{
  "query": "<original query>",
  "summary": {
    "total_results": <total hits>,
    "results_shown": <number of files returned>,
    "repositories_found": <number of distinct repos>,
    "top_languages": [
      {"language": "Python", "count": 42},
      {"language": "JavaScript", "count": 15}
    ],
    "top_repositories": [
      {"repository": "owner/repo", "count": 25}
    ]
  },
  "results_by_repository": [...]
}

The JSON is serialized with json.dumps(..., indent=2) and returned to the MCP tool caller, providing both high-level statistical insights and granular code matches.

Implementation in src/grep_mcp/server.py

The core logic resides in the _format_grep_response helper function located at lines 66-90 of src/grep_mcp/server.py. This function orchestrates the extraction of language and repository statistics from the API facets and manages the grouping of individual search results.

def _format_grep_response(data: dict, query: str) -> str:
    """
    Transform raw grep.app API response into structured statistics.
    """
    facets = data.get("facets", {})
    total_count = facets.get("count", 0)
    
    # Extract top 5 languages

    lang_buckets = facets.get("lang", {}).get("buckets", [])
    top_languages = [
        {"language": b.get("val", "Unknown"), "count": b.get("count", 0)}
        for b in lang_buckets[:5]
    ]
    
    # Extract top 5 repositories

    repo_buckets = facets.get("repo", {}).get("buckets", [])
    top_repositories = [
        {"repository": b.get("val", "Unknown"), "count": b.get("count", 0)}
        for b in repo_buckets[:5]
    ]
    
    # ... aggregation logic for results_by_repository ...

    
    summary = {
        "total_results": total_count,
        "top_languages": top_languages,
        "top_repositories": top_repositories,
        # ... other fields ...

    }
    
    return json.dumps({"query": query, "summary": summary, ...}, indent=2)

The grep_query function (lines 51-52) serves as the public MCP tool entry point, validating parameters and invoking the API before delegating to _format_grep_response for statistical processing.

Summary

  • grep-mcp extracts statistics by parsing the "facets" object returned by the grep.app API, which pre-aggregates matches by language and repository.
  • The _format_grep_response function in src/grep_mcp/server.py (lines 66-90) processes these facets to generate top 5 lists for both programming languages and repositories.
  • Individual search results are aggregated by repository using a grouping dictionary, sorted by match count, and limited to the first 10 hits to balance detail with readability.
  • The final output combines high-level summary statistics with granular, syntax-highlighted code snippets in a structured JSON format suitable for MCP tool consumers.

Frequently Asked Questions

How does grep-mcp determine the top 5 programming languages?

The tool extracts the "lang" buckets from the API's "facets" object, which contains pre-sorted language counts. It slices the first five entries using lang_buckets[:5] and maps each bucket's "val" (language name) and "count" (match count) into a structured list stored under summary["top_languages"].

What is the difference between top_repositories and results_by_repository?

top_repositories provides a high-level statistical summary of the five repositories with the most matches across the entire search, derived from the API's "repo" facets. In contrast, results_by_repository contains the detailed, granular search results—including file paths, line numbers, and code snippets—grouped by repository and sorted by total match count within each repository.

How does grep-mcp identify the language for syntax highlighting in code snippets?

The tool uses the _get_language_from_extension helper function located at lines 55-63 in src/grep_mcp/server.py. This function maps file extensions extracted from the file path to language identifiers (e.g., .py to python, .js to javascript), which are then used to format the code snippet with appropriate markdown code fences.

Where is the statistics formatting logic implemented?

The core statistics extraction and formatting logic resides in the _format_grep_response function within src/grep_mcp/server.py (lines 66-90). This private helper method orchestrates the parsing of language and repository facets, limits results to the top five entries for each category, aggregates individual hits by repository, and assembles the final structured JSON response.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →