Why Do Express Queries Return 0 Hits in Search Quality Benchmarks?

Express queries return zero hits because the benchmark's keyword search indexes full identifiers, but Express uses hyphen-separated module-pattern names that tokenize differently than how queries are matched.

In the tirth8205/code-review-graph repository, the search quality evaluation reveals a specific failure mode for Express-related codebases. While most language frameworks produce simple identifiers like UserService.getUser, Express generates node names following a distinctive naming convention that breaks the default search matching logic.

How the Keyword Search Index Works

The benchmark's search component operates in two stages: indexing and querying.

First, it extracts node identifiers—function names, class names, and module names—from the codebase. These identifiers feed into a plain-keyword lookup index. For a typical JavaScript project, this works as expected.

However, the tokenization process splits identifiers on hyphens. An identifier like express-create-application-callees becomes the tokens express, create, application, and callees in the index.

The Module-Pattern Naming Problem

Express follows a module-pattern naming convention that prefixes generated nodes with the package name and hyphen-separated descriptors:


express-create-application-callees
express-router-get-handler
express-middleware-chain

The critical mismatch occurs at query time. The query engine searches for the full identifier string, not the individual tokens, unless the query explicitly contains the complete hyphenated pattern. A plain query like express does not match any of these hyphen-joined identifiers.

As documented in README.md at lines 285-286:

"Search quality (MRR 0.35): Keyword search finds the right result in the top-4 for most queries, but ranking needs improvement. Express queries return 0 hits due to module-pattern naming."

The evaluation CSVs in evaluate/results/express_impact_accuracy_2026-08-02.csv confirm this outcome: Express benchmarks yield no predicted files.

Root Causes of the Zero-Hit Behavior

Factor Explanation
Hybrid naming Express nodes use hyphen-joined module patterns (express-*).
Keyword matcher The search index matches whole identifiers, not split tokens.
No token fallback Without token-level fallback, simple queries cannot locate Express nodes.
Design choice Modules are deliberately excluded from the primary search index to reduce noise, aggravating the problem.

Reproducing the Search Behavior

The following commands demonstrate the issue and its workaround:


# Plain keyword search returns no results for Express

$ code-review-graph search express

# → (no output)

# Exact module-pattern query matches the generated node

$ code-review-graph search express-create-application-callees

# → found 1 node

# Token-aware fallback with the --token-search flag

$ code-review-graph search --token-search express

# → now returns Express-related nodes

The --token-search flag exists as a partial mitigation, but it is not enabled by default in the benchmark evaluation, preserving the zero-hit results for Express.

Files Involved in the Search Quality Issue

File Relevance
README.md Contains the benchmark summary and explicit statement about Express zero hits
docs/REPRODUCING.md Explains evaluation harness index building and benchmark execution
evaluate/results/express_impact_accuracy_2026-08-02.csv Shows concrete zero predicted files outcome

Summary

  • Express uses hyphen-separated module-pattern naming (express-*) that diverges from typical identifier formats
  • The search index tokenizes on hyphens but queries match full strings, creating a mismatch
  • Plain queries like express return 0 hits because they cannot match the complete hyphenated identifiers
  • The --token-search flag provides a workaround by enabling token-level matching
  • Benchmark results reflect this design limitation with MRR 0.35 and explicit zero-hit documentation

Frequently Asked Questions

Why doesn't the search index handle Express names automatically?

The search system prioritizes exact identifier matching for precision. The tokenization exists for index compression, not query expansion. Without explicit token-level search enabled, the engine cannot bridge the gap between split tokens and whole-identifier queries.

Can the benchmark be fixed to support Express queries?

Yes. Enabling --token-search by default in the evaluation harness would resolve the immediate zero-hit problem. Alternatively, modifying the indexing to store both full identifiers and token n-grams would provide fallback matching without flag dependence.

Does this affect other frameworks besides Express?

Any framework using hyphen-heavy naming conventions would encounter similar issues. The README specifically calls out Express because its module-pattern naming is particularly prevalent and consistent across the generated graph nodes.

What is MRR and why is it 0.35 in the benchmarks?

MRR (Mean Reciprocal Rank) measures how high the correct result appears in search results. A value of 0.35 indicates the correct file typically appears in the top 3-4 results for most queries. Express queries drag this down because they contribute zero to the numerator—no correct results exist when the search returns nothing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →