How Relevance‑Weighted Cutting Is Implemented for Overflowing Content in ai‑job‑search

Relevance‑weighted cutting is implemented in the search_company() function in salary_lookup.py by scoring each match (0‑100), sorting by descending score, and filtering against a hard min_score = 30 threshold.

The ai‑job‑search repository provides a salary lookup tool that handles potentially large result sets by trimming low‑relevance matches before they overflow the output. This relevance‑weighted cutting mechanism ensures users receive only the most pertinent company matches without manual pagination or arbitrary limits.

The Relevance Scoring Pipeline

The cutting mechanism rests on three sequential steps: relevance scoring, descending sort, and threshold-based filtering.

Step 1: Compute Per‑Entry Relevance Scores

Each candidate entry receives a normalized score via match_score_optimized(). This function combines string matching, anglicised word variants, and word‑overlap heuristics to produce a score between 0 and 100.

The scoring logic lives in salary_lookup.py:

def match_score_optimized(query, text):
    # Normalized matching with anglicised forms and word overlap

    # Returns integer score 0-100

Higher scores indicate stronger relevance to the query.

Step 2: Sort by Relevance

After scoring, entries are sorted to prioritize high‑relevance results. The sort key uses negative score for descending order, with company name as tiebreaker:

scored.sort(key=lambda x: (-x[0], x[1]["company"]))

This ensures the most relevant matches appear first while maintaining deterministic ordering for equal scores.

Step 3: Apply the Relevance Cut‑Off

The actual "cut" occurs through a configurable hard threshold. Only entries meeting the minimum score survive:

min_score = 30
return [entry for score, entry in scored if score >= min_score]

This three‑line filter in search_company() implements relevance‑weighted cutting by discarding all entries below the relevance floor, effectively preventing result set overflow.

Core Implementation in search_company()

The complete relevance‑weighted cutting logic is encapsulated in salary_lookup.py. Here is the annotated structure:

def search_company(data, query, city=None):
    scored = []
    # Score every candidate entry

    for entry in data:
        score = match_score_optimized(query, entry["company"])
        if city and entry.get("city") != city:
            continue  # Optional city filter applied before scoring

        scored.append((score, entry))
    
    # Sort by relevance descending, then alphabetically

    scored.sort(key=lambda x: (-x[0], x[1]["company"]))
    
    # Relevance‑weighted cutting: keep only above-threshold matches

    min_score = 30
    return [entry for score, entry in scored if score >= min_score]

The min_score = 30 constant provides the cutting boundary—roughly the bottom 70% of the possible score range is discarded.

Practical Usage Examples

Basic Search with Automatic Cutting

from salary_lookup import load_data, search_company

data = load_data()
results = search_company(data, "Microsoft")

print(f"Retained {len(results)} high‑relevance matches")
for entry in results[:5]:  # Top 5 after relevance cutting

    print(f"{entry['company']}: {entry.get('city', 'N/A')}")

Custom Threshold for Different Cutting Strictness

Copy the function and adjust min_score for looser or stricter cutting:

def search_company_lenient(data, query, city=None, min_score=15):
    scored = [(match_score_optimized(query, e["company"]), e) 
              for e in data 
              if not city or e.get("city") == city]
    scored.sort(key=lambda x: (-x[0], x[1]["company"]))
    return [entry for score, entry in scored if score >= min_score]

Lower values retain more results; higher values enforce stricter relevance‑weighted cutting.

Key Source Files

File Purpose
salary_lookup.py Contains match_score_optimized(), search_company(), and the min_score = 30 cutting logic
tests/test_salary_lookup.py Unit tests verifying sorted descending order and threshold behavior
tools/README_SALARY_TOOL.md Documentation describing relevance matching and result filtering

Relationship to Overflow Prevention

The relevance‑weighted cutting approach solves overflow differently than pagination or hard result limits. By coupling output size to match quality, the system:

  • Preserves high‑relevance matches regardless of total dataset size
  • Automatically shrinks result sets when queries are vague (few high scores)
  • Expands result sets only when many strong matches exist

This quality‑adaptive behavior makes min_score more principled than a fixed LIMIT clause.

Summary

  • Relevance scoring: match_score_optimized() in salary_lookup.py assigns 0‑100 scores based on query‑text similarity
  • Descending sort: Results ordered by (-score, company_name) prioritize best matches
  • Threshold cutting: min_score = 30 filter removes low‑relevance entries before return
  • Configurable: The cutting boundary can be adjusted by modifying the min_score constant or wrapping the function

Frequently Asked Questions

What happens if no matches exceed the 30‑point threshold?

search_company() returns an empty list. The calling code should handle this as a "no relevant matches found" condition rather than an error.

Can the relevance threshold be changed without modifying source code?

Not directly—the min_score is hardcoded in salary_lookup.py. However, you can import the scoring function and reimplement the cutting logic with a custom threshold, as shown in the lenient search example above.

How does match_score_optimized() handle spelling variations?

The function normalises both query and target text, applies anglicised forms (e.g., "résumé" → "resume"), and computes word overlap—enabling robust matching despite minor spelling differences without requiring exact string equality.

Is there a performance cost to scoring all entries before cutting?

The O(n log n) sort dominates for large datasets. The cutting filter itself is O(n). For typical company datasets, this overhead is negligible; for very large collections, pre‑filtering by city or other attributes before scoring reduces the candidate pool.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →