Extracting, Matching, and Labeling Keypoints with Fuzzy String Matching in Prettymaps

Prettymaps uses the match_keypoint function in prettymaps/draw.py to find the best-matching point of interest from a text query using thefuzz.token_sort_ratio scoring.

This guide explains how marceloprates/prettymaps implements fuzzy string matching to locate and label geographic keypoints without requiring exact name matches. Whether you're highlighting landmarks on a city map or building custom POI visualizations, the match_keypoint helper streamlines the workflow.

How Keypoint Matching Works in Prettymaps

The keypoint matching system follows a clean four-step pipeline. When you call prettymaps.plot() with a keypoints argument, the library builds a pandas DataFrame where each row represents a point of interest with a required name column.

The matching logic lives entirely in prettymaps/draw.py at lines 71–99. Here's the breakdown:

Step 1: Data Cleaning

The function first sanitizes the input by removing rows without names and discarding empty geometries.

keypoints_df = keypoints_df.dropna(subset=["name"])
keypoints_df = keypoints_df[~keypoints_df.geometry.is_empty]

This ensures only valid, named points proceed to scoring. Source: lines 85–88.

Step 2: Fuzzy Scoring with token_sort_ratio

Each remaining name gets compared against your query using thefuzz.fuzz.token_sort_ratio.

keypoints_df.loc[:, "match"] = keypoints_df.name.apply(
    lambda x: fuzz.token_sort_ratio(x, query)
)

The token_sort_ratio comparator normalizes word order before scoring, so "Central Park" and "Park Central" receive high similarity. Source: lines 90–92.

Step 3: Select Top Matches

Rows with the maximum score are retained.

matches = keypoints_df[keypoints_df.match == max(keypoints_df.match)]

Source: line 95.

Step 4: Return Requested Match Index

The function returns the row at your specified index (default 0), capped to the available results.

return matches.iloc[[min(index, len(matches) - 1)]]

Source: line 98.

The match_keypoint Function Signature

Located at line 71 in prettymaps/draw.py, the function signature is:

def match_keypoint(keypoints_df: GeoDataFrame, query: str, index: int = 0) -> GeoDataFrame:
    """
    Find best matching keypoint from query
    ...
    """

Parameters:

  • keypoints_df — A GeoDataFrame with name and geometry columns
  • query — Free-text string to match against keypoint names
  • index — Which matching result to return (0 for best match, 1 for second-best, etc.)

Returns: A one-row GeoDataFrame containing the selected keypoint.

Practical Code Examples

Basic Keypoint Query

import prettymaps as pm
import geopandas as gp
from shapely import Point

# Define custom landmarks

keypoints = {
    "name": ["Central Park", "Statue of Liberty", "Brooklyn Bridge"],
    "geometry": [
        Point(-73.9654, 40.7829),
        Point(-74.0445, 40.6892),
        Point(-73.9969, 40.7061),
    ],
}

gdf = gp.GeoDataFrame(keypoints).set_geometry("geometry")

# Find match for loose query

matched = pm.match_keypoint(gdf, query="statue of libertyy")  # typo intentional

print(matched.name.iloc[0])  # → "Statue of Liberty"

Plotting with Matched Keypoints

import prettymaps as pm

# Match and visualize in one flow

keypoints = pm.match_keypoint(gdf, query="central park")

pm.plot(
    "Manhattan, New York, USA",
    keypoints=keypoints.to_dict('list'),
    style={
        "keypoints": {
            "ms": 20,           # marker size

            "mec": "crimson",   # marker edge color

            "mfc": "gold",      # marker face color

        }
    },
)

Handling Multiple Matches


# Get the second-best match for ambiguous queries

second_match = pm.match_keypoint(gdf, query="park", index=1)

Why token_sort_ratio?

thefuzz (formerly fuzzywuzzy) provides several comparators. Prettymaps specifically chose token_sort_ratio because it:

  • Normalizes word order — "Empire State Building" matches "Building, Empire State"
  • Ignores case and punctuation — robust against messy input data
  • Performs well on short phrases — ideal for POI names

Alternative comparators like ratio (strict sequence matching) or partial_ratio (substring matching) would fail on common real-world variations in geographic naming.

Integration with the Prettymaps Pipeline

The match_keypoint function integrates seamlessly with prettymaps.plot(). When you pass a keypoints dictionary, the plotting engine:

  1. Validates geometry types in prettymaps/draw.py
  2. Applies any style configuration from the style parameter
  3. Renders markers using matplotlib's scatter plot behind the map layers

For advanced use cases, combine keypoint matching with GPX track loading (prettymaps/gpx.py) to align landmarks with recorded routes.

Key Files and References

File Purpose Critical Lines
prettymaps/draw.py Core plotting and matching logic 71–99 (match_keypoint)
prettymaps/utils.py Geometry and logging utilities Helper decorators
prettymaps/gpx.py Track file handling Route-based queries
docs/api.md Official API documentation Full parameter reference

Direct source: github.com/marceloprates/prettymaps/blob/main/prettymaps/draw.py#L71-L99

Summary

  • match_keypoint in prettymaps/draw.py enables fuzzy keypoint lookup without exact name matches
  • Uses thefuzz.fuzz.token_sort_ratio for word-order-tolerant scoring
  • Returns a GeoDataFrame ready for plotting or further processing
  • Handles data cleaning (null names, empty geometries) automatically
  • Supports index-based selection when multiple matches exist

Frequently Asked Questions

How does Prettymaps handle typos in keypoint queries?

The token_sort_ratio comparator from thefuzz is inherently tolerant of minor typos and word variations. For example, "statue of libertyy" still matches "Statue of Liberty" with high confidence because the algorithm focuses on token presence rather than exact character alignment.

Can I use match_keypoint without plotting a map?

Yes. match_keypoint operates independently of the visualization pipeline. Import and use it directly for any geographic data matching task—just pass a GeoDataFrame with name and geometry columns.

What happens if multiple keypoints have the same match score?

The function returns all top-scoring matches, then selects by your index parameter. If you request index=0 (default), you get the first of the tied results; index=1 gives the second, and so on, capped automatically to available matches.

Does Prettymaps require thefuzz as a dependency?

Yes. The repository specifies thefuzz in its dependency list for fuzzy string matching functionality. Install with pip install prettymaps to pull all required packages including thefuzz and python-Levenshtein for performance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →