Extracting, Matching, and Labeling Keypoints with Fuzzy String Matching in Prettymaps
Prettymaps uses the match_keypoint function in prettymaps/draw.py to find the best-matching point of interest from a text query using thefuzz.token_sort_ratio scoring.
This guide explains how marceloprates/prettymaps implements fuzzy string matching to locate and label geographic keypoints without requiring exact name matches. Whether you're highlighting landmarks on a city map or building custom POI visualizations, the match_keypoint helper streamlines the workflow.
How Keypoint Matching Works in Prettymaps
The keypoint matching system follows a clean four-step pipeline. When you call prettymaps.plot() with a keypoints argument, the library builds a pandas DataFrame where each row represents a point of interest with a required name column.
The matching logic lives entirely in prettymaps/draw.py at lines 71–99. Here's the breakdown:
Step 1: Data Cleaning
The function first sanitizes the input by removing rows without names and discarding empty geometries.
keypoints_df = keypoints_df.dropna(subset=["name"])
keypoints_df = keypoints_df[~keypoints_df.geometry.is_empty]
This ensures only valid, named points proceed to scoring. Source: lines 85–88.
Step 2: Fuzzy Scoring with token_sort_ratio
Each remaining name gets compared against your query using thefuzz.fuzz.token_sort_ratio.
keypoints_df.loc[:, "match"] = keypoints_df.name.apply(
lambda x: fuzz.token_sort_ratio(x, query)
)
The token_sort_ratio comparator normalizes word order before scoring, so "Central Park" and "Park Central" receive high similarity. Source: lines 90–92.
Step 3: Select Top Matches
Rows with the maximum score are retained.
matches = keypoints_df[keypoints_df.match == max(keypoints_df.match)]
Source: line 95.
Step 4: Return Requested Match Index
The function returns the row at your specified index (default 0), capped to the available results.
return matches.iloc[[min(index, len(matches) - 1)]]
Source: line 98.
The match_keypoint Function Signature
Located at line 71 in prettymaps/draw.py, the function signature is:
def match_keypoint(keypoints_df: GeoDataFrame, query: str, index: int = 0) -> GeoDataFrame:
"""
Find best matching keypoint from query
...
"""
Parameters:
keypoints_df— A GeoDataFrame withnameandgeometrycolumnsquery— Free-text string to match against keypoint namesindex— Which matching result to return (0 for best match, 1 for second-best, etc.)
Returns: A one-row GeoDataFrame containing the selected keypoint.
Practical Code Examples
Basic Keypoint Query
import prettymaps as pm
import geopandas as gp
from shapely import Point
# Define custom landmarks
keypoints = {
"name": ["Central Park", "Statue of Liberty", "Brooklyn Bridge"],
"geometry": [
Point(-73.9654, 40.7829),
Point(-74.0445, 40.6892),
Point(-73.9969, 40.7061),
],
}
gdf = gp.GeoDataFrame(keypoints).set_geometry("geometry")
# Find match for loose query
matched = pm.match_keypoint(gdf, query="statue of libertyy") # typo intentional
print(matched.name.iloc[0]) # → "Statue of Liberty"
Plotting with Matched Keypoints
import prettymaps as pm
# Match and visualize in one flow
keypoints = pm.match_keypoint(gdf, query="central park")
pm.plot(
"Manhattan, New York, USA",
keypoints=keypoints.to_dict('list'),
style={
"keypoints": {
"ms": 20, # marker size
"mec": "crimson", # marker edge color
"mfc": "gold", # marker face color
}
},
)
Handling Multiple Matches
# Get the second-best match for ambiguous queries
second_match = pm.match_keypoint(gdf, query="park", index=1)
Why token_sort_ratio?
thefuzz (formerly fuzzywuzzy) provides several comparators. Prettymaps specifically chose token_sort_ratio because it:
- Normalizes word order — "Empire State Building" matches "Building, Empire State"
- Ignores case and punctuation — robust against messy input data
- Performs well on short phrases — ideal for POI names
Alternative comparators like ratio (strict sequence matching) or partial_ratio (substring matching) would fail on common real-world variations in geographic naming.
Integration with the Prettymaps Pipeline
The match_keypoint function integrates seamlessly with prettymaps.plot(). When you pass a keypoints dictionary, the plotting engine:
- Validates geometry types in
prettymaps/draw.py - Applies any style configuration from the
styleparameter - Renders markers using matplotlib's scatter plot behind the map layers
For advanced use cases, combine keypoint matching with GPX track loading (prettymaps/gpx.py) to align landmarks with recorded routes.
Key Files and References
| File | Purpose | Critical Lines |
|---|---|---|
prettymaps/draw.py |
Core plotting and matching logic | 71–99 (match_keypoint) |
prettymaps/utils.py |
Geometry and logging utilities | Helper decorators |
prettymaps/gpx.py |
Track file handling | Route-based queries |
docs/api.md |
Official API documentation | Full parameter reference |
Direct source: github.com/marceloprates/prettymaps/blob/main/prettymaps/draw.py#L71-L99
Summary
match_keypointinprettymaps/draw.pyenables fuzzy keypoint lookup without exact name matches- Uses
thefuzz.fuzz.token_sort_ratiofor word-order-tolerant scoring - Returns a GeoDataFrame ready for plotting or further processing
- Handles data cleaning (null names, empty geometries) automatically
- Supports index-based selection when multiple matches exist
Frequently Asked Questions
How does Prettymaps handle typos in keypoint queries?
The token_sort_ratio comparator from thefuzz is inherently tolerant of minor typos and word variations. For example, "statue of libertyy" still matches "Statue of Liberty" with high confidence because the algorithm focuses on token presence rather than exact character alignment.
Can I use match_keypoint without plotting a map?
Yes. match_keypoint operates independently of the visualization pipeline. Import and use it directly for any geographic data matching task—just pass a GeoDataFrame with name and geometry columns.
What happens if multiple keypoints have the same match score?
The function returns all top-scoring matches, then selects by your index parameter. If you request index=0 (default), you get the first of the tied results; index=1 gives the second, and so on, capped automatically to available matches.
Does Prettymaps require thefuzz as a dependency?
Yes. The repository specifies thefuzz in its dependency list for fuzzy string matching functionality. Install with pip install prettymaps to pull all required packages including thefuzz and python-Levenshtein for performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →