# How to Cache GeoDataFrames in Prettymaps for Faster Map Generation

> Speed up map generation with prettymaps by caching GeoDataFrames. Learn how this feature eliminates redundant requests and cuts down processing time significantly.

- Repository: [Marcelo de Oliveira Rosa Prates/prettymaps](https://github.com/marceloprates/prettymaps)
- Tags: performance
- Published: 2026-08-20

---

**Prettymaps provides a built-in file-system cache that stores OpenStreetMap GeoDataFrames as GeoJSON files, eliminating redundant network requests and reducing map generation time from minutes to milliseconds.**

Every time you generate a map with prettymaps, the library queries OpenStreetMap (OSM) and converts the response into **GeoDataFrames** (GDFs) using GeoPandas. Since OSM data changes slowly but styling experiments happen constantly, repeatedly downloading identical geometry wastes bandwidth and blocks your creative flow. The caching mechanism in [`prettymaps/fetch.py`](https://github.com/marceloprates/prettymaps/blob/main/prettymaps/fetch.py) solves this by persisting raw GDFs to disk using deterministic hash-based keys.

## How the Caching System Works

The cache implementation lives entirely in [`prettymaps/fetch.py`](https://github.com/marceloprates/prettymaps/blob/main/prettymaps/fetch.py) and operates through three coordinated functions: hash generation, disk I/O, and request integration.

### Hash Generation: Creating Deterministic Cache Keys

The cache key combines two hashed components to guarantee uniqueness for every distinct query.

```python

# From prettymaps/fetch.py – write_to_cache (lines 32-38)

perimeter_hash = hash(perimeter["geometry"].to_json())
kwargs_hash = hash(str(layer_kwargs))
hash_str = f"{perimeter_hash}_{kwargs_hash}"

```

- **perimeter_hash**: Derived from the geometry JSON of your area of interest (address, bounding box, or polygon)
- **kwargs_hash**: Captures layer-specific parameters including OSM tags, filters, and query modifiers

This dual-hash approach ensures that moving your map 100 meters east or swapping "building" tags for "highway" tags produces distinct cache entries.

### Writing GeoDataFrames to Disk

When fresh OSM data arrives, `write_to_cache()` serializes the GDF to GeoJSON format with logging suppressed for clean output.

```python

# From prettymaps/fetch.py – lines 41-45

gdf.to_file(cache_path, driver="GeoJSON")

```

Cached files land in `prettymaps_cache/<hash_str>.geojson` by default. The GeoJSON driver preserves coordinate reference systems and attribute tables without additional compression overhead.

### Reading from Cache Before Network Requests

The `read_from_cache()` function short-circuits expensive OSM queries when a matching file exists.

```python

# From prettymaps/fetch.py – lines 73-77

if os.path.exists(cache_path):
    return gp.read_file(cache_path)

```

GeoPandas reads the GeoJSON back into a functional GDF instantly, bypassing OSM's Overpass API entirely.

### Integration with the Fetch Pipeline

The high-level `get_gdfs()` entry point coordinates cache lookups through `unified_osm_request()`:

1. Generate hash from perimeter + layer kwargs
2. Attempt `read_from_cache()` — return immediately on hit
3. On miss: query OSM, receive GDF, call `write_to_cache()`, return result

While the current release has cache invocation commented in `unified_osm_request()` (lines 95-103), the helper functions remain fully functional and can be called directly or enabled with minor modifications.

### Legacy Decorator Pattern

A commented `cache_geometry` decorator (lines 75-111 in [`fetch.py`](https://github.com/marceloprates/prettymaps/blob/main/fetch.py)) demonstrates an alternative approach: wrapping any GDF-returning function to automatically handle cache read/write cycles. This pattern suits custom fetch pipelines or extended OSM query builders.

## Performance Benefits of Caching GeoDataFrames

| Benefit | Mechanism | Typical Impact |
|--------|-----------|--------------|
| **Network elimination** | Local file reads replace HTTP requests | Seconds to minutes → milliseconds |
| **Reproducible builds** | Identical inputs produce bit-identical outputs | CI pipelines succeed without connectivity |
| **Concurrent safety** | Stateless file reads impose no locks | Multi-process batch jobs scale linearly |
| **Offline operation** | Cache persistence enables disconnected work | Design iterations continue without internet |

Large perimeters with complex building footprints can trigger 30-60 second Overpass timeouts. A cached read completes in under 100ms regardless of geometry complexity.

## Practical Code Examples

### Basic Cached Workflow

This pattern handles first-time download and subsequent cache hits transparently:

```python
from prettymaps import get_gdfs

layers = {
    "streets": {"tags": {"highway": True}},
    "building": {"tags": {"building": True}},
    "green": {"tags": {"landuse": ["grass", "forest"]}},
}

# First execution: downloads from OSM, writes to cache

gdfs = get_gdfs(
    query="Porto, Portugal",
    layers_dict=layers,
    radius=1500,
    logging=True,
)

# Second execution: reads from prettymaps_cache/*.geojson

gdfs = get_gdfs(
    query="Porto, Portugal",
    layers_dict=layers,
    radius=1500,
    logging=True,
)

```

### Custom Cache Directory

For shared team caches or managed storage locations:

```python
from prettymaps.fetch import write_to_cache, read_from_cache
from shapely.geometry import Polygon
import geopandas as gpd

perimeter = gpd.GeoDataFrame(
    geometry=[Polygon([
        (-8.62, 41.15), (-8.62, 41.17),
        (-8.60, 41.17), (-8.60, 41.15)
    ])],
    crs="EPSG:4326",
)

layer_gdf = gpd.GeoDataFrame(
    geometry=[Polygon([(-8.615, 41.155), (-8.61, 41.155), (-8.61, 41.16)])],
    crs="EPSG:4326"
)

# Write to project-specific cache

write_to_cache(
    perimeter,
    layer_gdf,
    {"tags": {"building": True}},
    cache_dir="/shared/prettymaps_cache"
)

# Retrieve later or from another machine with access to same path

cached = read_from_cache(
    perimeter,
    {"tags": {"building": True}},
    cache_dir="/shared/prettymaps_cache"
)

```

## Key Source Files

Understanding the cache implementation requires familiarity with these modules:

- **[`prettymaps/fetch.py`](https://github.com/marceloprates/prettymaps/blob/main/prettymaps/fetch.py)** — Core caching logic with `write_to_cache()`, `read_from_cache()`, and `unified_osm_request()` according to the marceloprates/prettymaps source code
- **[`prettymaps/draw.py`](https://github.com/marceloprates/prettymaps/blob/main/prettymaps/draw.py)** — Rendering engine that consumes cached or fresh GDFs
- **[`prettymaps/utils.py`](https://github.com/marceloprates/prettymaps/blob/main/prettymaps/utils.py)** — Supporting decorators including execution timing utilities

## Summary

- **Cache keys are deterministic hashes** combining perimeter geometry and layer parameters, implemented in [`prettymaps/fetch.py`](https://github.com/marceloprates/prettymaps/blob/main/prettymaps/fetch.py)
- **GeoJSON serialization** provides fast, interoperable storage without external dependencies
- **Manual or automatic activation** — call helpers directly or integrate with `get_gdfs()` pipeline
- **Three primary use cases**: iterative styling, batch map generation, and offline development

## Frequently Asked Questions

### How do I enable caching in the current prettymaps release?

The cache helper functions in [`prettymaps/fetch.py`](https://github.com/marceloprates/prettymaps/blob/main/prettymaps/fetch.py) are fully implemented but not automatically invoked in recent versions. Import `write_to_cache` and `read_from_cache` directly, wrap your fetch calls, or uncomment the cache integration in `unified_osm_request()` for transparent operation.

### What happens if OSM data updates while I have cached files?

Cached GeoDataFrames remain valid indefinitely. The cache has no TTL mechanism — you must manually delete `prettymaps_cache/` entries or specific hash-named files to force re-fetch. For rapidly changing features like construction zones, implement timestamp-based cache invalidation or periodic manual clears.

### Can multiple Python processes share the same cache directory safely?

Yes. The cache uses simple file existence checks and GeoPandas read operations without file locking. Concurrent reads of identical cache entries succeed without corruption. Simultaneous writes to the same hash (rare in practice due to deterministic keys) follow filesystem atomicity semantics.

### Does caching reduce memory usage or only network traffic?

Caching primarily eliminates network I/O. Loaded GeoDataFrames occupy identical memory whether sourced from cache or fresh download. For memory-constrained workflows, consider GeoParquet export or geometry simplification before caching.