How to Parse a Google Maps Timeline JSON Export in Python

Parse a Google Maps Timeline JSON export by loading the file with UTF-8 decoding, normalizing heterogeneous coordinate formats into decimal degrees, and extracting chronological points while merging duplicate semantic intervals.

The mahlernim/google-timeline-visualizer repository provides a robust parser that handles the complexity of Google's Timeline data structure. Whether you're analyzing location history for visualization or data science projects, understanding how to correctly interpret the nested JSON format is essential for accurate results.

Understanding the Google Maps Timeline JSON Structure

Google Maps Timeline exports arrive in two distinct shapes depending on when and how they were generated. Newer Android and iOS exports typically provide a raw array of segments, while older exports wrap data in a semanticSegments array nested within an object. Both formats contain timeline paths, visit data, and activity segments with varying timestamp and coordinate representations.

Loading and Validating the Export

The entry point for parsing is parse_timeline in visualizer.py. This function opens the JSON file, decodes it as UTF-8, and validates that the data is readable.

from visualizer import parse_timeline

# Load and parse the export for a specific year

timestamps, xs, ys, cum_dist, lats, lons = parse_timeline(
    "path/to/Timeline.json",
    2024,
)

If the file is malformed or cannot be decoded, the function raises a TimelineParseError at lines 98-105 rather than failing silently. This ensures downstream processing only receives valid data structures.

Normalizing Coordinate Formats

Google stores geographic coordinates inconsistently across export versions. The parse_coordinate function (lines 42-61 in visualizer.py) consolidates these into standard (lat, lon) tuples.

The parser handles:

  • latLng and point objects with explicit latitude/longitude keys
  • Plain string values requiring splitting and parsing
  • E7 integer format (multiplied by 10,000,000), which requires division to restore decimal degrees
  • geo: URI schemes embedding coordinates

Any coordinates outside valid geographic ranges (-90 to 90 latitude, -180 to 180 longitude) are discarded during normalization.

Extracting Timeline Points

The heavy lifting occurs in extract_timeline_points, which transforms raw JSON into a clean chronological dataset.

Handling Export Format Variations

At lines 64-71, the function detects the export structure by checking for the presence of semanticSegments. For each segment encountered, it parses start and end timestamps using parse_timestamp, then extracts activity or visit coordinates. If semantic coordinates exist, the segment's path points are flagged as canonical; otherwise, they remain as stand-alone path points for later deduplication.

Timestamp Resolution

Individual path points use two timestamp conventions: absolute time fields or relative offsets via durationMinutesOffsetFromStartTime. The path_timestamp helper (lines 87-107) resolves these offsets into absolute datetimes from the segment start time, guarding against negative values and overflow errors.

Deduplication and Merging Logic

After collecting points, the parser merges overlapping semantic intervals—time spans covered by activities or visits—and removes any stand-alone path points falling within these intervals (lines 66-90). This prevents duplicate entries where a visit coordinate and a GPS path point represent the same physical location. Finally, the function collapses duplicate (timestamp, lat, lon) triples and sorts the list chronologically (lines 92-96).

Working with Parsed Data

Once parsed, the data is returned as separate arrays suitable for analysis or visualization. The parse_timeline function also computes cumulative Haversine distance for each point at lines 122-138, enabling distance-based calculations and video timing.


# Inspect the first five points

for t, lat, lon in zip(timestamps[:5], lats[:5], lons[:5]):
    print(f"{t.isoformat()} → ({lat:.5f}, {lon:.5f})")

# Plot the route using matplotlib

import matplotlib.pyplot as plt

plt.figure(figsize=(8, 8))
plt.plot(xs, ys, marker="o", linewidth=2, markersize=3)
plt.title("Extracted Route for 2024")
plt.axis("equal")
plt.show()

Unit tests in tests/test_parser.py validate both legacy and modern export formats against expected JSON shapes, ensuring compatibility across Timeline versions.

Summary

  • parse_timeline in visualizer.py serves as the main entry point, handling file I/O and orchestrating the extraction pipeline.
  • parse_coordinate normalizes heterogeneous coordinate formats including E7 integers and geo: URIs into decimal degrees.
  • extract_timeline_points handles both semanticSegments and flat array export formats while merging duplicate intervals and resolving relative timestamps.
  • The parser returns timestamp arrays, Mercator meter projections, and cumulative distances suitable for immediate visualization or analysis.

Frequently Asked Questions

What coordinate formats does Google Maps Timeline JSON use?

Google Maps Timeline exports store coordinates in multiple formats including latLng objects, point objects, plain strings, E7-encoded integers (values multiplied by 10,000,000), and geo: URI strings. The parse_coordinate function in visualizer.py detects and converts all these variations into standard (latitude, longitude) decimal degree tuples.

How does the parser handle different Timeline export versions?

The parser detects the export structure at lines 64-71 by checking whether the JSON root is an object containing semanticSegments (older format) or a raw array of segments (newer Android/iOS exports). It processes both structures identically after normalization, extracting points from each segment while preserving chronological order.

What is the E7 coordinate format and why does it need scaling?

E7 is Google's integer representation for geographic coordinates where values are stored as integers multiplied by 10,000,000. For example, 37.774900° becomes 377749000. The parser applies a scaling factor of 10^-7 to convert these back to decimal degrees, discarding any values that fall outside valid geographic bounds after conversion.

How are duplicate points removed from the timeline?

The parser implements a two-stage deduplication process. First, it merges overlapping semantic intervals (activities and visits) and removes stand-alone path points that fall within these intervals. Second, it collapses exact duplicate (timestamp, latitude, longitude) triples and sorts the remaining points chronologically before returning the final dataset.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →