How TimelineParser Handles Different Google Maps Timeline Export Formats
The TimelineParser in visualizer.py automatically detects and processes three Google Maps Timeline export formats: raw array exports from Android/iOS, semanticSegments-wrapped objects from Google Takeout, and legacy segment variants.
Google Maps Timeline data can be exported in several JSON structures depending on your device and export method. The google-timeline-visualizer repository implements a resilient parsing layer that normalizes these disparate formats into a unified point stream for visualization. This article examines how the extract_timeline_points function and its helpers handle format detection, coordinate normalization, and timestamp resolution.
Root-Level Format Detection
The parser begins by inspecting the top-level JSON structure in extract_timeline_points (visualizer.py, lines 31-36). This branching logic determines how to access the actual timeline segments:
- List format (Android/iOS direct export) — When
isinstance(data, list)returns True, the parser treats the entire array as the segment list. - Dictionary with
semanticSegments(Google Takeout) — When the top level is a dict, the parser extractsdata.get("semanticSegments", []). - Malformed or empty exports — Missing keys fall back to an empty list, which propagates a clear "no data" error downstream.
This detection requires no user configuration. The parser adapts automatically based on the file's structure.
Coordinate String Normalization
Google emits location data in three distinct string formats across platforms. The parse_coordinate helper (visualizer.py, lines 85-105) handles all variants:
| Format | Example | Handling |
|---|---|---|
| Android geo URI | "geo:37.422,-122.084?z=15" |
Strips "geo:" prefix and query parameters |
| iOS plain | "37.422,-122.084" |
Splits directly on comma |
| Scaled integers | "374220000,-1220840000" |
Detects values outside valid lat/lon range, divides by 10,000,000 |
The normalization pipeline:
- Removes degree symbols and whitespace
- Splits on commas into latitude/longitude components
- Converts to float and rescales when magnitude exceeds 180 (longitude) or 90 (latitude)
# Example: parsing coordinates from different export sources
from visualizer import parse_coordinate
android = parse_coordinate("geo:40.7128,-74.0060?z=12") # (40.7128, -74.0060)
ios = parse_coordinate("40.7128,-74.0060") # (40.7128, -74.0060)
scaled = parse_coordinate("407128000,-740060000") # (40.7128, -74.0060)
Flexible Timestamp Resolution
Timeline segments contain timestamps in multiple locations and formats. The parser extracts these through three pathways:
Semantic record timestamps
activity.startandactivity.endfor movement segmentsvisit.topCandidate.placeLocationfor stationary points
Raw path timestamps
timelinePathentries withdurationMinutesOffsetFromStartTime- Calculated as:
segment_start_time + timedelta(minutes=offset)
Timestamp parsing strategy
The parse_timestamp helper accepts:
- Native Python
datetimeobjects (passed through) - ISO-8601 strings
- RFC-2822 formats
- Other textual representations via
dateutil.parser.parse
from datetime import date
from visualizer import parse_timeline
# Filter exports by date range while parsing
timestamps, xs, ys, cum_dist, lats, lons = parse_timeline(
Path("timeline.json"),
start_date=date(2023, 6, 1),
end_date=date(2023, 6, 30),
filter_outliers="conservative"
)
Merging Semantic and Raw Path Data
A single Google Maps segment may contain both high-level semantic records and raw GPS traces. The parser handles this in extract_timeline_points (lines 40-55, 70-84) by:
- Extracting points from
activityandvisitrecords as semantic intervals - Extracting points from
timelinePatharrays with calculated timestamps - Merging both sets while respecting semantic intervals to prevent duplicate entries
- Filtering to user-specified date ranges via
_date_in_range - De-duplicating and sorting chronologically
The final output format is a standardized list of dictionaries:
{
"dt": datetime.datetime(2023, 6, 15, 14, 30, 0),
"lat": 40.7128,
"lon": -74.0060
}
Using the Parser Directly
For programmatic access without the full visualization pipeline:
import json
from pathlib import Path
from visualizer import extract_timeline_points
from datetime import date
# Load raw Google Takeout export
with open("takeout_2023.json") as f:
raw_data = json.load(f)
# Extract points with date filtering
points = extract_timeline_points(
raw_data,
start_date=date(2023, 1, 1),
end_date=date(2023, 12, 31)
)
print(f"Extracted {len(points)} timeline points")
# Each point: {"dt": datetime, "lat": float, "lon": float}
This low-level interface accepts any of the three supported export formats and returns normalized data ready for custom analysis or animation.
Summary
- Automatic format detection — The parser distinguishes list-based Android/iOS exports from
semanticSegmentsdictionaries without configuration. - Three coordinate formats — Geo URIs, plain strings, and scaled integers normalize through
parse_coordinate. - Multiple timestamp sources — Direct datetime objects, ISO strings, RFC-2822, and minute-offset calculations resolve through
parse_timestamp. - Unified output — Semantic records and raw path points merge into a chronological, de-duplicated stream of latitude/longitude/datetime dictionaries.
Frequently Asked Questions
How do I know which export format I have?
Open your JSON file in any text editor. If the first non-whitespace character is [, you have the list format from Android/iOS. If it starts with {, check for a "semanticSegments" key — this indicates Google Takeout format according to the visualizer.py source code.
Why are some coordinates appearing in the wrong location?
Check if your export contains scaled integer coordinates. The parse_coordinate function automatically rescales values exceeding normal latitude/longitude ranges, but malformed strings without commas or with degree symbols may require manual inspection. The parser handles degree symbols and removes them during normalization.
Can I parse Timeline data without installing dependencies?
The core parsing functions in visualizer.py require python-dateutil for timestamp resolution and numpy for coordinate processing. For minimal dependency usage, you could extract only the parse_coordinate and extract_timeline_points functions, though dateutil remains essential for flexible date parsing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →