How Python Timeline Point Extraction Normalizes Data in the Google Timeline Visualizer
The extract_timeline_points function in visualizer.py normalizes raw Google Location History JSON into a clean, chronologically-ordered list of datetime-coordinate dictionaries by resolving multiple export formats, merging semantic intervals, and eliminating duplicate path points.
Python timeline point extraction is the critical first stage of the google-timeline-visualizer pipeline. The extract_timeline_points function, implemented in [visualizer.py lines 24-62](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L24-L62), transforms highly variable Google Timeline exports into a canonical data structure that downstream components can trust. This article breaks down each normalization step with reference to the actual source implementation.
Resolving Multiple Export Formats
Google Timeline exports arrive in two possible structures. The extraction code detects and unifies these at L31-L36:
# The raw JSON may be either:
# - A top-level list of semanticSegments
# - A dict with a 'semanticSegments' key containing the list
if isinstance(raw_data, list):
segments = raw_data
else:
segments = raw_data.get('semanticSegments', [])
This defensive pattern ensures the pipeline handles both export variants without external preprocessing.
Timestamp Normalization Helpers
parse_timestamp
Located at L42-L51, this helper converts heterogeneous timestamp inputs into Python datetime objects:
def parse_timestamp(ts):
if ts is None:
return None
if isinstance(ts, datetime.datetime):
return ts
try:
return datetime.datetime.fromisoformat(ts)
except (ValueError, TypeError):
return None # Malformed timestamps are filtered downstream
path_timestamp
Timeline path points present a more complex case: they may carry absolute time fields or relative durationMinutesOffsetFromStartTime values. The path_timestamp helper (L52-L73) resolves these offsets against segment start times, clamps negative values, and validates against segment end boundaries.
Semantic vs. Raw Path Data Handling
The extraction distinguishes between semantic points (structured activities and visits) and raw timeline path points:
- Activity coordinates:
activity.startandactivity.endlocations - Visit coordinates:
visit.topCandidate.placeLocation - Timeline path points: GPS traces from
timelinePath
At L100-L104, the code decides storage strategy: if a segment has usable semantic coordinates, its path points are diverted to standalone_path_points; otherwise they populate canonical_points directly. This prevents premature duplication when both data sources exist.
Merging Intervals and Deduplicating
Building Disjoint Semantic Intervals
All segment start/end times with semantic records are collected and merged into minimal disjoint intervals at L122-L139. This creates a temporal map of "covered" periods where high-quality semantic data exists.
Filtering Covered Path Points
Standalone path points whose timestamps fall inside any merged semantic interval are discarded at L140-L154. This eliminates redundancy while preserving semantic locations, which encode place names and structured visit data that raw GPS lacks.
Final Deduplication and Sorting
Any remaining identical points—matched by (dt, lat, lon) tuple—are removed via dictionary keying at L155-L158. The final list is sorted by timestamp at L159-L160 before return.
Temporal Filtering with add_point
The add_point helper (L77-L82) enforces user-specified date constraints. Points are only appended when:
- A valid timestamp exists
- Valid coordinates (
lat,lon) are present - The timestamp falls within
year,start_date, orend_dateparameters if provided
Complete Usage Examples
from visualizer import extract_timeline_points
import json
import datetime
# Load raw Timeline export
with open("Timeline.json", "r", encoding="utf-8") as f:
raw = json.load(f)
# Extract all normalized points
all_points = extract_timeline_points(raw)
# Extract with date filtering for 2022
points_2022 = extract_timeline_points(
raw,
start_date=datetime.date(2022, 1, 1),
end_date=datetime.date(2022, 12, 31)
)
# Result structure: [{'dt': datetime, 'lat': float, 'lon': float}, ...]
print(f"Extracted {len(points_2022)} points")
first = points_2022[0]
print(f"First point: {first['dt']} at ({first['lat']:.4f}, {first['lon']:.4f})")
Source Files
| File | Purpose |
|---|---|
[visualizer.py](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py) |
Core extract_timeline_points implementation and all normalization helpers |
[tests/test_parser.py](https://github.com/mahlernim/google-timeline-visualizer/blob/main/tests/test_parser.py) |
Unit tests validating normalization against representative Timeline JSON |
[tests/test_release_workflow.py](https://github.com/mahlernim/google-timeline-visualizer/blob/main/tests/test_release_workflow.py) |
CI validation ensuring extraction runs without binary dependencies |
Summary
- Format normalization: Handles both list and dict-rooted Timeline exports transparently
- Timestamp resilience: Converts ISO strings, datetime objects, and relative offsets with graceful error handling
- Semantic priority: Merges activity/visit intervals and excludes redundant path points that fall within them
- Temporal filtering: Applies
year,start_date, andend_dateconstraints at extraction time - Deduplication: Removes exact coordinate-timestamp duplicates before returning sorted output
Python timeline point extraction in the Google Timeline Visualizer produces a canonical, time-ordered series of geospatial points ready for outlier filtering, distance calculations, and animation rendering.
Frequently Asked Questions
How does extract_timeline_points handle malformed timestamps?
The parse_timestamp helper catches ValueError and TypeError exceptions and returns None for any unparseable input. These entries are then filtered out by add_point, which requires a valid timestamp before inclusion.
Why are timeline path points stored separately from semantic points?
This separation, implemented at L100-L104, enables downstream deduplication. Path points from segments with semantic data are held in standalone_path_points pending interval comparison, then filtered against merged semantic coverage to avoid duplicate location records.
Can I extract points for a specific year without loading everything into memory?
Yes. Pass year=2022 (or start_date/end_date parameters) to extract_timeline_points. The temporal filter is applied during extraction at L77-L82, so only matching points are retained in the output list.
What coordinate precision does the normalized output preserve?
The extraction preserves the raw latitude and longitude floats from the Google export without truncation. The add_point helper validates presence but does not round or quantize values, leaving precision decisions to downstream processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →