# How Python Timeline Point Extraction Normalizes Data in the Google Timeline Visualizer

> Learn how Python timeline point extraction normalizes Google Location History JSON. This function cleans raw data, merges intervals, and removes duplicates for a chronological list.

- Repository: [mahlernim/google-timeline-visualizer](https://github.com/mahlernim/google-timeline-visualizer)
- Tags: how-to-guide
- Published: 2026-08-23

---

**The `extract_timeline_points` function in [`visualizer.py`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py) normalizes raw Google Location History JSON into a clean, chronologically-ordered list of datetime-coordinate dictionaries by resolving multiple export formats, merging semantic intervals, and eliminating duplicate path points.**

Python timeline point extraction is the critical first stage of the **google-timeline-visualizer** pipeline. The `extract_timeline_points` function, implemented in [[`visualizer.py`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py) lines 24-62](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L24-L62), transforms highly variable Google Timeline exports into a canonical data structure that downstream components can trust. This article breaks down each normalization step with reference to the actual source implementation.

## Resolving Multiple Export Formats

Google Timeline exports arrive in two possible structures. The extraction code detects and unifies these at [`L31-L36`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L31-L36):

```python

# The raw JSON may be either:

# - A top-level list of semanticSegments

# - A dict with a 'semanticSegments' key containing the list

if isinstance(raw_data, list):
    segments = raw_data
else:
    segments = raw_data.get('semanticSegments', [])

```

This defensive pattern ensures the pipeline handles both export variants without external preprocessing.

## Timestamp Normalization Helpers

### parse_timestamp

Located at [`L42-L51`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L42-L51), this helper converts heterogeneous timestamp inputs into Python `datetime` objects:

```python
def parse_timestamp(ts):
    if ts is None:
        return None
    if isinstance(ts, datetime.datetime):
        return ts
    try:
        return datetime.datetime.fromisoformat(ts)
    except (ValueError, TypeError):
        return None  # Malformed timestamps are filtered downstream

```

### path_timestamp

Timeline path points present a more complex case: they may carry absolute `time` fields or relative `durationMinutesOffsetFromStartTime` values. The `path_timestamp` helper ([`L52-L73`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L52-L73)) resolves these offsets against segment start times, clamps negative values, and validates against segment end boundaries.

## Semantic vs. Raw Path Data Handling

The extraction distinguishes between **semantic points** (structured activities and visits) and **raw timeline path points**:

- **Activity coordinates**: `activity.start` and `activity.end` locations
- **Visit coordinates**: `visit.topCandidate.placeLocation`
- **Timeline path points**: GPS traces from `timelinePath`

At [`L100-L104`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L100-L104), the code decides storage strategy: if a segment has usable semantic coordinates, its path points are diverted to `standalone_path_points`; otherwise they populate `canonical_points` directly. This prevents premature duplication when both data sources exist.

## Merging Intervals and Deduplicating

### Building Disjoint Semantic Intervals

All segment start/end times with semantic records are collected and merged into minimal disjoint intervals at [`L122-L139`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L122-L139). This creates a temporal map of "covered" periods where high-quality semantic data exists.

### Filtering Covered Path Points

Standalone path points whose timestamps fall inside any merged semantic interval are discarded at [`L140-L154`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L140-L154). This eliminates redundancy while preserving semantic locations, which encode place names and structured visit data that raw GPS lacks.

### Final Deduplication and Sorting

Any remaining identical points—matched by `(dt, lat, lon)` tuple—are removed via dictionary keying at [`L155-L158`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L155-L158). The final list is sorted by timestamp at [`L159-L160`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L159-L160) before return.

## Temporal Filtering with add_point

The `add_point` helper ([`L77-L82`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py#L77-L82)) enforces user-specified date constraints. Points are only appended when:

- A valid timestamp exists
- Valid coordinates (`lat`, `lon`) are present
- The timestamp falls within `year`, `start_date`, or `end_date` parameters if provided

## Complete Usage Examples

```python
from visualizer import extract_timeline_points
import json
import datetime

# Load raw Timeline export

with open("Timeline.json", "r", encoding="utf-8") as f:
    raw = json.load(f)

# Extract all normalized points

all_points = extract_timeline_points(raw)

# Extract with date filtering for 2022

points_2022 = extract_timeline_points(
    raw,
    start_date=datetime.date(2022, 1, 1),
    end_date=datetime.date(2022, 12, 31)
)

# Result structure: [{'dt': datetime, 'lat': float, 'lon': float}, ...]

 print(f"Extracted {len(points_2022)} points")
first = points_2022[0]
print(f"First point: {first['dt']} at ({first['lat']:.4f}, {first['lon']:.4f})")

```

## Source Files

| File | Purpose |
|------|---------|
| [[`visualizer.py`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py)](https://github.com/mahlernim/google-timeline-visualizer/blob/main/visualizer.py) | Core `extract_timeline_points` implementation and all normalization helpers |
| [[`tests/test_parser.py`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/tests/test_parser.py)](https://github.com/mahlernim/google-timeline-visualizer/blob/main/tests/test_parser.py) | Unit tests validating normalization against representative Timeline JSON |
| [[`tests/test_release_workflow.py`](https://github.com/mahlernim/google-timeline-visualizer/blob/main/tests/test_release_workflow.py)](https://github.com/mahlernim/google-timeline-visualizer/blob/main/tests/test_release_workflow.py) | CI validation ensuring extraction runs without binary dependencies |

## Summary

- **Format normalization**: Handles both list and dict-rooted Timeline exports transparently
- **Timestamp resilience**: Converts ISO strings, datetime objects, and relative offsets with graceful error handling
- **Semantic priority**: Merges activity/visit intervals and excludes redundant path points that fall within them
- **Temporal filtering**: Applies `year`, `start_date`, and `end_date` constraints at extraction time
- **Deduplication**: Removes exact coordinate-timestamp duplicates before returning sorted output

Python timeline point extraction in the Google Timeline Visualizer produces a **canonical, time-ordered series of geospatial points** ready for outlier filtering, distance calculations, and animation rendering.

## Frequently Asked Questions

### How does extract_timeline_points handle malformed timestamps?

The `parse_timestamp` helper catches `ValueError` and `TypeError` exceptions and returns `None` for any unparseable input. These entries are then filtered out by `add_point`, which requires a valid timestamp before inclusion.

### Why are timeline path points stored separately from semantic points?

This separation, implemented at `L100-L104`, enables downstream deduplication. Path points from segments with semantic data are held in `standalone_path_points` pending interval comparison, then filtered against merged semantic coverage to avoid duplicate location records.

### Can I extract points for a specific year without loading everything into memory?

Yes. Pass `year=2022` (or `start_date`/`end_date` parameters) to `extract_timeline_points`. The temporal filter is applied during extraction at `L77-L82`, so only matching points are retained in the output list.

### What coordinate precision does the normalized output preserve?

The extraction preserves the raw latitude and longitude floats from the Google export without truncation. The `add_point` helper validates presence but does not round or quantize values, leaving precision decisions to downstream processing.