How Path Deduplication Logic Works in Google Timeline Visualizer
The path deduplication logic in Google Timeline Visualizer eliminates duplicate geographic coordinates by mapping each point to a composite key of timestamp, latitude, and longitude, ensuring only the latest occurrence of identical location-time pairs persists in the final timeline.
The Google Timeline Visualizer converts raw Google Takeout exports into clean, renderable geographic datasets. Because exported timelines frequently contain overlapping path segments alongside activity and visit records, the open-source repository mahlernim/google-timeline-visualizer implements a robust path deduplication logic in web/src/timeline.ts to produce a canonical list of unique points while strictly preserving temporal accuracy.
The Deduplication Pipeline in timeline.ts
The deduplication process runs inside the parseTimelineJson function and executes in four distinct stages to sanitize raw location data containing activities, visits, and standalone path points.
Pre-Filtering: Handling Semantic Coverage
Before creating the uniqueness map, the code filters standalone path points that fall within existing semantic coverage areas (activities or visits). The isCovered function checks if a path point overlaps with semantic intervals, while mergeIntervals consolidates overlapping time ranges to prevent double-counting spatial data already represented by visit or activity markers. This step ensures that path segments do not duplicate data already captured by higher-level semantic events.
Creating the Uniqueness Map (Lines 71-75)
At lines 71-75 of web/src/timeline.ts, the algorithm instantiates a Map object to track unique points using a deterministic key strategy:
const unique = new Map<string, GeoPoint>();
for (const point of points) {
const key = `${point.instant.getTime()}:${point.latitude}:${point.longitude}`;
unique.set(key, point);
}
Each GeoPoint is keyed by a composite string combining millisecond timestamp, latitude, and longitude. Because Map.set() overwrites existing entries with identical keys, only the last processed occurrence of any duplicate coordinate-time combination survives. This approach achieves O(1) lookup complexity for duplicate detection, which is essential when processing timelines containing tens of thousands of points.
Chronological Normalization (Lines 76-79)
After extracting deduplicated values into an array, the code applies conditional sorting logic at lines 76-79:
const deduplicated = [...unique.values()];
const normalized = deduplicated.some(p => p.timeZoneMissing)
? deduplicated
: deduplicated.sort((a, b) => a.instant.getTime() - b.instant.getTime());
If any point lacks timezone information (timeZoneMissing === true), the original insertion order is preserved to prevent temporal shifts caused by sorting incomplete timestamps. Otherwise, the array is sorted strictly by epoch time to guarantee a monotonically increasing timeline for visualization components.
Practical Implementation Examples
The following patterns demonstrate how to leverage the deduplication logic in production applications.
Parsing Raw Timeline Data
import { parseTimelineJson } from './timeline';
// rawData represents the JSON payload exported from Google Takeout
const points = parseTimelineJson(rawData);
console.log(`Loaded ${points.length} unique location points`);
The parseTimelineJson function automatically executes the deduplication pipeline, returning a clean array free of overlapping path artifacts.
Integrating with React Components
import React from 'react';
import { parseTimelineJson } from './timeline';
import { MapViewer } from './MapViewer';
function TimelineMap({ rawJson }: { rawJson: unknown }) {
const points = React.useMemo(() => parseTimelineJson(rawJson), [rawJson]);
return <MapViewer points={points} />;
}
By passing the deduplicated array to MapViewer, the component avoids visual artifacts such as stacked markers or redundant path drawing that would occur if duplicate coordinates were rendered.
Design Rationale and Performance
The path deduplication logic prioritizes determinism and processing efficiency. The composite key strategy ensures that identical coordinates recorded at identical timestamps—common when path data overlaps with activity records in Google Takeout exports—are collapsed into single entries regardless of their source type.
The time-zone awareness built into the sorting logic respects the integrity of incomplete export data. When timezone metadata is absent, preserving the original order prevents inadvertent chronological displacement that could misrepresent the actual sequence of movements.
Summary
- The core deduplication logic resides in
web/src/timeline.tswithin theparseTimelineJsonfunction. - A JavaScript
Maputilizes composite keys combining millisecond timestamp, latitude, and longitude to detect duplicates in constant time. - Path points are pre-filtered against semantic coverage using
isCoveredandmergeIntervalsto eliminate redundancy with visit and activity data. - The algorithm preserves insertion order when timezones are missing, otherwise enforces strict epoch-time sorting.
- This implementation handles large-scale timeline datasets efficiently while maintaining geographical and temporal accuracy.
Frequently Asked Questions
What defines a duplicate point in this system?
A duplicate is any geographic point sharing the exact same millisecond timestamp, latitude, and longitude as another entry in the dataset. The path deduplication logic treats these as identical regardless of whether they originate from standalone path segments, activities, or visit records, collapsing them into a single canonical entry.
Why does the implementation retain the last occurrence of duplicates?
The code uses Map.set() which naturally overwrites previous values when identical keys are inserted. Since the input array is processed sequentially, the last occurrence of a duplicate coordinate-time pair persists. This behavior ensures that if subsequent entries contain more complete metadata or corrected coordinates, that enriched data survives the deduplication process.
How does the visualizer handle datasets with missing timezone information?
When any point in the collection has timeZoneMissing set to true, the algorithm bypasses chronological sorting and preserves the original array order derived from the Map iteration. This defensive pattern prevents temporal misalignment that could occur if the code attempted to sort timestamps lacking timezone context.
Where can I find the specific implementation of the deduplication algorithm?
The core deduplication logic is implemented at lines 71-79 of web/src/timeline.ts in the mahlernim/google-timeline-visualizer repository, with preliminary filtering functions (isCovered and mergeIntervals) located earlier in the same file to handle path-to-semantic coverage relationships.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →