How Plausible Imports Google Analytics 4 Historical Data
Plausible migrates Google Analytics 4 historical data through a four-stage Elixir pipeline that authenticates via OAuth, schedules Oban background jobs, paginates through the GA4 Reporting API, and buffers results into ClickHouse with automatic resume capabilities.
Plausible's open-source analytics platform provides a complete migration path from Google Analytics 4 through a robust import system implemented in the plausible/analytics repository. The architecture separates concerns between the Phoenix controller layer, background job processing, API client resilience, and high-throughput database ingestion. This technical deep dive examines the exact function calls, file paths, and data transformations that power the GA4 historical data import process.
OAuth Authentication and Property Selection
The import process begins in PlausibleWeb.GoogleAnalyticsController, where users authenticate with Google OAuth and select their GA4 property and date range.
Token Handling and Property Listing
When a user completes OAuth authentication, the frontend posts access_token, refresh_token, and expires_at to the property picker endpoint. The controller's property_form/2 function validates these credentials and retrieves available properties:
# lib/plausible_web/controllers/google_analytics_controller.ex
result = Google.API.list_properties(access_token)
After property selection, the property/2 action validates the requested date range against GA4's actual data availability using Google.API.get_analytics_start_date/2 and Google.API.get_analytics_end_date/2.
Date Validation and Import Initiation
Once validated, the import/2 action creates the import job through Imported.GoogleAnalytics4.new_import/3:
{:ok, _} <- Imported.GoogleAnalytics4.new_import(site, current_user, import_opts)
This immediately returns to the user while the actual data migration proceeds asynchronously.
Job Scheduling with Oban
Plausible.Imported.GoogleAnalytics4.new_import/3 wraps the generic Plausible.Imported.Importer behavior, storing the property ID, date range, and OAuth tokens as an Oban job payload:
def new_import(site, user, opts), do: Plausible.Imported.Importer.new_import(name(), site, user, opts, &before_start/2)
When Oban processes the job, it delegates to import_data/2, which coordinates the actual data retrieval and persistence.
Data Retrieval and Transformation
The core import logic resides in Plausible.Google.GA4.API.import_analytics/4, which orchestrates API calls, handles pagination, and manages rate limiting across the reporting datasets.
Orchestrating the GA4 Reporting API
The import process first refreshes the OAuth access token via Google.API.maybe_refresh_token/1, then determines whether this is a fresh import or a resumed job:
{:ok, access_token} <- Google.API.maybe_refresh_token(auth)
For fresh imports, the system iterates over all reports defined in GA4.ReportRequest.full_report/0. For resumed imports (after rate limiting), it starts from the saved resume_from_dataset and resume_from_offset parameters.
Pagination and Rate Limit Resilience
The fetch_and_persist/2 function handles pagination by comparing the current offset against the total row count:
if report_request.offset + @per_page < row_count do
fetch_and_persist(%GA4.ReportRequest{report_request | offset: report_request.offset + @per_page}, opts)
When the GA4 API returns {:error, {:rate_limit_exceeded, details}}, the error propagates to import_data/2, which creates a resumable job rather than failing permanently.
Schema Translation from GA4 to ClickHouse
Plausible.Imported.GoogleAnalytics4.from_report/4 transforms GA4 API rows into Plausible's ClickHouse schema, filtering out GA4's "missing" dimension values:
if Map.get(row.dimensions, "date") in @missing_values do
acc
else
[new_from_report(site_id, import_id, table, row) | acc]
end
The new_from_report/4 function handles specific table mappings. For visitor metrics, it translates GA4 fields like totalUsers and screenPageViews into Plausible's imported_visitors structure:
%{
date: get_date(row),
visitors: row.metrics["totalUsers"] |> parse_number(),
pageviews: row.metrics["screenPageViews"] |> parse_number(),
bounces: row.metrics["bounces"] |> parse_number(),
visits: row.metrics["sessions"] |> parse_number(),
visit_duration: row.metrics["userEngagementDuration"] |> parse_number()
}
Bulk Persistence and Resume Logic
High-throughput ingestion relies on Plausible.Imported.Buffer, which accumulates records in memory and flushes batches to ClickHouse every flush_interval_ms (default 1000ms).
In-Memory Buffering
The import_data/2 function starts a buffer process and creates a persistence closure:
{:ok, buffer} = Plausible.Imported.Buffer.start_link(flush_interval_ms: flush_interval_ms)
persist_fn = fn table, rows ->
records = from_report(rows, site_import.site_id, site_import.id, table)
Plausible.Imported.Buffer.insert_many(buffer, table, records)
end
This buffering strategy reduces database round-trips while preventing memory bloat during large historical imports.
Resumable Imports After Failures
When recoverable errors (rate limits, socket failures, or server errors) occur, import_data/2 schedules a resumed job with the current dataset and offset:
new_import(site_import.site, site_import.imported_by,
resume_from_import_id: site_import.id,
resume_from_dataset: dataset,
resume_from_offset: offset,
job_opts: [schedule_in: {65, :minutes}, unique: nil])
The skip_mark_failed?: true option ensures the import remains in a resumable state rather than being marked as permanently failed.
Code Examples
Triggering Imports via CLI
For testing or administrative tasks, you can manually trigger imports from an IEx console:
auth = {"ya29.A0AfH6SM...", "1//0g...", "2026-07-01T12:00:00Z"}
import_opts = [
property: "123456789",
date_range: Date.range(~D[2022-01-01], ~D[2022-12-31]),
auth: auth,
flush_interval_ms: 2_000
]
{:ok, site} = Plausible.Sites.get_by_domain("example.com")
{:ok, user} = Plausible.Auth.get_user_by_email("admin@example.com")
{:ok, _job} = Plausible.Imported.GoogleAnalytics4.new_import(site, user, import_opts)
Direct API Usage for Debugging
To inspect raw GA4 data without persisting to the database:
date_range = Date.range(~D[2022-01-01], ~D[2022-01-07])
property = "123456789"
auth = {"access", "refresh", "2026-07-01T12:00:00Z"}
persist_fn = fn table, rows ->
IO.inspect({:persist, table, length(rows)})
end
Plausible.Google.GA4.API.import_analytics(
date_range,
property,
auth,
persist_fn: persist_fn,
fetch_opts: [max_attempts: 3]
)
Resuming Failed Imports
When a rate limit interrupts an import, resume from the specific dataset and offset:
resume_opts = [
dataset: "users",
offset: 200_000
]
Plausible.Google.GA4.API.import_analytics(
date_range,
property,
auth,
persist_fn: persist_fn,
resume_opts: resume_opts
)
Summary
- Four-stage pipeline: OAuth authentication in
GoogleAnalyticsController, Oban job scheduling inImported.Importer, API orchestration inGA4.API, and bulk insertion viaImported.Buffer. - Resilient pagination: The
fetch_and_persist/2function recursively pages through GA4 reports with configurable batch sizes, automatically resuming from the last successful offset after rate limits. - Schema mapping:
from_report/4translates GA4 metrics liketotalUsersandscreenPageViewsinto Plausible's ClickHouse tables (imported_visitors,imported_sources, etc.), filtering invalid dimension values. - Atomic batching: The
Bufferprocess collects records in memory and flushes to ClickHouse every 1000ms, balancing throughput with memory constraints. - Resume capability: Failed imports automatically schedule resumption jobs with preserved state (
resume_from_dataset,resume_from_offset), ensuring no data loss during temporary API outages.
Frequently Asked Questions
What OAuth permissions are required for GA4 import?
Plausible requires Google Analytics read permissions to list properties and access the Reporting API. The application requests these scopes during the OAuth flow handled by PlausibleWeb.GoogleAnalyticsController, specifically using the tokens to call Google.API.list_properties/1 and the Data API v1 for historical reports.
How does Plausible handle GA4 API rate limits?
The system implements exponential backoff through Oban job rescheduling. When import_analytics/4 encounters a rate_limit_exceeded error, import_data/2 catches the failure and creates a new job scheduled 65 minutes in the future with the current dataset and offset preserved, allowing seamless continuation without duplicate data.
Can imports be resumed if they fail mid-process?
Yes. The architecture supports resumable imports through the resume_from_dataset and resume_from_offset options passed to GA4.API.import_analytics/4. When a job fails due to network issues or rate limiting, Plausible schedules a follow-up job that automatically restarts from the exact report page where the interruption occurred.
Which historical data tables does Plausible migrate from GA4?
The importer extracts data into multiple ClickHouse tables including imported_visitors (session metrics), imported_sources (referral data), and imported_pages (pageview details). The GA4.ReportRequest module defines the specific dimensions and metrics requested for each table, while from_report/4 handles the row-level transformation logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →