How the Plausible CSV Importer Parses and Validates Data: A Deep Dive into the Elixir Implementation
Plausible Analytics processes CSV imports through a multi-stage pipeline that validates filenames against strict naming conventions, verifies column structures against whitelisted schemas, and streams data via NimbleCSV to minimize memory usage.
The Plausible.Imported.CSVImporter module orchestrates the entire CSV ingestion workflow in the plausible/analytics repository. When administrators initiate an import via the LiveView interface at lib/plausible_web/live/csv_import.ex, the system executes a rigorous validation sequence before accepting any analytics data into the underlying PostgreSQL tables.
Entry Point and Storage Dispatch
The import process begins with parse_args/1, which expects a map containing two critical keys: "uploads" (a list of %Plug.Upload{} structs) and "storage" (either :local or :s3). This function constructs a canonical import description containing the site ID, target table reference, and file-specific metadata.
The import_data/2 function then acts as the central dispatcher, routing to import_local/3 or import_s3/3 based on the storage backend:
import_local/3streams files directly from disk usingFile.stream!import_s3/3retrieves objects viaExAws.S3.download_stream
Both functions open a readable stream and pipe it into the CSV parser without loading entire files into memory, ensuring the system handles multi-gigabyte imports efficiently.
Filename Validation and Date Extraction
Before parsing any rows, the importer validates the filename structure through three key functions defined in lib/plausible/imported/csv_importer.ex:
valid_filename?/1enforces the pattern<site-id>_<table>.csv(e.g.,123_visitors.csvor20240301_20240331_visitors.csv)extract_table/1parses the table name segment to identify valid targets likevisitors,pages, orreferrersparse_filename!/1extracts date ranges embedded in filenames, usingparse_date!/1to safely convertYYYYMMDDstrings intoDatestructs
If any filename fails these checks, the importer aborts immediately with an informative error message before consuming compute resources.
Column Structure Validation
Each supported analytics table maintains a strict whitelisted schema defined through metaprogramming in the importer module. The system relies on two core functions:
input_structure!/1returns a map of required column names to their expected types (currently unified as strings, but structured for future type enforcement)input_columns!/1returns the exact column order the parser expects
The importer reads the CSV header row and verifies it matches the whitelisted structure exactly. Missing columns, extra columns, or misordered headers trigger immediate validation failures that rollback the transaction and surface errors to the UI.
Streaming Parser and Row Processing
The actual parsing leverages NimbleCSV for high-performance, pure-Elixir RFC4180 compliance:
stream
|> NimbleCSV.RFC4180.parse_stream(skip_headers: true)
|> Enum.each(&process_row/1)
For each row in the stream, the importer performs:
- Field count validation – ensuring the row length matches the expected column count
- Sanity checks – verifying non-empty timestamps and numeric values for counters
- Batch insertion – calling
Plausible.Imported.CSVImporter.insert_row/2(wrappingRepo.insert_all/3) to populate tables likesite_visitorsandpageviews
This streaming approach guarantees constant memory usage regardless of file size.
Success Handling and Error Recovery
After processing all files, on_success/2 executes the completion workflow:
- Creates a site import record tracking the ingestion metadata
- Updates the site's
imported_attimestamp - Triggers email notifications using the template at
lib/plausible_web/templates/email/csv_import.html.heex
If any stage raises an exception—whether from malformed CSV syntax, validation failures, or database constraints—the importer rolls back the entire transaction and returns a structured error payload that lib/plausible_web/live/csv_import.ex surfaces to the user interface.
Practical Implementation Examples
To trigger an import programmatically from a Phoenix controller:
def create(conn, %{"uploads" => uploads, "storage" => storage}) do
{:ok, args} = Plausible.Imported.CSVImporter.parse_args(%{
"uploads" => uploads,
"storage" => String.to_atom(storage)
})
case Plausible.Imported.CSVImporter.import_data(args, %{}) do
{:ok, %{site_import: import}} ->
conn
|> put_flash(:info, "CSV import completed")
|> redirect(to: Routes.site_path(conn, :show, import.site_id))
{:error, reason} ->
conn
|> put_flash(:error, "Import failed: #{reason}")
|> render("new.html")
end
end
For LiveView integration, the UI component passes uploads to the backend:
<.live_component
module={PlausibleWeb.Live.CSVImport}
id="csv-import"
site={@site}
uploads={@uploads}
storage={@storage} />
The test suite at test/plausible/imported/csv_importer_test.exs provides comprehensive coverage of parsing edge cases, validation failures, and successful ingestion paths.
Summary
- Strict filename enforcement via
valid_filename?/1andparse_filename!/1prevents malformed inputs before parsing begins - Schema validation through
input_columns!/1ensures CSV headers match expected table structures exactly - Memory-efficient streaming using NimbleCSV.RFC4180 allows processing of large files without loading them entirely into RAM
- Transaction safety guarantees that validation failures or parsing errors roll back partial data imports
- Multi-backend support seamlessly handles both local filesystem and S3 storage sources
Frequently Asked Questions
What CSV format does Plausible expect for imports?
Plausible expects RFC4180-compliant CSV files with specific naming conventions: <site-id>_<table>.csv or <start-date>_<end-date>_<table>.csv where dates use YYYYMMDD format. The header row must match the exact column order defined in input_columns!/1 for the specified table (visitors, pages, referrers, etc.), and all files must be uploaded through the LiveView interface at lib/plausible_web/live/csv_import.ex.
How does Plausible handle large CSV files during import?
The importer uses Elixir streams via File.stream! or ExAws.S3.download_stream combined with NimbleCSV.RFC4180.parse_stream/2 to process files line-by-line. This streaming architecture maintains constant memory usage regardless of file size, allowing the system to import multi-gigabyte analytics exports without exhausting server resources.
What happens if my CSV has extra or missing columns?
The importer validates the header row against the whitelisted schema returned by input_columns!/1 before processing any data rows. If columns are missing, unexpected, or incorrectly ordered, the entire import transaction aborts immediately and returns a descriptive error to the UI. This strict validation prevents partial or corrupted data from entering the analytics tables.
Can I import CSVs directly from cloud storage?
Yes. The import_data/2 function dispatches to import_s3/3 when the storage parameter is set to :s3, enabling direct ingestion from AWS S3 buckets via ExAws.S3.download_stream. The system treats S3 streams identically to local filesystem streams, applying the same validation, parsing, and error handling regardless of the storage backend.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →