How the Plausible CSV Importer Parses and Validates Data: A Deep Dive into the Elixir Implementation

Plausible Analytics processes CSV imports through a multi-stage pipeline that validates filenames against strict naming conventions, verifies column structures against whitelisted schemas, and streams data via NimbleCSV to minimize memory usage.

The Plausible.Imported.CSVImporter module orchestrates the entire CSV ingestion workflow in the plausible/analytics repository. When administrators initiate an import via the LiveView interface at lib/plausible_web/live/csv_import.ex, the system executes a rigorous validation sequence before accepting any analytics data into the underlying PostgreSQL tables.

Entry Point and Storage Dispatch

The import process begins with parse_args/1, which expects a map containing two critical keys: "uploads" (a list of %Plug.Upload{} structs) and "storage" (either :local or :s3). This function constructs a canonical import description containing the site ID, target table reference, and file-specific metadata.

The import_data/2 function then acts as the central dispatcher, routing to import_local/3 or import_s3/3 based on the storage backend:

  • import_local/3 streams files directly from disk using File.stream!
  • import_s3/3 retrieves objects via ExAws.S3.download_stream

Both functions open a readable stream and pipe it into the CSV parser without loading entire files into memory, ensuring the system handles multi-gigabyte imports efficiently.

Filename Validation and Date Extraction

Before parsing any rows, the importer validates the filename structure through three key functions defined in lib/plausible/imported/csv_importer.ex:

  • valid_filename?/1 enforces the pattern <site-id>_<table>.csv (e.g., 123_visitors.csv or 20240301_20240331_visitors.csv)
  • extract_table/1 parses the table name segment to identify valid targets like visitors, pages, or referrers
  • parse_filename!/1 extracts date ranges embedded in filenames, using parse_date!/1 to safely convert YYYYMMDD strings into Date structs

If any filename fails these checks, the importer aborts immediately with an informative error message before consuming compute resources.

Column Structure Validation

Each supported analytics table maintains a strict whitelisted schema defined through metaprogramming in the importer module. The system relies on two core functions:

  • input_structure!/1 returns a map of required column names to their expected types (currently unified as strings, but structured for future type enforcement)
  • input_columns!/1 returns the exact column order the parser expects

The importer reads the CSV header row and verifies it matches the whitelisted structure exactly. Missing columns, extra columns, or misordered headers trigger immediate validation failures that rollback the transaction and surface errors to the UI.

Streaming Parser and Row Processing

The actual parsing leverages NimbleCSV for high-performance, pure-Elixir RFC4180 compliance:

stream
|> NimbleCSV.RFC4180.parse_stream(skip_headers: true)
|> Enum.each(&process_row/1)

For each row in the stream, the importer performs:

  1. Field count validation – ensuring the row length matches the expected column count
  2. Sanity checks – verifying non-empty timestamps and numeric values for counters
  3. Batch insertion – calling Plausible.Imported.CSVImporter.insert_row/2 (wrapping Repo.insert_all/3) to populate tables like site_visitors and pageviews

This streaming approach guarantees constant memory usage regardless of file size.

Success Handling and Error Recovery

After processing all files, on_success/2 executes the completion workflow:

  • Creates a site import record tracking the ingestion metadata
  • Updates the site's imported_at timestamp
  • Triggers email notifications using the template at lib/plausible_web/templates/email/csv_import.html.heex

If any stage raises an exception—whether from malformed CSV syntax, validation failures, or database constraints—the importer rolls back the entire transaction and returns a structured error payload that lib/plausible_web/live/csv_import.ex surfaces to the user interface.

Practical Implementation Examples

To trigger an import programmatically from a Phoenix controller:

def create(conn, %{"uploads" => uploads, "storage" => storage}) do
  {:ok, args} = Plausible.Imported.CSVImporter.parse_args(%{
    "uploads" => uploads,
    "storage" => String.to_atom(storage)
  })

  case Plausible.Imported.CSVImporter.import_data(args, %{}) do
    {:ok, %{site_import: import}} ->
      conn
      |> put_flash(:info, "CSV import completed")
      |> redirect(to: Routes.site_path(conn, :show, import.site_id))

    {:error, reason} ->
      conn
      |> put_flash(:error, "Import failed: #{reason}")
      |> render("new.html")
  end
end

For LiveView integration, the UI component passes uploads to the backend:

<.live_component
  module={PlausibleWeb.Live.CSVImport}
  id="csv-import"
  site={@site}
  uploads={@uploads}
  storage={@storage} />

The test suite at test/plausible/imported/csv_importer_test.exs provides comprehensive coverage of parsing edge cases, validation failures, and successful ingestion paths.

Summary

  • Strict filename enforcement via valid_filename?/1 and parse_filename!/1 prevents malformed inputs before parsing begins
  • Schema validation through input_columns!/1 ensures CSV headers match expected table structures exactly
  • Memory-efficient streaming using NimbleCSV.RFC4180 allows processing of large files without loading them entirely into RAM
  • Transaction safety guarantees that validation failures or parsing errors roll back partial data imports
  • Multi-backend support seamlessly handles both local filesystem and S3 storage sources

Frequently Asked Questions

What CSV format does Plausible expect for imports?

Plausible expects RFC4180-compliant CSV files with specific naming conventions: <site-id>_<table>.csv or <start-date>_<end-date>_<table>.csv where dates use YYYYMMDD format. The header row must match the exact column order defined in input_columns!/1 for the specified table (visitors, pages, referrers, etc.), and all files must be uploaded through the LiveView interface at lib/plausible_web/live/csv_import.ex.

How does Plausible handle large CSV files during import?

The importer uses Elixir streams via File.stream! or ExAws.S3.download_stream combined with NimbleCSV.RFC4180.parse_stream/2 to process files line-by-line. This streaming architecture maintains constant memory usage regardless of file size, allowing the system to import multi-gigabyte analytics exports without exhausting server resources.

What happens if my CSV has extra or missing columns?

The importer validates the header row against the whitelisted schema returned by input_columns!/1 before processing any data rows. If columns are missing, unexpected, or incorrectly ordered, the entire import transaction aborts immediately and returns a descriptive error to the UI. This strict validation prevents partial or corrupted data from entering the analytics tables.

Can I import CSVs directly from cloud storage?

Yes. The import_data/2 function dispatches to import_s3/3 when the storage parameter is set to :s3, enabling direct ingestion from AWS S3 buckets via ExAws.S3.download_stream. The system treats S3 streams identically to local filesystem streams, applying the same validation, parsing, and error handling regardless of the storage backend.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →