# How the Plausible CSV Importer Parses and Validates Data: A Deep Dive into the Elixir Implementation

> Discover how Plausible Analytics parses and validates CSV data through a multi-stage Elixir pipeline. Learn about filename validation, schema checks, and efficient NimbleCSV streaming.

- Repository: [Plausible Analytics/analytics](https://github.com/plausible/analytics)
- Tags: deep-dive
- Published: 2026-05-19

---

**Plausible Analytics processes CSV imports through a multi-stage pipeline that validates filenames against strict naming conventions, verifies column structures against whitelisted schemas, and streams data via NimbleCSV to minimize memory usage.**

The `Plausible.Imported.CSVImporter` module orchestrates the entire CSV ingestion workflow in the plausible/analytics repository. When administrators initiate an import via the LiveView interface at [`lib/plausible_web/live/csv_import.ex`](https://github.com/plausible/analytics/blob/main/lib/plausible_web/live/csv_import.ex), the system executes a rigorous validation sequence before accepting any analytics data into the underlying PostgreSQL tables.

## Entry Point and Storage Dispatch

The import process begins with `parse_args/1`, which expects a map containing two critical keys: `"uploads"` (a list of `%Plug.Upload{}` structs) and `"storage"` (either `:local` or `:s3`). This function constructs a canonical **import description** containing the site ID, target table reference, and file-specific metadata.

The `import_data/2` function then acts as the central dispatcher, routing to `import_local/3` or `import_s3/3` based on the storage backend:

- **`import_local/3`** streams files directly from disk using `File.stream!`
- **`import_s3/3`** retrieves objects via `ExAws.S3.download_stream`

Both functions open a **readable stream** and pipe it into the CSV parser without loading entire files into memory, ensuring the system handles multi-gigabyte imports efficiently.

## Filename Validation and Date Extraction

Before parsing any rows, the importer validates the filename structure through three key functions defined in [`lib/plausible/imported/csv_importer.ex`](https://github.com/plausible/analytics/blob/main/lib/plausible/imported/csv_importer.ex):

- **`valid_filename?/1`** enforces the pattern `<site-id>_<table>.csv` (e.g., `123_visitors.csv` or `20240301_20240331_visitors.csv`)
- **`extract_table/1`** parses the table name segment to identify valid targets like `visitors`, `pages`, or `referrers`
- **`parse_filename!/1`** extracts **date ranges** embedded in filenames, using `parse_date!/1` to safely convert `YYYYMMDD` strings into `Date` structs

If any filename fails these checks, the importer aborts immediately with an informative error message before consuming compute resources.

## Column Structure Validation

Each supported analytics table maintains a strict **whitelisted schema** defined through metaprogramming in the importer module. The system relies on two core functions:

- **`input_structure!/1`** returns a map of required column names to their expected types (currently unified as strings, but structured for future type enforcement)
- **`input_columns!/1`** returns the exact column order the parser expects

The importer reads the CSV header row and verifies it matches the whitelisted structure exactly. Missing columns, extra columns, or misordered headers trigger immediate validation failures that rollback the transaction and surface errors to the UI.

## Streaming Parser and Row Processing

The actual parsing leverages **NimbleCSV** for high-performance, pure-Elixir RFC4180 compliance:

```elixir
stream
|> NimbleCSV.RFC4180.parse_stream(skip_headers: true)
|> Enum.each(&process_row/1)

```

For each row in the stream, the importer performs:

1. **Field count validation** – ensuring the row length matches the expected column count
2. **Sanity checks** – verifying non-empty timestamps and numeric values for counters
3. **Batch insertion** – calling `Plausible.Imported.CSVImporter.insert_row/2` (wrapping `Repo.insert_all/3`) to populate tables like `site_visitors` and `pageviews`

This streaming approach guarantees constant memory usage regardless of file size.

## Success Handling and Error Recovery

After processing all files, `on_success/2` executes the completion workflow:

- Creates a **site import** record tracking the ingestion metadata
- Updates the site's `imported_at` timestamp
- Triggers email notifications using the template at `lib/plausible_web/templates/email/csv_import.html.heex`

If any stage raises an exception—whether from malformed CSV syntax, validation failures, or database constraints—the importer rolls back the entire transaction and returns a structured error payload that [`lib/plausible_web/live/csv_import.ex`](https://github.com/plausible/analytics/blob/main/lib/plausible_web/live/csv_import.ex) surfaces to the user interface.

## Practical Implementation Examples

To trigger an import programmatically from a Phoenix controller:

```elixir
def create(conn, %{"uploads" => uploads, "storage" => storage}) do
  {:ok, args} = Plausible.Imported.CSVImporter.parse_args(%{
    "uploads" => uploads,
    "storage" => String.to_atom(storage)
  })

  case Plausible.Imported.CSVImporter.import_data(args, %{}) do
    {:ok, %{site_import: import}} ->
      conn
      |> put_flash(:info, "CSV import completed")
      |> redirect(to: Routes.site_path(conn, :show, import.site_id))

    {:error, reason} ->
      conn
      |> put_flash(:error, "Import failed: #{reason}")
      |> render("new.html")
  end
end

```

For LiveView integration, the UI component passes uploads to the backend:

```elixir
<.live_component
  module={PlausibleWeb.Live.CSVImport}
  id="csv-import"
  site={@site}
  uploads={@uploads}
  storage={@storage} />

```

The test suite at [`test/plausible/imported/csv_importer_test.exs`](https://github.com/plausible/analytics/blob/main/test/plausible/imported/csv_importer_test.exs) provides comprehensive coverage of parsing edge cases, validation failures, and successful ingestion paths.

## Summary

- **Strict filename enforcement** via `valid_filename?/1` and `parse_filename!/1` prevents malformed inputs before parsing begins
- **Schema validation** through `input_columns!/1` ensures CSV headers match expected table structures exactly
- **Memory-efficient streaming** using NimbleCSV.RFC4180 allows processing of large files without loading them entirely into RAM
- **Transaction safety** guarantees that validation failures or parsing errors roll back partial data imports
- **Multi-backend support** seamlessly handles both local filesystem and S3 storage sources

## Frequently Asked Questions

### What CSV format does Plausible expect for imports?

Plausible expects RFC4180-compliant CSV files with specific naming conventions: `<site-id>_<table>.csv` or `<start-date>_<end-date>_<table>.csv` where dates use `YYYYMMDD` format. The header row must match the exact column order defined in `input_columns!/1` for the specified table (visitors, pages, referrers, etc.), and all files must be uploaded through the LiveView interface at [`lib/plausible_web/live/csv_import.ex`](https://github.com/plausible/analytics/blob/main/lib/plausible_web/live/csv_import.ex).

### How does Plausible handle large CSV files during import?

The importer uses Elixir streams via `File.stream!` or `ExAws.S3.download_stream` combined with `NimbleCSV.RFC4180.parse_stream/2` to process files line-by-line. This streaming architecture maintains constant memory usage regardless of file size, allowing the system to import multi-gigabyte analytics exports without exhausting server resources.

### What happens if my CSV has extra or missing columns?

The importer validates the header row against the whitelisted schema returned by `input_columns!/1` before processing any data rows. If columns are missing, unexpected, or incorrectly ordered, the entire import transaction aborts immediately and returns a descriptive error to the UI. This strict validation prevents partial or corrupted data from entering the analytics tables.

### Can I import CSVs directly from cloud storage?

Yes. The `import_data/2` function dispatches to `import_s3/3` when the storage parameter is set to `:s3`, enabling direct ingestion from AWS S3 buckets via `ExAws.S3.download_stream`. The system treats S3 streams identically to local filesystem streams, applying the same validation, parsing, and error handling regardless of the storage backend.