# DBX File Validator Mechanism for CSV, Excel, and Parquet Uploads

> Explore the DBX file validator mechanism for CSV, Excel, and Parquet uploads. Learn how DBX ensures data integrity with its three-layer extension-based system, preventing malformed file processing.

- Repository: [skyler/dbx](https://github.com/t8y2/dbx)
- Tags: how-to-guide
- Published: 2026-07-04

---

**DBX validates uploaded data files through a three-layer extension-based system that inspects file suffixes via `import_file_kind`, parses content through `parse_import_file`, and generates safe preview queries using `build_dropped_file_preview_sql` to block malformed or unsupported files before processing.**

DBX implements a robust file validator mechanism to ensure data integrity when handling CSV, Excel (.xlsx), and Parquet uploads. Located in the core import module, this validator prevents malformed files from entering the processing pipeline by combining strict extension detection, content structure validation, and secure SQL generation for file previews.

## Extension Detection and File Type Validation

The validation process begins in [`crates/dbx-core/src/table_import.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/table_import.rs) with the `import_file_kind` function (lines 13-26). This utility inspects the lower-cased file name and returns an `ImportFileKind` enum variant—`Csv`, `Tsv`, `Json`, or `Xlsx`. 

If the suffix does not match one of these supported types, the function returns an error immediately, preventing the upload from proceeding to content parsing. This first layer acts as a gatekeeper that rejects unsupported file types before any resource-intensive processing occurs.

## Content Parsing and Structure Validation

Once the extension is verified, the `parse_import_file` function (lines 40-68 in [`table_import.rs`](https://github.com/t8y2/dbx/blob/main/table_import.rs)) orchestrates the second validation layer. This function branches to specialized parsers—including `parse_csv_bytes`, `parse_delimited_bytes`, `parse_json_bytes`, and `parse_xlsx_file`—which validate the file's internal structure.

These parsers check for critical formatting requirements such as CSV header presence and enforce row limits while simultaneously preparing a preview of rows and columns for the UI. If the file structure does not match the expected format—for example, if a CSV lacks headers or contains malformed delimiters—the parser returns an error before the data reaches the database.

## Safe Preview Generation for Drag-and-Drop Uploads

For drag-and-drop interactions including Parquet files, DBX implements an additional validation layer in [`crates/dbx-core/src/query_execution_sql.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/query_execution_sql.rs). The `build_dropped_file_preview_sql` function (lines 55-72) constructs safe `SELECT * FROM read_…` queries for supported formats.

This function explicitly checks the lower-cased file path and only creates preview queries for `.parquet`, `.csv`, `.tsv`, or `.json` files. If the suffix is not recognized, the function returns `None`, causing the UI to reject the file immediately. This approach leverages DuckDB's native read functions (such as `read_parquet`) to safely preview file contents without fully importing them.

## Implementation Examples

### Validating File Extensions

Use the `import_file_kind` utility to validate uploads before processing:

```rust
use dbx_core::table_import::import_file_kind;

fn validate_upload(path: &str) -> Result<(), String> {
    // Returns ImportFileKind::Csv, ::Xlsx, etc. or an error for unsupported types
    import_file_kind(path).map(|_| ())
}

```

### Parsing CSV Content for Preview

The `parse_import_file` dispatcher selects the correct parser based on the detected extension:

```rust
use dbx_core::table_import::{parse_import_file, TableImportRequest};

async fn preview_csv(req: TableImportRequest) -> Result<TableImportPreview, String> {
    // `parse_import_file` selects the correct parser based on the extension
    let preview = dbx_core::table_import::preview_table_import_file_core(&req.file_path).await?;
    Ok(preview)
}

```

### Generating Safe Parquet Previews

For drag-and-drop Parquet files, construct safe preview queries using the SQL builder:

```rust
use dbx_core::query_execution_sql::build_dropped_file_preview_sql;

fn parquet_preview(path: &str) -> Option<String> {
    // Returns a safe DuckDB query like:
    // SELECT * FROM read_parquet('/tmp/data.parquet') LIMIT 1000
    build_dropped_file_preview_sql(
        dbx_core::query_execution_sql::DroppedFilePreviewSqlOptions {
            path: path.to_string(),
            limit: Some(1000),
        },
    )
}

```

## Key Source Files

- **[`crates/dbx-core/src/table_import.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/table_import.rs)**: Defines `ImportFileKind`, `import_file_kind`, and content parsers (`parse_csv_bytes`, `parse_xlsx_file`, etc.) that validate file contents and structure.
- **[`crates/dbx-core/src/query_execution_sql.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/query_execution_sql.rs)**: Provides `build_dropped_file_preview_sql`, the gatekeeper for drag-and-drop previews of Parquet, CSV, TSV, and JSON files via DuckDB functions.
- **[`crates/dbx-web/src/routes/table_import.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-web/src/routes/table_import.rs)**: Exposes the HTTP endpoint that calls `preview_table_import_file_core` and utilizes the validator logic for web uploads.

## Summary

- **Extension validation** occurs first via `import_file_kind` in [`table_import.rs`](https://github.com/t8y2/dbx/blob/main/table_import.rs) (lines 13-26), rejecting unsupported file types before parsing begins.
- **Content structure validation** happens through `parse_import_file`, which dispatches to format-specific parsers (lines 40-68) that verify headers, delimiters, and row limits.
- **Drag-and-drop security** is enforced by `build_dropped_file_preview_sql` in [`query_execution_sql.rs`](https://github.com/t8y2/dbx/blob/main/query_execution_sql.rs) (lines 55-72), which generates safe DuckDB queries only for recognized extensions.
- The validator supports CSV, TSV, JSON, XLSX, and Parquet formats, ensuring only well-structured data enters the DBX processing pipeline.

## Frequently Asked Questions

### What file formats does DBX support for upload validation?

DBX validates CSV, TSV, JSON, Excel (.xlsx), and Parquet files. The `import_file_kind` function handles CSV, TSV, JSON, and XLSX extensions, while Parquet support is primarily implemented in the drag-and-drop preview system via `build_dropped_file_preview_sql`.

### How does DBX prevent unsupported files from being processed?

The validator checks file extensions immediately upon upload through `import_file_kind` in [`table_import.rs`](https://github.com/t8y2/dbx/blob/main/table_import.rs). If the lower-cased file suffix does not match `Csv`, `Tsv`, `Json`, or `Xlsx`, the function returns an error, halting the upload before content parsing occurs.

### Where is the Parquet file validation logic implemented?

Parquet validation for drag-and-drop uploads resides in [`crates/dbx-core/src/query_execution_sql.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/query_execution_sql.rs) within the `build_dropped_file_preview_sql` function (lines 55-72). This function validates the file extension and generates a safe `SELECT * FROM read_parquet()` query for preview purposes.

### Does DBX validate file content or just extensions?

DBX performs both extension and content validation. After the initial suffix check, `parse_import_file` validates actual file structure—including CSV header presence, delimiter correctness, and row limits—before generating UI previews or importing data.