# DuckDB Worker Process Architecture in DBX: Isolated File Preview System

> Explore the DuckDB worker process architecture in DBX, isolating file previews for a responsive UI and preventing crashes with heavy operations. Learn how DBX ensures stability.

- Repository: [skyler/dbx](https://github.com/t8y2/dbx)
- Tags: architecture
- Published: 2026-07-05

---

**DBX runs all DuckDB queries for file previews inside a separate, isolated worker process to keep the main UI responsive and prevent crashes from heavy CPU or memory operations.**

The **DuckDB worker process architecture** in DBX consists of three tightly-coupled components that enable safe, sandboxed execution of SQL queries against CSV, Parquet, and other file formats. This design isolates potentially expensive operations from the main application thread while maintaining low-latency communication through JSON-encoded messages over STDIN/STDOUT.

## Core Components of the Architecture

### The Worker Process

The worker is a minimal executable—actually the same DBX binary—started with the `--duckdb-worker` flag. When this flag is detected in [`src-tauri/src/main.rs`](https://github.com/t8y2/dbx/blob/main/src-tauri/src/main.rs), the binary bypasses the UI and enters a dedicated runtime that creates a DuckDB connection and listens for JSON-encoded requests on STDIN. This single-binary approach simplifies deployment while ensuring the worker code remains identical to the client's expectations.

### The Client Wrapper (`DuckDbWorkerClient`)

Running inside the main DBX process, the client manages the worker lifecycle. Implemented in [`crates/dbx-core/src/db/duckdb_worker_process.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/db/duckdb_worker_process.rs), it handles process spawning, request serialization, and response routing. A critical feature is the **process-limit semaphore** (`process_limiter`) that enforces a maximum number of concurrent workers (defaulting to `DUCKDB_WORKER_MAX_PROCESSES_DEFAULT`) to prevent resource exhaustion.

### The Communication Protocol

Both sides share message structures defined in [`crates/dbx-core/src/db/duckdb_worker_protocol.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/db/duckdb_worker_protocol.rs). The protocol uses a simple JSON-RPC style format with `method`, `params`, and `id` fields. Because the client and worker share the same Rust structs (`DuckDbWorkerRequest` and `DuckDbWorkerResponse`), protocol changes automatically propagate to both sides, eliminating version mismatches.

## Process Isolation and Lifecycle Management

When DBX needs to preview a file, it invokes `DuckDbWorkerClient::open` (or `open_with_process_limit` for explicit concurrency control). The client first checks the semaphore to ensure the system hasn't exceeded the configured worker limit.

If a worker is not already running, `ensure_started_locked` spawns a new child process using `tokio::process::Command` with the `--duckdb-worker` flag. The child's **stdin** and **stdout** are piped, and a background task (`spawn_stdout_reader`) continuously reads JSON responses from stdout. The client tracks pending requests in an `Arc<Mutex<HashMap<id, PendingRequest>>>` to match incoming responses with their original callers.

## Request-Response Flow and Protocol Design

**Client to Worker:**
The `send_request` method generates a unique request ID, constructs a `DuckDbWorkerRequest`, serializes it to a JSON line, and writes it to the child's STDIN. The request is stored in `inner.pending` along with a generation number to detect stale responses from restarted workers.

**Worker to Client:**
The worker reads each line from STDIN, parses the JSON, executes the corresponding DuckDB operation (such as `Execute`, `ListTables`, or `AttachDatabase`), and writes a `DuckDbWorkerResponse` JSON line to STDOUT. The `spawn_stdout_reader` task reads this line, looks up the pending request by ID, and forwards the result through a `oneshot` channel to the awaiting caller.

**Error Handling:**
If the worker process exits unexpectedly or fails to connect, `ensure_started_locked` automatically restarts it. All errors are wrapped in `DuckDbWorkerError` and propagated to the UI layer.

## Query Execution and Cancellation

The `execute` method wraps each request in a `tokio::select!` block that honors both `CancellationToken` signals and per-query timeouts. This allows DBX to handle user-initiated cancellations gracefully.

When cancellation is requested, `cancel_or_kill` first attempts a graceful shutdown by sending a `Cancel` request to the worker. If the worker fails to respond or terminate, the client forcibly kills the child process using the `kill` method on the process handle. This two-phase approach prevents zombie processes while ensuring the UI remains responsive even when queries hang.

## Configuration and Process Isolation Settings

DBX exposes a UI toggle (`duckdb_worker_process_isolation`) stored in [`crates/dbx-core/src/connection.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/connection.rs) within the `DuckDbWorkerHandle` type. When enabled, this flag forces every DuckDB operation—including file previews—to run in its own isolated worker process rather than sharing a connection. This setting persists in local storage via [`crates/dbx-core/src/storage.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/storage.rs), allowing users to prioritize stability over resource efficiency when working with untrusted or extremely large files.

## Implementation Example

The following patterns demonstrate how DBX utilizes the worker architecture for file previews:

```rust
// Open a DuckDB worker for a temporary file preview
let client = DuckDbWorkerClient::open(
    "/tmp/preview.duckdb".to_string(),
    vec![],                     // no attached databases
).await?;

// Execute a preview query with row limits
let result = client
    .execute(
        None,                                 // default database
        "SELECT * FROM my_table LIMIT 100".into(),
        Some(100),                            // max rows
        None,                                 // no cancellation token
        None,                                 // no per-query timeout
    )
    .await?;

```

For external data sources, attach additional files as virtual tables:

```rust
// Attach an external CSV as a virtual table
client.attach_database(AttachedDatabaseConfig {
    name: "external".to_string(),
    path: "/data/sales.csv".to_string(),
}).await?;

```

To handle user cancellations or timeouts:

```rust
// Cancel a long-running query gracefully
let cancel_token = CancellationToken::new();
let handle = tokio::spawn(async move {
    client.execute(
        None,
        "SELECT * FROM huge_table".into(),
        None,
        Some(cancel_token.clone()),
        None,
    ).await
});

// Trigger cancellation from UI interaction
cancel_token.cancel();

```

## Summary

- **DBX isolates DuckDB operations** in a separate worker process to protect the main UI from crashes and blocking.
- **Three components** form the architecture: the worker binary (`--duckdb-worker`), the client wrapper (`DuckDbWorkerClient`), and the shared JSON protocol.
- **STDIN/STDOUT piping** enables low-latency communication without network overhead, while a semaphore enforces process limits.
- **Automatic restarts and cancellation** ensure robust handling of hung queries or process failures.
- **Process isolation** can be toggled per connection, forcing every query into its own sandboxed worker for maximum safety.

## Frequently Asked Questions

### How does DBX prevent the DuckDB worker from crashing the main application?

DBX spawns the DuckDB engine as a separate child process using `tokio::process::Command` with the `--duckdb-worker` flag. Because the worker runs as an isolated OS process, memory exhaustion or segmentation faults in DuckDB terminate only the worker, not the main DBX UI. The client detects the exit via the `spawn_stdout_reader` task and can automatically restart the worker for subsequent queries.

### What communication protocol does the DuckDB worker use?

The worker uses a JSON-RPC style protocol defined in [`crates/dbx-core/src/db/duckdb_worker_protocol.rs`](https://github.com/t8y2/dbx/blob/main/crates/dbx-core/src/db/duckdb_worker_protocol.rs). Messages include a `method` field (such as `Execute` or `ListTables`), a `params` object, and a unique `id` for request-response correlation. Both client and worker share the same Rust structs, ensuring type safety and eliminating serialization mismatches.

### How does DBX handle concurrent file previews?

DBX enforces concurrency limits through a semaphore (`process_limiter`) in `DuckDbWorkerClient`. The default maximum is defined by `DUCKDB_WORKER_MAX_PROCESSES_DEFAULT`. When `open_with_process_limit` is called, the client blocks until a permit is available, preventing resource exhaustion from too many simultaneous DuckDB processes.

### Can users cancel a running DuckDB query in DBX?

Yes. The `execute` method accepts an optional `CancellationToken`. When cancelled, the client first sends a graceful `Cancel` request to the worker. If the worker fails to terminate, `cancel_or_kill` forcibly kills the child process. This mechanism ensures the UI remains responsive even when querying massive CSV or Parquet files.