How to Register Parquet, JSON, and CSV Data in WrenAI’s Browser-Based WASM Engine
You must register external datasets using engine.registerParquet(), engine.registerJson(), or engine.registerCsv() before loading the MDL model when running WrenAI in local source mode.
WrenAI provides a WebAssembly (WASM) engine that executes entirely within the browser, enabling client-side analytical queries without a backend server. To make external data files available for querying, you need to register them with the engine instance prior to loading your semantic model. This article explains the three registration methods implemented in the wren-core-wasm SDK and demonstrates how to use them with Parquet, JSON, and CSV formats.
Understanding the WASM Engine Architecture
Engine Initialization
The WASM engine lifecycle begins with WrenEngine.init(), which loads the Rust-compiled WASM binary from either a remote URL or an in-memory ArrayBuffer. According to the source code in core/wren-core-wasm/sdk/src/index.ts, this factory method returns a WrenEngine instance that wraps the underlying WasmEngine and exposes high-level TypeScript methods for data registration.
Source Modes: URL vs. Local
The engine operates in two distinct modes controlled by the source parameter passed to loadMDL() (see lines 188-193 in src/index.ts):
- URL mode: When
sourcepoints to a remote endpoint (e.g.,https://cdn.example.com/data/), the engine automatically resolves Parquet files as{source}/{table_name}.parquetwithout requiring explicit registration. - Local mode: When
sourcereferences a local directory (e.g.,./data/), you must pre-register every table using the dedicated registration methods before callingloadMDL().
Registration Methods Overview
The SDK exposes three specific methods for ingesting data into the WASM heap, each defined in core/wren-core-wasm/sdk/src/index.ts:
| Method | Signature | Location |
|---|---|---|
| Parquet | registerParquet(name: string, data: BufferSource) |
Lines 200-212 |
| JSON | registerJson(name: string, data: unknown[]) |
Lines 215-220 |
| CSV | registerCsv(name: string, data: string | BufferSource, options?: CsvReadOptions) |
Lines 247-262 |
Each method converts the input into a Uint8Array (or Arrow representation) and passes it to the Rust engine, where it becomes queryable via SQL.
Step-by-Step Registration Guide
Registering Parquet Files
Parquet registration accepts any BufferSource (typically an ArrayBuffer from a fetch response). The engine parses the binary and creates an Arrow table representation internally.
import { WrenEngine } from "wren-core-wasm/sdk";
async function setupParquet() {
const engine = await WrenEngine.init();
// Fetch Parquet bytes from remote or local source
const response = await fetch("https://example.com/data/orders.parquet");
const parquetBytes = await response.arrayBuffer();
// Register with a table name matching your MDL schema
await engine.registerParquet("orders", parquetBytes);
// Load MDL referencing the registered table
await engine.loadMDL(mdlManifest, { source: "./data/" });
}
Registering JSON Data
JSON registration expects a JavaScript array of objects. The engine infers the Arrow schema from the first record and converts the entire array into a columnar format.
async function setupJson() {
const engine = await WrenEngine.init();
const customers = [
{ id: 1, name: "Alice", tier: "gold", signup_date: "2023-01-15" },
{ id: 2, name: "Bob", tier: "silver", signup_date: "2023-06-22" }
];
await engine.registerJson("customers", customers);
await engine.loadMDL(mdlManifest, { source: "./data/" });
}
Registering CSV Data with Options
CSV registration offers the most flexibility through the optional CsvReadOptions parameter. As implemented in lines 247-262, you can specify header detection, delimiters, and explicit schema definitions. When omitted, the engine infers the schema from the first 1,000 rows using the CsvColumnType mappings defined in lines 82-116.
async function setupCsv() {
const engine = await WrenEngine.init();
const csvString = "id;amount;category\n1;100.5;electronics\n2;200.75;furniture";
const options = {
header: true,
delimiter: ";",
schema: [
{ name: "id", type: "int64" },
{ name: "amount", type: "float64" },
{ name: "category", type: "string" }
]
};
await engine.registerCsv("metrics", csvString, options);
await engine.loadMDL(mdlManifest, { source: "./data/" });
}
Complete Working Example
Here is a complete pattern for initializing the engine, registering multiple data formats, and executing a query:
import { WrenEngine } from "wren-core-wasm/sdk";
async function initializeAnalytics() {
// 1. Initialize WASM engine
const engine = await WrenEngine.init();
try {
// 2. Register heterogeneous data sources
const parquetResp = await fetch("/data/sales.parquet");
await engine.registerParquet("sales", await parquetResp.arrayBuffer());
await engine.registerJson("targets", [
{ region: "North", q1_target: 100000 },
{ region: "South", q1_target: 85000 }
]);
await engine.registerCsv("returns", "order_id,reason\n1001,defective\n1002,wrong_item", {
header: true
});
// 3. Load semantic model
await engine.loadMDL(mdlManifest, { source: "./data/" });
// 4. Query across registered tables
const results = await engine.query(`
SELECT s.region, SUM(s.amount) as total, t.q1_target
FROM sales s
JOIN targets t ON s.region = t.region
LEFT JOIN returns r ON s.order_id = r.order_id
WHERE r.order_id IS NULL
GROUP BY s.region, t.q1_target
`);
return results;
} finally {
// 5. Release WASM heap memory
engine.free();
}
}
Memory Management Best Practices
Because the WASM engine allocates memory within the browser's sandbox, you should explicitly release resources when analytics operations complete. Call engine.free() to deallocate the underlying WasmEngine instance and purge registered Arrow tables from the heap. This is particularly important when registering large Parquet files or frequently reinitializing the engine in single-page applications.
Summary
- Register before loading: Always call
registerParquet(),registerJson(), orregisterCsv()beforeloadMDL()when using local source mode. - File locations: The registration APIs are defined in
core/wren-core-wasm/sdk/src/index.tsat lines 200-212, 215-220, and 247-262 respectively. - Data conversion: The engine automatically converts Parquet buffers, JSON arrays, and CSV strings into Arrow tables for query execution.
- CSV flexibility: Use
CsvReadOptionsto handle custom delimiters, header rows, and explicit schema definitions when automatic inference is insufficient. - Resource cleanup: Invoke
engine.free()to prevent memory leaks in long-running browser sessions.
Frequently Asked Questions
What data types can I pass to registerParquet()?
The registerParquet() method accepts any BufferSource, which includes ArrayBuffer, Uint8Array, or DataView objects. Typically, you obtain this by calling response.arrayBuffer() on a fetch Response object or by reading a file via the File API in the browser.
Does the WASM engine support streaming data registration?
No, the current implementation in src/index.ts requires the complete dataset to be loaded into memory as a byte array or string before registration. The engine converts the entire input into an Arrow table during the registration call, making it suitable for analytical datasets that fit within the browser's memory constraints.
How do I handle CSV files with special characters or encodings?
Pass a configuration object as the third argument to registerCsv(). According to the CsvReadOptions interface defined in the SDK, you can specify the delimiter (e.g., "\t" for TSV), set header: true to skip header rows, and provide an explicit schema array with column names and Arrow types to override automatic inference.
Can I register data after calling loadMDL()?
No, registration must occur before loadMDL() executes. The semantic model loaded via loadMDL() binds table references to the registered datasets immediately. If you attempt to register additional data after loading the MDL, the engine will not recognize the new tables for existing queries, requiring you to reinitialize the engine and reload the model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →