How GeoLibre Handles Vector File Imports via DuckDB-WASM Spatial: Formats, Pipeline, and Code Examples
GeoLibre imports vector files client-side using DuckDB-WASM Spatial, supporting formats from GeoJSON and Shapefile to GeoParquet and CAD files through a unified pipeline that registers buffers, loads the Spatial extension, and materializes GeoJSON via SQL. This browser-based approach eliminates server round-trips while leveraging GDAL's format breadth through DuckDB's WebAssembly build.
In the opengeos/GeoLibre repository, the entire import system lives in apps/geolibre-desktop/src/lib/duckdb-vector-loader.ts. This article breaks down the exact pipeline, every supported format, and practical TypeScript examples you can run today.
The DuckDB-WASM Spatial Import Pipeline
GeoLibre's loader orchestrates eight distinct stages from file drop to GeoJSON FeatureCollection. Each stage is implemented as a concrete function call with specific error handling and fallback paths.
1. File Registration in DuckDB-WASM
Before any SQL executes, the file buffer enters DuckDB's virtual filesystem via db.registerFileBuffer. Auxiliary "sibling" files (like Shapefile sidecars) register simultaneously.
// From duckdb-vector-loader.ts#L23-L28
// Registers main file + any .shx, .dbf, .prj companions
await db.registerFileBuffer(file.name, file.data);
2. Spatial Extension Installation
The Spatial extension loads exactly once per database instance. The ensureSpatialExtension function guards against redundant loads.
// From duckdb-vector-loader.ts#L81-L89
await db.query(`INSTALL spatial`);
await db.query(`LOAD spatial`);
You can override the extension source via VITE_DUCKDB_SPATIAL_EXTENSION_PATH—useful for air-gapped deployments or custom builds.
3. SQL Source Generation
The loader branches based on file type:
- Parquet/GeoParquet: Uses native
read_parquetfor columnar performance - All other formats: Uses
ST_Readwith optionallayer=parameter
// From duckdb-vector-loader.ts#L33-L47
if (extension === 'parquet' || extension === 'geoparquet') {
return `SELECT * FROM read_parquet('${fileName}')`;
} else {
const layerClause = layer ? `, layer='${layer}'` : '';
return `SELECT * FROM ST_Read('${fileName}'${layerClause})`;
}
4. Geometry Column Detection
The detectGeometryColumn function inspects DESCRIBE output to find geometry columns and determine if base64-WKB decoding is required.
// From duckdb-vector-loader.ts#L18-L23
const describeResult = await db.query(`DESCRIBE ${sql}`);
const geomCol = detectGeometryColumn(describeResult);
5. CRS Resolution with Fallbacks
Spatial reference systems resolve through multiple strategies:
ST_Read_Metafor embedded metadata- Side-car
.prjfiles for Shapefiles - WKT string parsing when EPSG codes are absent
// From duckdb-vector-loader.ts#L76-L84 and #L103-L108
const meta = await db.query(`SELECT * FROM ST_Read_Meta('${fileName}')`);
const prjContent = await readSiblingFile(`${baseName}.prj`);
6. Large Dataset Guard
Before materializing massive files, confirmLargeDataset checks row counts against user-defined thresholds.
// From duckdb-vector-loader.ts#L44-L48
const countResult = await db.query(`SELECT count(*) FROM (${sql})`);
const featureCount = countResult.get(0)!.count;
await confirmLargeDataset({ name: file.name, featureCount });
7. GeoJSON Materialization
The geometryGeoJsonSql helper builds ST_AsGeoJSON expressions, then toFeatureCollection converts DuckDB rows to standard GeoJSON.
// From duckdb-vector-loader.ts#L50-L58
const geojsonSql = `SELECT ${geometryGeoJsonSql(geomCol)}, * EXCLUDE ${geomCol} FROM (${sql})`;
const features = await toFeatureCollection(await db.query(geojsonSql));
8. Fallback for Complex Surface Geometries
When ST_Read fails on ESRI MultiPatch TIN or similar surface types, the loader falls back to WKB parsing in JavaScript.
// From duckdb-vector-loader.ts#L86-L92 and #L104-L108
const wkbResult = await db.query(`SELECT ST_AsBinary(geom) as wkb FROM ST_Read('${fileName}', keep_wkb=true)`);
const features = decodeWkb(wkbResult);
const reprojected = await reprojectFeatureCollectionToWgs84(features, sourceCrs);
Supported Vector File Formats
GeoLibre's format support spans native DuckDB readers, GDAL drivers via Spatial extension, and specialized handlers for edge cases.
| Extension | DuckDB-WASM Path | Notes |
|---|---|---|
.geojson / .json |
ST_Read (GDAL) |
Standard GeoJSON with legacy crs member reprojection to WGS84 |
.shp (Shapefile) |
ST_Read (GDAL) |
.prj side-car consulted for CRS resolution |
.gpkg (GeoPackage) |
Custom gpkg-reader.ts |
Avoids single-threaded WASM crash; falls back to ST_Read for non-SQLite files |
.parquet / .geoparquet |
read_parquet (native) |
Direct columnar read without GDAL overhead |
.csv |
read_csv_auto |
Requires explicit longitude/latitude column mapping via GeoParquetConversionOptions.csv |
.dxf / .dwg (CAD) |
ST_Read with layer= |
Layer enumeration via readCadLayers |
.kml, .gpx, .osm, others |
ST_Read (GDAL) |
Any format GDAL opens works through DuckDB-WASM Spatial |
The GeoPackage handler deserves special attention. Due to a WASM threading limitation, apps/geolibre-desktop/src/lib/gpkg-reader.ts implements loadGeoPackageVectorFile as a dedicated path that bypasses ST_Read for genuine SQLite databases, falling back only when the file is misidentified.
Working Code Examples
Load a Generic Vector File
Drop this into any GeoLibre-compatible environment to load GeoJSON, Shapefile, or KML:
import { loadDuckDbVectorFile } from "./duckdb-vector-loader";
const file = {
name: "mydata.geojson",
extension: "geojson",
data: await fetch("mydata.geojson").then(r => r.arrayBuffer()),
};
loadDuckDbVectorFile(file, {
onLargeDataset: async ({ name, featureCount }) => {
if (featureCount > 1_000_000) {
if (!confirm(`Load ${featureCount} features from ${name}?`)) {
throw new Error("User cancelled large dataset load");
}
}
},
})
.then(fc => {
console.log("Loaded", fc.features.length, "features");
})
.catch(err => console.error("Vector import failed:", err));
Convert CSV to GeoParquet
Explicit column mapping for point data in delimited files:
import { convertDuckDbVectorToGeoParquet } from "./duckdb-vector-loader";
const csvFile = {
name: "points.csv",
extension: "csv",
data: await fetch("points.csv").then(r => r.arrayBuffer()),
};
convertDuckDbVectorToGeoParquet(csvFile, {
csv: { lonColumn: "lon", latColumn: "lat" },
}).then(result => {
console.log("GeoParquet size:", result.data.byteLength);
});
Enumerate CAD Layers Before Loading
Multi-layer CAD files require layer selection before full import:
import { readCadLayers } from "./duckdb-vector-loader";
const cadFile = {
name: "plan.dwg",
extension: "dwg",
data: await fetch("plan.dwg").then(r => r.arrayBuffer()),
};
readCadLayers(cadFile).then(layers => {
console.log("Available CAD layers:", layers);
// Subsequent loadDuckDbVectorFile call with selected layer
});
Key Source Files
Understanding the module structure helps with debugging and extension:
| File | Purpose |
|---|---|
apps/geolibre-desktop/src/lib/duckdb-vector-loader.ts |
Main orchestration: registration, extension loading, SQL generation, CRS handling, guards, conversion |
apps/geolibre-desktop/src/lib/duckdb-geometry.ts |
Geometry detection, ST_AsGeoJSON builders, row-to-GeoJSON conversion |
apps/geolibre-desktop/src/lib/duckdb-vector-guard.ts |
Large dataset confirmation logic |
apps/geolibre-desktop/src/lib/gpkg-reader.ts |
Specialized GeoPackage handling to avoid WASM crashes |
apps/geolibre-desktop/src/lib/spatial-extension-config.ts |
Environment variable parsing for custom extension paths |
apps/geolibre-desktop/src/hooks/useAddData.tsx |
UI integration point for drag-drop and file picker flows |
Summary
- DuckDB-WASM Spatial powers all client-side vector imports in GeoLibre, combining DuckDB's SQL engine with GDAL's format breadth in WebAssembly.
- Eight pipeline stages—from
registerFileBuffertoST_AsGeoJSONmaterialization—handle everything from CRS resolution to large dataset guards. - Native Parquet/GeoParquet reads skip GDAL entirely for superior performance, while
ST_Readcovers 100+ GDAL-supported formats. - Specialized handlers exist for GeoPackage (crash avoidance), CAD (layer enumeration), and complex surface geometries (WKB fallback).
- Environment variable
VITE_DUCKDB_SPATIAL_EXTENSION_PATHenables custom Spatial extension deployment.
Frequently Asked Questions
Does GeoLibre require a server to process vector files?
No. All processing happens in the browser through DuckDB-WASM Spatial. The file buffer registers directly with the in-memory DuckDB instance, executes SQL client-side, and returns a GeoJSON FeatureCollection without network transmission.
Why does GeoPackage use a separate reader instead of ST_Read?
The standard ST_Read path triggers a single-threaded WASM crash with GeoPackage files due to SQLite locking behavior. The custom gpkg-reader.ts implementation opens the database directly, avoiding this limitation while maintaining format compatibility.
Can I import vector files larger than available RAM?
The confirmLargeDataset guard warns before materializing large datasets, but DuckDB-WASM currently operates in-memory. For files exceeding browser memory limits, the loader will fail during registerFileBuffer or query execution. Consider converting to GeoParquet and using row group filtering for partial reads.
How do I add support for a new vector format?
If GDAL supports it, DuckDB-WASM Spatial already does—no changes needed. For formats requiring special handling (like GeoPackage), implement a dedicated reader in apps/geolibre-desktop/src/lib/ and branch the loader logic based on file extension detection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →