How GeoLibre Handles Vector File Imports via DuckDB-WASM Spatial: Formats, Pipeline, and Code Examples

GeoLibre imports vector files client-side using DuckDB-WASM Spatial, supporting formats from GeoJSON and Shapefile to GeoParquet and CAD files through a unified pipeline that registers buffers, loads the Spatial extension, and materializes GeoJSON via SQL. This browser-based approach eliminates server round-trips while leveraging GDAL's format breadth through DuckDB's WebAssembly build.

In the opengeos/GeoLibre repository, the entire import system lives in apps/geolibre-desktop/src/lib/duckdb-vector-loader.ts. This article breaks down the exact pipeline, every supported format, and practical TypeScript examples you can run today.

The DuckDB-WASM Spatial Import Pipeline

GeoLibre's loader orchestrates eight distinct stages from file drop to GeoJSON FeatureCollection. Each stage is implemented as a concrete function call with specific error handling and fallback paths.

1. File Registration in DuckDB-WASM

Before any SQL executes, the file buffer enters DuckDB's virtual filesystem via db.registerFileBuffer. Auxiliary "sibling" files (like Shapefile sidecars) register simultaneously.

// From duckdb-vector-loader.ts#L23-L28
// Registers main file + any .shx, .dbf, .prj companions
await db.registerFileBuffer(file.name, file.data);

2. Spatial Extension Installation

The Spatial extension loads exactly once per database instance. The ensureSpatialExtension function guards against redundant loads.

// From duckdb-vector-loader.ts#L81-L89
await db.query(`INSTALL spatial`);
await db.query(`LOAD spatial`);

You can override the extension source via VITE_DUCKDB_SPATIAL_EXTENSION_PATH—useful for air-gapped deployments or custom builds.

3. SQL Source Generation

The loader branches based on file type:

  • Parquet/GeoParquet: Uses native read_parquet for columnar performance
  • All other formats: Uses ST_Read with optional layer= parameter
// From duckdb-vector-loader.ts#L33-L47
if (extension === 'parquet' || extension === 'geoparquet') {
  return `SELECT * FROM read_parquet('${fileName}')`;
} else {
  const layerClause = layer ? `, layer='${layer}'` : '';
  return `SELECT * FROM ST_Read('${fileName}'${layerClause})`;
}

4. Geometry Column Detection

The detectGeometryColumn function inspects DESCRIBE output to find geometry columns and determine if base64-WKB decoding is required.

// From duckdb-vector-loader.ts#L18-L23
const describeResult = await db.query(`DESCRIBE ${sql}`);
const geomCol = detectGeometryColumn(describeResult);

5. CRS Resolution with Fallbacks

Spatial reference systems resolve through multiple strategies:

  • ST_Read_Meta for embedded metadata
  • Side-car .prj files for Shapefiles
  • WKT string parsing when EPSG codes are absent
// From duckdb-vector-loader.ts#L76-L84 and #L103-L108
const meta = await db.query(`SELECT * FROM ST_Read_Meta('${fileName}')`);
const prjContent = await readSiblingFile(`${baseName}.prj`);

6. Large Dataset Guard

Before materializing massive files, confirmLargeDataset checks row counts against user-defined thresholds.

// From duckdb-vector-loader.ts#L44-L48
const countResult = await db.query(`SELECT count(*) FROM (${sql})`);
const featureCount = countResult.get(0)!.count;
await confirmLargeDataset({ name: file.name, featureCount });

7. GeoJSON Materialization

The geometryGeoJsonSql helper builds ST_AsGeoJSON expressions, then toFeatureCollection converts DuckDB rows to standard GeoJSON.

// From duckdb-vector-loader.ts#L50-L58
const geojsonSql = `SELECT ${geometryGeoJsonSql(geomCol)}, * EXCLUDE ${geomCol} FROM (${sql})`;
const features = await toFeatureCollection(await db.query(geojsonSql));

8. Fallback for Complex Surface Geometries

When ST_Read fails on ESRI MultiPatch TIN or similar surface types, the loader falls back to WKB parsing in JavaScript.

// From duckdb-vector-loader.ts#L86-L92 and #L104-L108
const wkbResult = await db.query(`SELECT ST_AsBinary(geom) as wkb FROM ST_Read('${fileName}', keep_wkb=true)`);
const features = decodeWkb(wkbResult);
const reprojected = await reprojectFeatureCollectionToWgs84(features, sourceCrs);

Supported Vector File Formats

GeoLibre's format support spans native DuckDB readers, GDAL drivers via Spatial extension, and specialized handlers for edge cases.

Extension DuckDB-WASM Path Notes
.geojson / .json ST_Read (GDAL) Standard GeoJSON with legacy crs member reprojection to WGS84
.shp (Shapefile) ST_Read (GDAL) .prj side-car consulted for CRS resolution
.gpkg (GeoPackage) Custom gpkg-reader.ts Avoids single-threaded WASM crash; falls back to ST_Read for non-SQLite files
.parquet / .geoparquet read_parquet (native) Direct columnar read without GDAL overhead
.csv read_csv_auto Requires explicit longitude/latitude column mapping via GeoParquetConversionOptions.csv
.dxf / .dwg (CAD) ST_Read with layer= Layer enumeration via readCadLayers
.kml, .gpx, .osm, others ST_Read (GDAL) Any format GDAL opens works through DuckDB-WASM Spatial

The GeoPackage handler deserves special attention. Due to a WASM threading limitation, apps/geolibre-desktop/src/lib/gpkg-reader.ts implements loadGeoPackageVectorFile as a dedicated path that bypasses ST_Read for genuine SQLite databases, falling back only when the file is misidentified.

Working Code Examples

Load a Generic Vector File

Drop this into any GeoLibre-compatible environment to load GeoJSON, Shapefile, or KML:

import { loadDuckDbVectorFile } from "./duckdb-vector-loader";

const file = {
  name: "mydata.geojson",
  extension: "geojson",
  data: await fetch("mydata.geojson").then(r => r.arrayBuffer()),
};

loadDuckDbVectorFile(file, {
  onLargeDataset: async ({ name, featureCount }) => {
    if (featureCount > 1_000_000) {
      if (!confirm(`Load ${featureCount} features from ${name}?`)) {
        throw new Error("User cancelled large dataset load");
      }
    }
  },
})
  .then(fc => {
    console.log("Loaded", fc.features.length, "features");
  })
  .catch(err => console.error("Vector import failed:", err));

Convert CSV to GeoParquet

Explicit column mapping for point data in delimited files:

import { convertDuckDbVectorToGeoParquet } from "./duckdb-vector-loader";

const csvFile = {
  name: "points.csv",
  extension: "csv",
  data: await fetch("points.csv").then(r => r.arrayBuffer()),
};

convertDuckDbVectorToGeoParquet(csvFile, {
  csv: { lonColumn: "lon", latColumn: "lat" },
}).then(result => {
  console.log("GeoParquet size:", result.data.byteLength);
});

Enumerate CAD Layers Before Loading

Multi-layer CAD files require layer selection before full import:

import { readCadLayers } from "./duckdb-vector-loader";

const cadFile = {
  name: "plan.dwg",
  extension: "dwg",
  data: await fetch("plan.dwg").then(r => r.arrayBuffer()),
};

readCadLayers(cadFile).then(layers => {
  console.log("Available CAD layers:", layers);
  // Subsequent loadDuckDbVectorFile call with selected layer
});

Key Source Files

Understanding the module structure helps with debugging and extension:

File Purpose
apps/geolibre-desktop/src/lib/duckdb-vector-loader.ts Main orchestration: registration, extension loading, SQL generation, CRS handling, guards, conversion
apps/geolibre-desktop/src/lib/duckdb-geometry.ts Geometry detection, ST_AsGeoJSON builders, row-to-GeoJSON conversion
apps/geolibre-desktop/src/lib/duckdb-vector-guard.ts Large dataset confirmation logic
apps/geolibre-desktop/src/lib/gpkg-reader.ts Specialized GeoPackage handling to avoid WASM crashes
apps/geolibre-desktop/src/lib/spatial-extension-config.ts Environment variable parsing for custom extension paths
apps/geolibre-desktop/src/hooks/useAddData.tsx UI integration point for drag-drop and file picker flows

Summary

  • DuckDB-WASM Spatial powers all client-side vector imports in GeoLibre, combining DuckDB's SQL engine with GDAL's format breadth in WebAssembly.
  • Eight pipeline stages—from registerFileBuffer to ST_AsGeoJSON materialization—handle everything from CRS resolution to large dataset guards.
  • Native Parquet/GeoParquet reads skip GDAL entirely for superior performance, while ST_Read covers 100+ GDAL-supported formats.
  • Specialized handlers exist for GeoPackage (crash avoidance), CAD (layer enumeration), and complex surface geometries (WKB fallback).
  • Environment variable VITE_DUCKDB_SPATIAL_EXTENSION_PATH enables custom Spatial extension deployment.

Frequently Asked Questions

Does GeoLibre require a server to process vector files?

No. All processing happens in the browser through DuckDB-WASM Spatial. The file buffer registers directly with the in-memory DuckDB instance, executes SQL client-side, and returns a GeoJSON FeatureCollection without network transmission.

Why does GeoPackage use a separate reader instead of ST_Read?

The standard ST_Read path triggers a single-threaded WASM crash with GeoPackage files due to SQLite locking behavior. The custom gpkg-reader.ts implementation opens the database directly, avoiding this limitation while maintaining format compatibility.

Can I import vector files larger than available RAM?

The confirmLargeDataset guard warns before materializing large datasets, but DuckDB-WASM currently operates in-memory. For files exceeding browser memory limits, the loader will fail during registerFileBuffer or query execution. Consider converting to GeoParquet and using row group filtering for partial reads.

How do I add support for a new vector format?

If GDAL supports it, DuckDB-WASM Spatial already does—no changes needed. For formats requiring special handling (like GeoPackage), implement a dedicated reader in apps/geolibre-desktop/src/lib/ and branch the loader logic based on file extension detection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →