How GeoLibre Queries Apache Iceberg Tables Using DuckDB-WASM
GeoLibre enables client-side SQL queries against Apache Iceberg tables by initializing the DuckDB-WASM engine, dynamically loading the Iceberg extension, and registering remote manifest files through custom file system handlers.
GeoLibre is an open-source geospatial platform that implements serverless analytics by running DuckDB directly in the browser via WebAssembly. According to the source code in the opengeos/GeoLibre repository, the application establishes a direct pipeline to Apache Iceberg tables using DuckDB-WASM's extension architecture, eliminating the need for backend query engines while processing partitioned Parquet data stored in cloud object storage.
DuckDB-WASM Initialization and Extension Loading
The query process begins with the instantiation of the DuckDB-WASM engine. When a user configures an Iceberg data source, GeoLibre creates a new DuckDB instance within the browser's JavaScript environment.
In apps/geolibre-desktop/src/lib/iceberg.ts, the initialization sequence installs the Iceberg extension that ships with the DuckDB-WASM build:
import { DuckDB } from '@duckdb/duckdb-wasm';
// Initialize the WebAssembly engine
const db = await DuckDB.create({});
// Load the Iceberg extension
await db.run(`INSTALL iceberg; LOAD iceberg;`);
This extension provides DuckDB with the capability to parse Iceberg metadata files and understand the table's partitioning scheme without requiring server-side drivers.
Registering Remote File Systems
Before querying can occur, GeoLibre must bridge the browser's security sandbox with remote storage endpoints. The apps/geolibre-desktop/src/lib/iceberg-loader.ts module handles fetching Iceberg manifest files from locations such as Amazon S3 or HTTP servers.
The loader registers a custom file system with the DuckDB-WASM instance via registerFileSystem() to enable transparent access to remote URLs:
// Register HTTP file system for remote manifests
await db.registerFileSystem('http', new HttpFileSystem());
This registration allows DuckDB-WASM to resolve file paths in the Iceberg metadata as HTTP requests, fetching only the necessary Parquet files during query execution rather than downloading entire datasets.
Creating Virtual Iceberg Tables
With the extension loaded and file system registered, GeoLibre creates a virtual table that points to the Iceberg table location. This step translates the Iceberg catalog structure into a DuckDB-native table representation.
The SQL command executed via db.run() follows this pattern:
CREATE TABLE iceberg_demo
USING iceberg
LOCATION 'https://my-bucket.s3.amazonaws.com/iceberg/table/';
According to the implementation in apps/geolibre-desktop/src/lib/iceberg.ts, this statement instructs DuckDB to read the Iceberg metadata, resolve the current snapshot, and map the underlying Parquet file paths to the virtual table schema.
Executing Client-Side Queries
Once the virtual table is established, GeoLibre issues standard SQL queries through the DuckDB-WASM API. The engine pushes down predicates and reads only required column chunks from the remote Parquet files referenced in the Iceberg manifests.
// Query the virtual table
const result = await db.query(`SELECT * FROM iceberg_demo LIMIT 5`);
// Result is returned as JavaScript objects
console.log(result);
The query results are returned as native JavaScript objects, which GeoLibre's rendering pipeline converts to GeoJSON for display on map layers. This entire workflow executes client-side, requiring only that the browser can access the remote storage endpoint.
Key Source Files in GeoLibre
The Iceberg integration is implemented across three main files in the codebase:
apps/geolibre-desktop/src/lib/iceberg.ts
This file contains core helper functions that manage the DuckDB-WASM connection lifecycle and execute Iceberg-specific SQL commands. It handles the INSTALL iceberg and LOAD iceberg sequence, virtual table creation, and query execution interfaces.
apps/geolibre-desktop/src/lib/iceberg-loader.ts
This module implements the manifest fetching logic, file system registration via registerFileSystem(), and preparation of virtual tables for DuckDB-WASM consumption. It bridges remote storage protocols (S3, HTTP) with DuckDB's file system abstraction.
tests/iceberg.test.ts
The test suite validates end-to-end Iceberg table loading and querying functionality within the DuckDB-WASM environment, ensuring that the client-side pipeline correctly parses metadata and returns accurate query results.
Summary
- GeoLibre uses DuckDB-WASM to run SQL queries directly in the browser without server-side components or data movement.
- The Iceberg extension is dynamically loaded into DuckDB-WASM to enable parsing of Iceberg metadata formats and snapshot isolation.
- Remote data access is achieved by registering custom file system handlers that translate file operations into HTTP requests against object storage.
- Virtual tables are created using the
CREATE TABLE ... USING icebergsyntax to map Iceberg catalogs to queryable DuckDB tables. - Query results are returned as JavaScript objects and rendered directly into GeoLibre's geospatial visualization layers.
Frequently Asked Questions
How does GeoLibre access Iceberg tables stored in private S3 buckets?
GeoLibre registers a custom file system implementation with DuckDB-WASM that can include authentication headers. When configuring an Iceberg source, users provide S3 credentials or pre-signed URLs that the iceberg-loader.ts module passes to the registerFileSystem() call, allowing the browser to authenticate requests to private buckets while maintaining query execution client-side.
What version of DuckDB-WASM does GeoLibre use to support Iceberg tables?
The GeoLibre source code bundles DuckDB-WASM with the Iceberg extension pre-compiled for WebAssembly. The specific version is defined in the project's dependency configuration, ensuring compatibility with Iceberg table specification v1 and v2 formats supported by the underlying DuckDB engine.
Can GeoLibre query Iceberg tables with partitioning?
Yes, GeoLibre leverages DuckDB's partition pruning capabilities through the Iceberg extension. When executing queries with WHERE clauses on partition columns, DuckDB-WASM reads only the relevant Iceberg manifests and underlying Parquet files, minimizing data transfer from remote storage and improving query performance in the browser.
Is there a limit to the size of Iceberg tables GeoLibre can query?
The practical limit depends on browser memory constraints and network latency rather than GeoLibre's architecture. Since DuckDB-WASM streams Parquet data and applies predicate pushdown, GeoLibre can query terabyte-scale Iceberg tables, though result sets must fit within the browser's available memory (typically 2-4 GB) for final materialization and visualization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →