# How GeoLibre Queries Apache Iceberg Tables Using DuckDB-WASM

> Discover how GeoLibre queries Apache Iceberg tables client-side with DuckDB-WASM. Learn to initialize the engine, load extensions, and register manifest files for seamless data access.

- Repository: [Open Geospatial Solutions/GeoLibre](https://github.com/opengeos/GeoLibre)
- Tags: how-to-guide
- Published: 2026-08-22

---

**GeoLibre enables client-side SQL queries against Apache Iceberg tables by initializing the DuckDB-WASM engine, dynamically loading the Iceberg extension, and registering remote manifest files through custom file system handlers.**

GeoLibre is an open-source geospatial platform that implements serverless analytics by running DuckDB directly in the browser via WebAssembly. According to the source code in the `opengeos/GeoLibre` repository, the application establishes a direct pipeline to Apache Iceberg tables using DuckDB-WASM's extension architecture, eliminating the need for backend query engines while processing partitioned Parquet data stored in cloud object storage.

## DuckDB-WASM Initialization and Extension Loading

The query process begins with the instantiation of the **DuckDB-WASM** engine. When a user configures an Iceberg data source, GeoLibre creates a new DuckDB instance within the browser's JavaScript environment.

In [`apps/geolibre-desktop/src/lib/iceberg.ts`](https://github.com/opengeos/GeoLibre/blob/main/apps/geolibre-desktop/src/lib/iceberg.ts), the initialization sequence installs the Iceberg extension that ships with the DuckDB-WASM build:

```typescript
import { DuckDB } from '@duckdb/duckdb-wasm';

// Initialize the WebAssembly engine
const db = await DuckDB.create({});

// Load the Iceberg extension
await db.run(`INSTALL iceberg; LOAD iceberg;`);

```

This extension provides DuckDB with the capability to parse Iceberg metadata files and understand the table's partitioning scheme without requiring server-side drivers.

## Registering Remote File Systems

Before querying can occur, GeoLibre must bridge the browser's security sandbox with remote storage endpoints. The [`apps/geolibre-desktop/src/lib/iceberg-loader.ts`](https://github.com/opengeos/GeoLibre/blob/main/apps/geolibre-desktop/src/lib/iceberg-loader.ts) module handles fetching Iceberg manifest files from locations such as Amazon S3 or HTTP servers.

The loader registers a custom file system with the DuckDB-WASM instance via `registerFileSystem()` to enable transparent access to remote URLs:

```typescript
// Register HTTP file system for remote manifests
await db.registerFileSystem('http', new HttpFileSystem());

```

This registration allows DuckDB-WASM to resolve file paths in the Iceberg metadata as HTTP requests, fetching only the necessary Parquet files during query execution rather than downloading entire datasets.

## Creating Virtual Iceberg Tables

With the extension loaded and file system registered, GeoLibre creates a virtual table that points to the Iceberg table location. This step translates the Iceberg catalog structure into a DuckDB-native table representation.

The SQL command executed via `db.run()` follows this pattern:

```sql
CREATE TABLE iceberg_demo
USING iceberg
LOCATION 'https://my-bucket.s3.amazonaws.com/iceberg/table/';

```

According to the implementation in [`apps/geolibre-desktop/src/lib/iceberg.ts`](https://github.com/opengeos/GeoLibre/blob/main/apps/geolibre-desktop/src/lib/iceberg.ts), this statement instructs DuckDB to read the Iceberg metadata, resolve the current snapshot, and map the underlying Parquet file paths to the virtual table schema.

## Executing Client-Side Queries

Once the virtual table is established, GeoLibre issues standard SQL queries through the DuckDB-WASM API. The engine pushes down predicates and reads only required column chunks from the remote Parquet files referenced in the Iceberg manifests.

```typescript
// Query the virtual table
const result = await db.query(`SELECT * FROM iceberg_demo LIMIT 5`);

// Result is returned as JavaScript objects
console.log(result);

```

The query results are returned as native JavaScript objects, which GeoLibre's rendering pipeline converts to GeoJSON for display on map layers. This entire workflow executes client-side, requiring only that the browser can access the remote storage endpoint.

## Key Source Files in GeoLibre

The Iceberg integration is implemented across three main files in the codebase:

### [`apps/geolibre-desktop/src/lib/iceberg.ts`](https://github.com/opengeos/GeoLibre/blob/main/apps/geolibre-desktop/src/lib/iceberg.ts)

This file contains core helper functions that manage the DuckDB-WASM connection lifecycle and execute Iceberg-specific SQL commands. It handles the `INSTALL iceberg` and `LOAD iceberg` sequence, virtual table creation, and query execution interfaces.

### [`apps/geolibre-desktop/src/lib/iceberg-loader.ts`](https://github.com/opengeos/GeoLibre/blob/main/apps/geolibre-desktop/src/lib/iceberg-loader.ts)

This module implements the manifest fetching logic, file system registration via `registerFileSystem()`, and preparation of virtual tables for DuckDB-WASM consumption. It bridges remote storage protocols (S3, HTTP) with DuckDB's file system abstraction.

### [`tests/iceberg.test.ts`](https://github.com/opengeos/GeoLibre/blob/main/tests/iceberg.test.ts)

The test suite validates end-to-end Iceberg table loading and querying functionality within the DuckDB-WASM environment, ensuring that the client-side pipeline correctly parses metadata and returns accurate query results.

## Summary

- GeoLibre uses **DuckDB-WASM** to run SQL queries directly in the browser without server-side components or data movement.
- The **Iceberg extension** is dynamically loaded into DuckDB-WASM to enable parsing of Iceberg metadata formats and snapshot isolation.
- Remote data access is achieved by registering **custom file system handlers** that translate file operations into HTTP requests against object storage.
- Virtual tables are created using the **`CREATE TABLE ... USING iceberg`** syntax to map Iceberg catalogs to queryable DuckDB tables.
- Query results are returned as **JavaScript objects** and rendered directly into GeoLibre's geospatial visualization layers.

## Frequently Asked Questions

### How does GeoLibre access Iceberg tables stored in private S3 buckets?

GeoLibre registers a custom file system implementation with DuckDB-WASM that can include authentication headers. When configuring an Iceberg source, users provide S3 credentials or pre-signed URLs that the [`iceberg-loader.ts`](https://github.com/opengeos/GeoLibre/blob/main/iceberg-loader.ts) module passes to the `registerFileSystem()` call, allowing the browser to authenticate requests to private buckets while maintaining query execution client-side.

### What version of DuckDB-WASM does GeoLibre use to support Iceberg tables?

The GeoLibre source code bundles DuckDB-WASM with the Iceberg extension pre-compiled for WebAssembly. The specific version is defined in the project's dependency configuration, ensuring compatibility with Iceberg table specification v1 and v2 formats supported by the underlying DuckDB engine.

### Can GeoLibre query Iceberg tables with partitioning?

Yes, GeoLibre leverages DuckDB's partition pruning capabilities through the Iceberg extension. When executing queries with `WHERE` clauses on partition columns, DuckDB-WASM reads only the relevant Iceberg manifests and underlying Parquet files, minimizing data transfer from remote storage and improving query performance in the browser.

### Is there a limit to the size of Iceberg tables GeoLibre can query?

The practical limit depends on browser memory constraints and network latency rather than GeoLibre's architecture. Since DuckDB-WASM streams Parquet data and applies predicate pushdown, GeoLibre can query terabyte-scale Iceberg tables, though result sets must fit within the browser's available memory (typically 2-4 GB) for final materialization and visualization.