# How viser_wrapper Enables Interactive 3D Visualization in the Browser

> Discover how viser_wrapper enables interactive 3D visualization in your browser. Stream point-cloud and camera data for real-time filtering and navigation.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-25

---

**The `viser_wrapper` function transforms GCT (global camera-tracking) predictions into a live, browser-based 3D scene by spawning a Viser server, building a reactive GUI, and streaming point-cloud and camera data over WebSockets to enable real-time filtering and navigation.**

The `viser_wrapper` module in the **Robbyant/lingbot-map** repository provides a lightweight interface for visualizing global camera-tracking predictions directly in a web browser. By leveraging the **Viser** library, it constructs an interactive 3D environment where users can inspect point clouds, toggle camera frustums, and navigate between frames without leaving their browser.

## Architectural Pipeline: Building the 3D Scene

The wrapper follows a ten-step pipeline defined in [`lingbot_map/vis/viser_wrapper.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/viser_wrapper.py) to convert raw prediction dictionaries into an explorable visual scene.

### Server Initialization and GUI Theme

The process begins by instantiating a `ViserServer` on the specified port and configuring a collapsible GUI theme. This establishes the WebSocket endpoint that serves the visualization to connected browsers.

```python

# Lines 60-61 in viser_wrapper.py

server = viser.ViserServer(port=port)

```

### Point Cloud Data Preparation

The wrapper accepts predictions containing either pre-computed `world_points` or raw depth maps requiring back-projection. When `use_point_map=False`, the function calls `unproject_depth_map_to_point_map` (defined in [`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py)) to convert depth and intrinsics into 3-D coordinates.

```python

# Lines 74-80 in viser_wrapper.py

if not use_point_map:
    world_points = unproject_depth_map_to_point_map(
        depth, extrinsic, intrinsic
    )

```

Optional sky masking via `apply_sky_segmentation` (lines 82-85) removes aerial points that clutter outdoor reconstructions. To maintain browser responsiveness, the implementation transposes images to `(S, H, W, 3)`, flattens point arrays, and applies random subsampling to cap the dataset at **6 million points** (lines 94-105).

### Scene Centering and Coordinate Optimization

To prevent floating-point precision issues in the renderer, the wrapper computes the centroid of all points and translates both the point cloud and camera poses to center the scene at the origin (lines 110-114).

```python

# Scene centering implementation (lines 110-114)

scene_center = np.mean(world_points, axis=0)
world_points -= scene_center
extrinsic[:, :3, 3] -= scene_center

```

### Interactive GUI Controls

The wrapper constructs three reactive widgets (lines 22-33) that allow users to manipulate the visualization without reloading the page:

- **Show Cameras**: A boolean checkbox toggling the visibility of camera frustums
- **Confidence Percent**: A slider filtering points based on their confidence scores
- **Show Points from Frames**: A dropdown selector isolating points from specific frame indices

### Scene Object Construction

The initial point cloud is added using `server.scene.add_point_cloud` with a crisp rendering configuration: a point size of `0.0005` and circular shape (lines 33-42). For each frame in the sequence, the wrapper creates:
- A `FrameHandle` displaying coordinate axes
- A `CameraFrustumHandle` textured with the original RGB image
- Click callbacks that teleport the viewer's camera to the selected pose

The `@frustum.on_click` decorator (implemented in the `visualize_frames` section, lines 47-94) registers handlers that update the client's camera `wxyz` (rotation) and `position`, enabling the "click-to-jump" navigation pattern.

### Reactive Update Mechanism

When GUI values change, the `on_update` callbacks (lines 95-27) recompute boolean masks based on the current confidence threshold and frame selection. Rather than rebuilding the scene, these callbacks mutate `point_cloud.points` and `point_cloud.colors` in-place, triggering Viser's WebSocket synchronization to push minimal deltas to all connected clients.

### Server Lifecycle Management

Finally, the wrapper handles execution flow based on the `background_mode` parameter. When `True`, a daemon thread maintains the server loop (lines 36-45), allowing the calling script to continue execution. Otherwise, the function blocks indefinitely to keep the browser session alive.

## Real-Time Interactivity Mechanisms

Three core technologies enable the seamless browser experience:

**WebSocket-Backed State Synchronization**
Viser maintains persistent WebSocket connections to every browser client. When Python code mutates scene objects (e.g., updating `point_cloud.points`), Viser serializes these changes and transmits them to the client in real-time, instantly updating the WebGL rendering without page refreshes.

**Click-to-Focus Camera Navigation**
Each camera frustum registers an `on_click` callback that overwrites the viewer's camera pose. This creates an intuitive interaction model where users click a camera icon in the 3D view to instantly adopt that perspective, facilitating inspection of specific viewpoints in the reconstruction.

**Performance Optimization Strategies**
The wrapper enforces a 6-million-point cap to ensure GPU memory and network bandwidth remain within browser limits. Pre-centering the scene at the origin reduces floating-point jitter during rendering, while the small point size (`0.0005`) minimizes overdraw on the client's GPU.

## Usage Examples

### Basic Visualization

Launch a viewer from a prediction dictionary containing `images`, `world_points`, `world_points_conf`, `depth`, `depth_conf`, `extrinsic`, and `intrinsic` keys:

```python
from lingbot_map.vis import viser_wrapper

# preds obtained from model inference

preds = {...}  

server = viser_wrapper(preds, port=8080)

# Navigate to http://localhost:8080

```

### Non-Blocking Background Mode

Run the server concurrently with other processing tasks:

```python
server = viser_wrapper(
    preds,
    port=8081,
    background_mode=True,  # Non-blocking execution

)

print("Visualization active at http://localhost:8081")

# Continue with post-processing...

```

### Advanced Configuration

Customize confidence thresholds and apply sky masking:

```python
server = viser_wrapper(
    preds,
    init_conf_threshold=30.0,   # Lower initial confidence cutoff

    mask_sky=True,              # Enable sky segmentation filtering

    image_folder="data/images", # Source images for sky detection

)

```

## Key Source Files

The visualization pipeline spans several modules in the **Robbyant/lingbot-map** repository:

- **[`lingbot_map/vis/viser_wrapper.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/viser_wrapper.py)**: Core orchestration logic that instantiates the Viser server, manages GUI state, and handles point-cloud updates.
- **[`lingbot_map/vis/point_cloud_viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/point_cloud_viewer.py)**: Extended viewer implementation providing a richer UI; the wrapper serves as a lightweight facade over this module.
- **[`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py)**: Mathematical utilities including `unproject_depth_map_to_point_map` and `closed_form_inverse_se3` for 3D coordinate transformations.
- **[`lingbot_map/vis/sky_segmentation.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/sky_segmentation.py)**: Computer vision routines for identifying and masking sky regions in outdoor captures.
- **[`benchmark/viewer.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/viewer.py)**: Reference implementation demonstrating additional GUI patterns and standalone Viser configurations.

## Summary

- The **`viser_wrapper`** converts GCT prediction dictionaries into interactive browser-based 3D scenes using the Viser library.
- It implements a **ten-step pipeline** spanning server initialization, point-cloud preprocessing (capped at 6M points), scene centering, and GUI construction.
- **Real-time interactivity** is achieved through WebSocket state synchronization, in-place point-cloud mutations, and click-to-focus camera callbacks.
- Users can filter points by **confidence threshold** and **frame index**, toggle camera frustums, and navigate poses via an intuitive web interface.
- The `background_mode` parameter controls whether the server blocks execution or runs in a daemon thread, enabling flexible integration into larger processing workflows.

## Frequently Asked Questions

### How does viser_wrapper handle large point clouds without crashing the browser?

The wrapper enforces a hard limit of approximately **6 million points** through random subsampling (lines 94-105 in [`viser_wrapper.py`](https://github.com/Robbyant/lingbot-map/blob/main/viser_wrapper.py)). It also centers the scene at the origin to reduce floating-point precision issues and uses a small point size (`0.0005`) to minimize GPU overdraw, ensuring smooth rendering on consumer hardware.

### Can I use viser_wrapper with depth maps instead of pre-computed 3D points?

Yes. Set `use_point_map=False` when calling the wrapper. The function will automatically invoke `unproject_depth_map_to_point_map` from [`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py) (lines 74-80) to convert your depth maps, extrinsics, and intrinsics into world-space point clouds before visualization.

### What triggers the camera "click-to-jump" functionality?

Each camera frustum created by `visualize_frames` (lines 47-94) registers an `on_click` callback using the `@frustum.on_click` decorator. When a user clicks a frustum in the browser, this callback updates the client's camera `wxyz` (quaternion rotation) and `position` vectors to match the selected frame's pose, instantly teleporting the view.

### Is it possible to run the visualization without blocking my Python script?

Yes. Pass `background_mode=True` to the wrapper (lines 36-45). This spawns a daemon thread to maintain the Viser server, allowing the function to return immediately while the browser session remains active. The default behavior (`background_mode=False`) blocks indefinitely to prevent the script from terminating the server prematurely.