# How to Use `manifest.json` for Programmatic File Discovery and Filtering in BBL Tracker

> Learn to use manifest.json for programmatic file discovery and filtering in BBL Tracker. Access Parquet files efficiently without listing entire storage buckets.

- Repository: [Nelson Chen/bbl-tracker-public-db](https://github.com/nelsonjchen/bbl-tracker-public-db)
- Tags: how-to-guide
- Published: 2026-03-08

---

**The BBL Tracker public dataset provides a [`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json) file that acts as a centralized index of all available Parquet files, enabling efficient discovery and filtering without listing the entire Cloudflare R2 storage bucket.**

The `nelsonjchen/bbl-tracker-public-db` repository hosts the Bambu Lab Store Filament Tracker dataset as hourly Parquet files. To simplify programmatic access, the dataset includes a **[`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json)** file that serves as a machine-readable index for file discovery and filtering, eliminating the need to enumerate the bucket directly.

## What Is the Manifest and Where to Find It

The [`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json) lives at the root of the public dataset and can be fetched directly from:

```

https://db-public.bbltracker.com/manifest.json

```

According to the source code in [`README.md`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/README.md) (lines 189-204), the manifest contains a JSON object with two top-level fields:

- **`generated`** – the ISO-8601 timestamp of when the manifest was created.
- **`files`** – a map whose keys are the filenames of the Parquet files (e.g., `2026-02-16-0000.parquet`) and whose values contain basic metadata such as the number of rows in each file.

### Manifest Structure and Schema

The concrete structure follows this pattern:

```json
{
  "generated": "2026-02-16T17:30:00.000Z",
  "files": {
    "2026-02-16-0000.parquet": { "rows": 1500 },
    "2026-02-16-0600.parquet": { "rows": 840 },
    "...": { }
  }
}

```

## Why Use [`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json) for File Discovery?

Using the manifest provides three distinct advantages over bucket listing:

- **Discovery:** The complete list of files is obtained with a single HTTP request, avoiding extra API calls and respecting rate limits imposed by the storage backend.
- **Filtering:** Since each key follows a timestamped naming convention (`YYYY-MM-DD-HHMM.parquet`), you can filter on date ranges, specific hours, or custom logic before downloading any data.
- **Metadata-Driven Decisions:** The `rows` count lets a client decide whether to skip very small files or prioritize larger ones, optimizing bandwidth and processing time.

## Programmatic File Discovery Workflow

The architectural flow implemented in [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py) and [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) follows five steps:

1. **Fetch the Manifest** – a single `GET` request returns the complete index.
2. **Parse JSON** – deserialize into a dictionary or object.
3. **Filter Keys** – apply date-range, naming pattern, or metadata filters.
4. **Construct Full URLs** – prepend `https://db-public.bbltracker.com/` to each selected filename.
5. **Pass URLs to DuckDB** – DuckDB can read a list of remote Parquet files directly, enabling bulk analysis with one query.

Both [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) and [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py) define a `manifest_url` constant pointing to `https://db-public.bbltracker.com/manifest.json`, demonstrating this pattern in production code.

### Fetching and Parsing the Manifest

The simplest implementation uses Python's standard library:

```python
import json
import urllib.request

MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"

with urllib.request.urlopen(MANIFEST_URL) as resp:
    manifest = json.load(resp)

print(f"Manifest generated at: {manifest['generated']}")
print(f"Total files indexed: {len(manifest['files'])}")

```

### Filtering Files by Date Range

Because filenames follow the pattern `YYYY-MM-DD-HHMM.parquet`, string prefix matching efficiently isolates specific time periods:

```python

# Filter for February 2026 files

feb_files = [
    fname for fname in manifest["files"]
    if fname.startswith("2026-02-")
]

```

### Filtering by Metadata (Row Count)

The manifest includes row counts to help clients avoid downloading insignificant files:

```python

# Keep only files with at least 1,000 rows

substantial_files = [
    fname for fname, meta in manifest["files"].items()
    if meta.get("rows", 0) >= 1000
]

```

## Practical Code Examples

### Python: List All February 2026 Files

This complete example demonstrates fetching, filtering, and constructing full URLs:

```python
import json
import urllib.request

MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"
BASE_URL = "https://db-public.bbltracker.com/"

# 1️⃣ Fetch the manifest

with urllib.request.urlopen(MANIFEST_URL) as resp:
    manifest = json.load(resp)

# 2️⃣ Filter for February 2026 files

feb_files = [
    BASE_URL + fname
    for fname in manifest["files"]
    if fname.startswith("2026-02-")
]

print(f"Found {len(feb_files)} files for February 2026")
print("\n".join(feb_files[:5]))   # show first few URLs

```

### Python: Skip Small Files Using Row Count Metadata

Optimize bandwidth by filtering out files with minimal data:

```python
import json
import urllib.request

MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"
BASE_URL = "https://db-public.bbltracker.com/"

manifest = json.load(urllib.request.urlopen(MANIFEST_URL))

# Keep only files with at least 1,000 rows

large_files = [
    BASE_URL + name
    for name, meta in manifest["files"].items()
    if meta.get("rows", 0) >= 1000
]

print(f"{len(large_files)} files have ≥1,000 rows")

```

### DuckDB: Direct Query Using Filtered URLs

DuckDB can consume a list of remote Parquet URLs directly, enabling analysis without intermediate downloads:

```python
import duckdb
import json
import urllib.request

MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"

manifest = json.load(urllib.request.urlopen(MANIFEST_URL))

# Example: all files for 2026-02-16

target_files = [
    "https://db-public.bbltracker.com/" + name
    for name in manifest["files"]
    if name.startswith("2026-02-16-")
]

# DuckDB can read a list of URLs in one call

df = duckdb.read_parquet(target_files).df()
print(df.head())

```

### Node.js: File Discovery with Fetch and DuckDB

The same pattern works in JavaScript environments:

```javascript
const fetch = require('node-fetch');
const duckdb = require('duckdb');

(async () => {
  const manifest = await fetch('https://db-public.bbltracker.com/manifest.json')
                       .then(r => r.json());

  const urls = Object.keys(manifest.files)
    .filter(name => name.startsWith('2026-02-16-'))
    .map(name => `https://db-public.bbltracker.com/${name}`);

  const db = new duckdb.Database();
  const con = await db.connect();
  const result = await con.query(`SELECT * FROM read_parquet('${urls.join("','")}') LIMIT 5`);
  console.log(result);
})();

```

## Key Repository Files

The `nelsonjchen/bbl-tracker-public-db` repository includes several reference implementations that demonstrate manifest usage:

| File | Role | Link |
|------|------|------|
| [`README.md`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/README.md) | Primary documentation; explains the manifest format and provides usage patterns. | [README.md](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/master/README.md) |
| [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) | Example script that references `manifest_url` and shows how to download a single Parquet file. | [script.py](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/master/script.py) |
| [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py) | Utility that pulls the manifest, filters a date range, and builds a local DuckDB file with the last 30 days of data. | [reconstruct_db.py](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/master/reconstruct_db.py) |
| [`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json) (remote) | Index of all dataset files; fetched from the public bucket. | <https://db-public.bbltracker.com/manifest.json> |

These files together illustrate the complete workflow: **fetch the manifest → filter → construct URLs → query with DuckDB**. By leveraging [`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json), clients can efficiently discover and select exactly the data they need without enumerating the entire storage bucket.

## Summary

- The **[`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json)** file at `https://db-public.bbltracker.com/manifest.json` serves as a centralized index for the BBL Tracker public dataset.
- It contains a **`generated`** timestamp and a **`files`** dictionary mapping Parquet filenames to metadata (including row counts).
- **Programmatic file discovery** involves fetching the JSON, filtering keys by date prefixes or metadata values, and prepending the base URL to construct full paths.
- **DuckDB** can consume a list of remote URLs directly via `read_parquet()`, enabling analysis without intermediate downloads.
- Reference implementations in [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) and [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py) demonstrate production-ready patterns for manifest consumption.

## Frequently Asked Questions

### What is the URL for the manifest.json file?

The manifest is hosted at `https://db-public.bbltracker.com/manifest.json`. This endpoint returns a JSON object containing the `generated` timestamp and the `files` dictionary that indexes all available Parquet files in the dataset.

### How often is the manifest.json updated?

According to the repository structure, the manifest is regenerated periodically to reflect new hourly Parquet files as they are added to the Cloudflare R2 bucket. The `generated` field in the JSON provides the exact ISO-8601 timestamp of when that specific manifest version was created.

### Can I use the manifest to filter files by specific date ranges?

Yes. Because filenames follow the pattern `YYYY-MM-DD-HHMM.parquet`, you can filter the `files` dictionary keys using string prefix matching (e.g., `name.startswith("2026-02-")` for February 2026). This allows precise date-range selection without downloading unnecessary files.

### What metadata is available for each file in the manifest?

Each file entry in the `files` dictionary currently includes a `rows` field indicating the number of data rows in that Parquet file. This metadata allows clients to skip very small files or prioritize larger datasets based on their specific analysis requirements.