How to Use `manifest.json` for Programmatic File Discovery and Filtering in BBL Tracker
The BBL Tracker public dataset provides a manifest.json file that acts as a centralized index of all available Parquet files, enabling efficient discovery and filtering without listing the entire Cloudflare R2 storage bucket.
The nelsonjchen/bbl-tracker-public-db repository hosts the Bambu Lab Store Filament Tracker dataset as hourly Parquet files. To simplify programmatic access, the dataset includes a manifest.json file that serves as a machine-readable index for file discovery and filtering, eliminating the need to enumerate the bucket directly.
What Is the Manifest and Where to Find It
The manifest.json lives at the root of the public dataset and can be fetched directly from:
https://db-public.bbltracker.com/manifest.json
According to the source code in README.md (lines 189-204), the manifest contains a JSON object with two top-level fields:
generated– the ISO-8601 timestamp of when the manifest was created.files– a map whose keys are the filenames of the Parquet files (e.g.,2026-02-16-0000.parquet) and whose values contain basic metadata such as the number of rows in each file.
Manifest Structure and Schema
The concrete structure follows this pattern:
{
"generated": "2026-02-16T17:30:00.000Z",
"files": {
"2026-02-16-0000.parquet": { "rows": 1500 },
"2026-02-16-0600.parquet": { "rows": 840 },
"...": { }
}
}
Why Use manifest.json for File Discovery?
Using the manifest provides three distinct advantages over bucket listing:
- Discovery: The complete list of files is obtained with a single HTTP request, avoiding extra API calls and respecting rate limits imposed by the storage backend.
- Filtering: Since each key follows a timestamped naming convention (
YYYY-MM-DD-HHMM.parquet), you can filter on date ranges, specific hours, or custom logic before downloading any data. - Metadata-Driven Decisions: The
rowscount lets a client decide whether to skip very small files or prioritize larger ones, optimizing bandwidth and processing time.
Programmatic File Discovery Workflow
The architectural flow implemented in reconstruct_db.py and script.py follows five steps:
- Fetch the Manifest – a single
GETrequest returns the complete index. - Parse JSON – deserialize into a dictionary or object.
- Filter Keys – apply date-range, naming pattern, or metadata filters.
- Construct Full URLs – prepend
https://db-public.bbltracker.com/to each selected filename. - Pass URLs to DuckDB – DuckDB can read a list of remote Parquet files directly, enabling bulk analysis with one query.
Both script.py and reconstruct_db.py define a manifest_url constant pointing to https://db-public.bbltracker.com/manifest.json, demonstrating this pattern in production code.
Fetching and Parsing the Manifest
The simplest implementation uses Python's standard library:
import json
import urllib.request
MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"
with urllib.request.urlopen(MANIFEST_URL) as resp:
manifest = json.load(resp)
print(f"Manifest generated at: {manifest['generated']}")
print(f"Total files indexed: {len(manifest['files'])}")
Filtering Files by Date Range
Because filenames follow the pattern YYYY-MM-DD-HHMM.parquet, string prefix matching efficiently isolates specific time periods:
# Filter for February 2026 files
feb_files = [
fname for fname in manifest["files"]
if fname.startswith("2026-02-")
]
Filtering by Metadata (Row Count)
The manifest includes row counts to help clients avoid downloading insignificant files:
# Keep only files with at least 1,000 rows
substantial_files = [
fname for fname, meta in manifest["files"].items()
if meta.get("rows", 0) >= 1000
]
Practical Code Examples
Python: List All February 2026 Files
This complete example demonstrates fetching, filtering, and constructing full URLs:
import json
import urllib.request
MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"
BASE_URL = "https://db-public.bbltracker.com/"
# 1️⃣ Fetch the manifest
with urllib.request.urlopen(MANIFEST_URL) as resp:
manifest = json.load(resp)
# 2️⃣ Filter for February 2026 files
feb_files = [
BASE_URL + fname
for fname in manifest["files"]
if fname.startswith("2026-02-")
]
print(f"Found {len(feb_files)} files for February 2026")
print("\n".join(feb_files[:5])) # show first few URLs
Python: Skip Small Files Using Row Count Metadata
Optimize bandwidth by filtering out files with minimal data:
import json
import urllib.request
MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"
BASE_URL = "https://db-public.bbltracker.com/"
manifest = json.load(urllib.request.urlopen(MANIFEST_URL))
# Keep only files with at least 1,000 rows
large_files = [
BASE_URL + name
for name, meta in manifest["files"].items()
if meta.get("rows", 0) >= 1000
]
print(f"{len(large_files)} files have ≥1,000 rows")
DuckDB: Direct Query Using Filtered URLs
DuckDB can consume a list of remote Parquet URLs directly, enabling analysis without intermediate downloads:
import duckdb
import json
import urllib.request
MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"
manifest = json.load(urllib.request.urlopen(MANIFEST_URL))
# Example: all files for 2026-02-16
target_files = [
"https://db-public.bbltracker.com/" + name
for name in manifest["files"]
if name.startswith("2026-02-16-")
]
# DuckDB can read a list of URLs in one call
df = duckdb.read_parquet(target_files).df()
print(df.head())
Node.js: File Discovery with Fetch and DuckDB
The same pattern works in JavaScript environments:
const fetch = require('node-fetch');
const duckdb = require('duckdb');
(async () => {
const manifest = await fetch('https://db-public.bbltracker.com/manifest.json')
.then(r => r.json());
const urls = Object.keys(manifest.files)
.filter(name => name.startsWith('2026-02-16-'))
.map(name => `https://db-public.bbltracker.com/${name}`);
const db = new duckdb.Database();
const con = await db.connect();
const result = await con.query(`SELECT * FROM read_parquet('${urls.join("','")}') LIMIT 5`);
console.log(result);
})();
Key Repository Files
The nelsonjchen/bbl-tracker-public-db repository includes several reference implementations that demonstrate manifest usage:
| File | Role | Link |
|---|---|---|
README.md |
Primary documentation; explains the manifest format and provides usage patterns. | README.md |
script.py |
Example script that references manifest_url and shows how to download a single Parquet file. |
script.py |
reconstruct_db.py |
Utility that pulls the manifest, filters a date range, and builds a local DuckDB file with the last 30 days of data. | reconstruct_db.py |
manifest.json (remote) |
Index of all dataset files; fetched from the public bucket. | https://db-public.bbltracker.com/manifest.json |
These files together illustrate the complete workflow: fetch the manifest → filter → construct URLs → query with DuckDB. By leveraging manifest.json, clients can efficiently discover and select exactly the data they need without enumerating the entire storage bucket.
Summary
- The
manifest.jsonfile athttps://db-public.bbltracker.com/manifest.jsonserves as a centralized index for the BBL Tracker public dataset. - It contains a
generatedtimestamp and afilesdictionary mapping Parquet filenames to metadata (including row counts). - Programmatic file discovery involves fetching the JSON, filtering keys by date prefixes or metadata values, and prepending the base URL to construct full paths.
- DuckDB can consume a list of remote URLs directly via
read_parquet(), enabling analysis without intermediate downloads. - Reference implementations in
script.pyandreconstruct_db.pydemonstrate production-ready patterns for manifest consumption.
Frequently Asked Questions
What is the URL for the manifest.json file?
The manifest is hosted at https://db-public.bbltracker.com/manifest.json. This endpoint returns a JSON object containing the generated timestamp and the files dictionary that indexes all available Parquet files in the dataset.
How often is the manifest.json updated?
According to the repository structure, the manifest is regenerated periodically to reflect new hourly Parquet files as they are added to the Cloudflare R2 bucket. The generated field in the JSON provides the exact ISO-8601 timestamp of when that specific manifest version was created.
Can I use the manifest to filter files by specific date ranges?
Yes. Because filenames follow the pattern YYYY-MM-DD-HHMM.parquet, you can filter the files dictionary keys using string prefix matching (e.g., name.startswith("2026-02-") for February 2026). This allows precise date-range selection without downloading unnecessary files.
What metadata is available for each file in the manifest?
Each file entry in the files dictionary currently includes a rows field indicating the number of data rows in that Parquet file. This metadata allows clients to skip very small files or prioritize larger datasets based on their specific analysis requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →