Understanding Data Ingestion Start Times and Regional Availability in BBL Tracker

Data ingestion for the BBL Tracker public dataset began on January 17, 2026, for the United States, with the European Union, United Kingdom, Australia, and Canada added on January 25, 2026, and global markets following on February 1, 2026.

The BBL Tracker is an open-source public database that streams hourly-to-30-minute snapshots of Bambu Lab store inventory into Parquet files. Understanding data ingestion start times and regional availability is critical for analysts who need to align historical stock queries with the actual coverage window for each market.

Data Ingestion Start Times by Region

The collection pipeline did not launch simultaneously worldwide. Instead, the crawler rolled out in phases, with each phase adding new regional shards to the manifest.json index.

United States (Initial Launch)

The dataset first appeared on January 17, 2026, containing exclusively United States inventory data. These earliest Parquet shards (located in paths like 2026-01-17-0000.parquet) contain rows only where region = 'us'. According to the repository's README.md (lines 64-73), this date marks the initial launch of the public streaming dataset.

European Union, United Kingdom, Australia, and Canada

On January 25, 2026, the ingestion pipeline expanded to include the European Union (eu), United Kingdom (uk), Australia (au), and Canada (ca). Consumers can identify the exact crossover point by scanning the manifest.json for the first shards dated January 25, 2026, and filtering for rows where the region column matches these new codes.

Global Markets and Asia-Pacific Expansion

The final major expansion occurred on February 1, 2026, adding global coverage for the rest of Asia (excluding Japan and Korea). This phase completed the regional rollout, ensuring that the dataset now captures inventory snapshots across all active Bambu Lab store regions.

Regional Availability in the Dataset Schema

Each Parquet file follows a consistent schema that encodes regional provenance explicitly, allowing for precise filtering without parsing file paths.

The Region Column and Identifiers

Every row contains a region column (documented in README.md lines 38-46) that stores lowercase ISO-style region codes: us, eu, uk, au, ca, and others. This column is the primary mechanism for isolating market-specific data. For example, to analyze only German inventory (served by the EU store), you would filter WHERE region = 'eu' and further refine by product availability.

Temporal Granularity and Sampling Rates

The dataset's sampling frequency varies by date, affecting the resolution of regional availability windows:

  • January 17 – February 14, 2026: Approximately 60-minute intervals
  • February 15, 2026 onwards: Approximately 30-minute intervals, aligned to the hour plus 5 minutes (e.g., 00:05, 00:35)

As noted in README.md (lines 92-96), this increased granularity improves the precision of stock-out detection for high-velocity items in each region.

Querying Regional Data with Python and DuckDB

The repository provides reference implementations in script.py and reconstruct_db.py that demonstrate how to filter by region using DuckDB's Parquet reader.

To query recent United States availability (mirroring the logic in script.py lines 38-49):

import duckdb
import urllib.request
import json

# Load the manifest index

manifest = json.loads(urllib.request.urlopen(
    "https://db-public.bbltracker.com/manifest.json").read())

# Select the 4 most recent shards (approximately last 24 hours)

urls = [f"https://db-public.bbltracker.com/{f}"
        for f in sorted(manifest['files'].keys())[-4:]]

query = f"""
SELECT
    product_name || ' - ' || variant_name AS sku,
    ROUND(100.0 * SUM(CASE WHEN stock > 0 THEN 1 END) /
          COUNT(*), 1) AS availability_pct,
    MAX(stock) AS max_stock_seen
FROM read_parquet({urls})
WHERE region = 'us'
GROUP BY sku
ORDER BY availability_pct ASC
LIMIT 10;
"""
df = duckdb.query(query).df()
print(df)

To build a persistent database for cross-region analysis (as implemented in reconstruct_db.py lines 56-65):

import duckdb
import json
import urllib.request
from datetime import datetime, timedelta, timezone

# Fetch manifest

manifest = json.loads(urllib.request.urlopen(
    "https://db-public.bbltracker.com/manifest.json").read())

# Filter to last 30 days

cutoff = (datetime.now(timezone.utc) - timedelta(days=30)).strftime("%Y-%m-%d")
files = [f for f in sorted(manifest['files']) if f >= cutoff]
urls = [f"https://db-public.bbltracker.com/{f}" for f in files]

# Build local DuckDB

con = duckdb.connect("bambu_stock.duckdb")
con.execute(f"CREATE TABLE stock_history AS SELECT * FROM read_parquet({urls});")

# Query EU-specific availability

eu_df = con.execute("""
SELECT timestamp, stock
FROM stock_history
WHERE region = 'eu'
  AND product_name = 'PLA Matte'
  AND variant_name = 'White';
""").df()
print(eu_df.head())

Summary

  • Data ingestion start times were staggered: US began on January 17, 2026; EU/UK/AU/CA on January 25, 2026; and global markets on February 1, 2026.
  • Regional availability is encoded in the region column of every Parquet file, using lowercase codes like us, eu, and au.
  • Temporal resolution improved from 60-minute to 30-minute intervals after February 15, 2026, enabling more precise stock tracking per region.
  • The manifest.json index and reference scripts (script.py, reconstruct_db.py) provide ready-made patterns for filtering by region using DuckDB.

Frequently Asked Questions

When did data ingestion begin for each region in the BBL Tracker dataset?

Data ingestion began on January 17, 2026, for the United States. The European Union, United Kingdom, Australia, and Canada were added on January 25, 2026. Global coverage for the rest of Asia (excluding Japan and Korea) began on February 1, 2026, completing the regional rollout.

How is regional availability represented in the Parquet files?

Each row in the Parquet schema contains a region column that stores a lowercase region code (e.g., us, eu, uk, au, ca). This column allows analysts to filter queries to specific markets without parsing filenames. The schema also includes timestamp, product_name, variant_name, stock, and is_flash_sale columns for complete inventory context.

What is the temporal granularity of the regional data?

The dataset uses two sampling rates. From January 17 to February 14, 2026, snapshots were taken approximately every 60 minutes. From February 15, 2026 onwards, the frequency increased to approximately 30 minutes, aligned to 5 minutes past the hour (e.g., 00:05, 00:35). This improved granularity allows for more precise detection of stock-out events in each region.

How can I query data for a specific region using DuckDB?

You can filter by the region column in SQL queries against the Parquet files. For quick analysis, use the pattern shown in script.py to read recent shards and filter with WHERE region = 'us'. For persistent analysis, use reconstruct_db.py to build a local DuckDB database from the last 30 days of shards, then run queries like SELECT * FROM stock_history WHERE region = 'eu' AND product_name = 'PLA Matte'.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →