Understanding Sampling Rate Changes in Bambu Lab Filament Tracker Data

The Bambu Lab Store Filament Tracker switched from approximately 60‑minute snapshots to 30‑minute snapshots on February 14 2026, aligning new samples to 5 minutes past each hour to improve temporal resolution without changing the underlying Parquet file structure.

The nelsonjchen/bbl-tracker-public-db repository maintains a public dataset of Bambu Lab filament stock levels. Understanding sampling rate changes is essential for analysts working with historical availability data, as the shift from hourly to half‑hourly collection affects how you detect brief stock events like flash‑sale restocks.

Sampling Rate Timeline and Alignment

The dataset contains two distinct sampling regimes separated by a transition date in February 2026.

Initial 60‑Minute Collection Period

From January 17 2026 through February 14 2026, the ingestion pipeline captured stock snapshots at approximately 60‑minute intervals. These samples carried no specific alignment to clock hours, resulting in irregular offsets depending on when the crawler initiated.

Transition to 30‑Minute Intervals

Beginning February 14 2026, the collection frequency doubled to approximately 30 minutes. The new schedule aligns samples to 5 minutes past each hour and half‑hour (e.g., 00:05, 00:35, 01:05). This alignment simplifies partitioning logic: each 6‑hour Parquet shard contains exactly 12 snapshots (or 6 for partial windows), maintaining consistent file sizes while doubling temporal resolution.

Why the Sampling Rate Changed

The shift from 60‑minute to 30‑minute sampling reflects a deliberate evolution in data collection strategy.

Initially, the one‑hour cadence minimized bandwidth consumption while the team evaluated data source stability and storage costs. Once the pipeline proved reliable, the frequency increased to capture finer‑grained insight into short‑lived stock events, such as filament variants that sell out within 45 minutes of restocking.

The hour + 5 minute alignment specifically supports the existing partitioning scheme. By fixing samples to predictable offsets, the reconstruct_db.py script and other downstream tools can deterministically map timestamps to Parquet filenames without parsing file contents.

Impact on Data Analysis and Querying

Mixed sampling rates introduce considerations for temporal analysis across the February 14 2026 boundary.

Temporal Resolution Differences

Queries spanning the early period will encounter fewer data points per hour, potentially smoothing over brief stock spikes that lasted less than 60 minutes. Post‑February 14 queries capture these transient events with twice the fidelity, revealing restock patterns invisible in the earlier data.

File Structure Consistency

Despite the sampling change, the storage format remains unchanged. Data continues to partition into 6‑hour Parquet files named YYYY‑MM‑DD‑HHMM.parquet. The internal row count per file roughly doubles after the transition, but the schema and UTC timestamp handling remain consistent, ensuring backward compatibility with existing DuckDB or Pandas workflows.

Working with Mixed Sampling Rates in Code

When analyzing data across both regimes, normalize the temporal granularity to ensure consistent visualizations.

Loading Recent Data Across Both Regimes

This Python example pulls the last 24 hours of data (which may span both sampling rates) and resamples to uniform 30‑minute bins:

import duckdb
import pandas as pd
import matplotlib.pyplot as plt

# Build the list of recent 6‑hour files (last 4 shards)

manifest_url = "https://db-public.bbltracker.com/manifest.json"
manifest = duckdb.read_json(manifest_url).fetchall()[0][0]
files = sorted(manifest['files'].keys())[-4:]
urls = [f"https://db-public.bbltracker.com/{f}" for f in files]

# Load all rows

df = duckdb.read_parquet(urls).df()

# Convert timestamp string to datetime (UTC)

df["ts"] = pd.to_datetime(df["timestamp"], utc=True)

# Resample to 30‑minute bins, counting in‑stock snapshots per variant

variant = "Black PETG"
variant_df = df[df["variant_name"] == variant]

# Count snapshots where stock > 0 per 30‑min bucket

availability = (
    variant_df.set_index("ts")
    .groupby("variant_name")
    .apply(lambda x: (x["stock"] > 0).resample("30T").sum())
    .reset_index(name="in_stock")
)

plt.figure(figsize=(10, 4))
plt.plot(availability["ts"], availability["in_stock"], drawstyle="steps-post")
plt.title(f"In‑stock snapshots for {variant} (30‑min bins)")
plt.xlabel("UTC time")
plt.ylabel("Snapshots with stock > 0")
plt.grid(True)
plt.show()

Normalizing Early‑Period Data

To compare pre‑February 14 data with later 30‑minute samples, forward‑fill the hourly data to match the finer granularity:


# Resample early‑period (1‑hour) data to 30‑minute bins by forward‑filling

early = df[df["timestamp"] < "2026-02-14T00:00:00Z"]
early["ts"] = pd.to_datetime(early["timestamp"], utc=True)
early_30 = (
    early.set_index("ts")
    .groupby("variant_name")["stock"]
    .resample("30T")
    .ffill()
    .reset_index()
)

# Combine with later (already 30‑min) data

later = df[df["timestamp"] >= "2026-02-14T00:00:00Z"]
combined = pd.concat([early_30, later], ignore_index=True)

Quick CLI Verification

Verify sampling intervals directly via DuckDB without downloading files locally:

duckdb -c "SELECT * FROM read_parquet('https://db-public.bbltracker.com/2026-02-16-0600.parquet') LIMIT 5;"

Summary

  • The dataset transitioned from 60‑minute sampling (Jan 17 – Feb 14 2026) to 30‑minute sampling (Feb 14 2026 onward) to improve detection of brief stock events.
  • New samples align to 5 minutes past the hour and half‑hour, maintaining compatibility with the existing 6‑hour Parquet partitioning scheme.
  • File structure remains unchanged: data resides in YYYY‑MM‑DD‑HHMM.parquet shards with UTC timestamps, though post‑transition files contain roughly twice as many rows.
  • Analysts should normalize mixed‑granularity data using Pandas resample() or DuckDB temporal functions to ensure consistent visualizations across the February 14 boundary.

Frequently Asked Questions

How do I detect which sampling rate a specific Parquet file uses?

Check the timestamp range within the file. Files covering periods before February 14 2026 contain roughly 6–12 rows per variant (hourly sampling), while files after that date contain 12–24 rows per variant (30‑minute sampling). You can verify this programmatically by counting distinct timestamps per variant using DuckDB: SELECT variant_name, COUNT(DISTINCT timestamp) FROM read_parquet('file.parquet') GROUP BY variant_name.

Will the sampling rate change again in the future?

According to the repository documentation in README.md, the current 30‑minute cadence is considered stable for production use. However, the ingestion pipeline supports arbitrary intervals, so future adjustments are possible if storage costs or API rate limits change. Monitor the repository’s commit history or the #sampling-rate section in README.md for announcements.

Does the 5‑minute offset affect time‑zone conversions?

No. All timestamps are stored in UTC, so the 5‑minute alignment (e.g., 00:05, 00:35) is consistent regardless of your local time zone. When converting to local time for display purposes, apply your time‑zone offset after parsing the UTC string. The offset exists solely to ensure Parquet files contain complete 30‑minute intervals without straddling hour boundaries.

How do I resample the early hourly data to match the 30‑minute granularity?

Use Pandas resample() with forward‑fill (ffill) or interpolation. First filter for timestamps before February 14 2026, convert to datetime index, then apply .resample('30T').ffill(). This creates synthetic 30‑minute rows that carry forward the last known stock value, allowing consistent visualization alongside native 30‑minute samples from the later period. See the code example in the "Normalizing Early‑Period Data" section above for the exact implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →