Understanding the Parquet File Naming Convention and 6-Hour Partition Structure in BBL Tracker
The BBL Tracker public dataset stores data in Parquet files partitioned into 6-hour UTC intervals, with filenames following the pattern YYYY-MM-DD-HHMM.parquet where HHMM is always 0000, 0600, 1200, or 1800.
The nelsonjchen/bbl-tracker-public-db repository provides a public database of BBL (Broadband Light) tracker data, organized as a time-series dataset stored in Apache Parquet format. Understanding the Parquet file naming convention and 6-hour partition structure is essential for efficiently querying specific time ranges and reconstructing the database locally.
The 6-Hour UTC Partitioning Strategy
The dataset is explicitly partitioned into 6-hour UTC intervals to balance data freshness with file management overhead. According to the repository's README.md, this partitioning strategy keeps the number of files manageable while still providing relatively fresh data updates【source L174】.
Each partition represents a half-day segment:
- 0000: Midnight to 6:00 AM UTC
- 0600: 6:00 AM to noon UTC
- 1200: Noon to 6:00 PM UTC
- 1800: 6:00 PM to midnight UTC
Why 6-Hour Intervals?
The 6-hour window represents a compromise between granularity and practicality. Shorter intervals would generate too many small files, increasing HTTP request overhead when querying remote data. Longer intervals would delay data availability and create unnecessarily large files for incremental updates.
Decoding the Parquet File Naming Convention
The Parquet file naming convention follows a strict temporal pattern documented in README.md【source L176】:
YYYY-MM-DD-HHMM.parquet
Where:
YYYY-MM-DDrepresents the calendar date in ISO 8601 formatHHMMrepresents the starting hour and minute (always0000,0600,1200, or1800)
File Name Examples
| File Name | Time Range Covered |
|---|---|
2026-02-15-1800.parquet |
18:00 UTC – 00:00 UTC on 15 Feb 2026 |
2026-02-16-0000.parquet |
00:00 UTC – 06:00 UTC on 16 Feb 2026 |
2026-02-16-0600.parquet |
06:00 UTC – 12:00 UTC on 16 Feb 2026 |
All files are hosted under the public base URL https://db-public.bbltracker.com/.
How URLs Are Constructed in the Source Code
The repository includes Python scripts that programmatically generate URLs following the 6-hour partition structure. Both script.py and reconstruct_db.py implement similar URL construction logic.
In script.py, the code builds filenames and converts them to full URLs using list comprehensions【source L30】:
# Generate URLs for specific 6-hour blocks
base_url = "https://db-public.bbltracker.com"
day = 16
hour = 0 # Represents 0000 (midnight)
url = f"https://db-public.bbltracker.com/2026-02-{day:02d}-{hour:02d}00.parquet"
The reconstruct_db.py script uses similar logic at line 46 to reconstruct the local database from remote Parquet files【source L46】.
Both scripts eventually feed these URLs into DuckDB queries using the read_parquet function【source L47】:
import duckdb
urls = [
"https://db-public.bbltracker.com/2026-02-16-0000.parquet",
"https://db-public.bbltracker.com/2026-02-16-0600.parquet"
]
query = f"SELECT * FROM read_parquet({urls})"
df = duckdb.query(query).df()
Querying Partitioned Parquet Files with DuckDB
DuckDB can query remote Parquet files directly over HTTP without downloading them first. The Parquet file naming convention allows you to target specific time windows precisely.
Query a Single 6-Hour Partition
SELECT *
FROM read_parquet('https://db-public.bbltracker.com/2026-02-16-0000.parquet');
Query Multiple Consecutive Partitions
To analyze a full 24-hour period, combine all four 6-hour partitions for that day:
import duckdb
date = "2026-02-16"
hours = ["0000", "0600", "1200", "1800"]
urls = [f"'https://db-public.bbltracker.com/{date}-{h}.parquet'" for h in hours]
query = f"SELECT * FROM read_parquet([{', '.join(urls)}])"
df = duckdb.query(query).df()
Programmatically Generate URLs for Date Ranges
import datetime
import duckdb
def generate_url(dt: datetime.datetime) -> str:
"""Generate Parquet URL for a given datetime, rounded to 6-hour partition."""
hour = (dt.hour // 6) * 6
return f"https://db-public.bbltracker.com/{dt:%Y-%m-%d}-{hour:02d}00.parquet"
# Generate URLs for a 3-day range
start = datetime.datetime(2026, 2, 15, 0, 0)
end = datetime.datetime(2026, 2, 17, 23, 59)
urls = []
current = start
while current <= end:
urls.append(generate_url(current))
current += datetime.timedelta(hours=6)
# Query with DuckDB
df = duckdb.query(f"SELECT * FROM read_parquet({urls})").df()
Summary
- The Parquet file naming convention follows
YYYY-MM-DD-HHMM.parquet, whereHHMMindicates the start of a 6-hour UTC interval. - Files are partitioned into four 6-hour blocks per day:
0000,0600,1200, and1800UTC. - The
script.pyandreconstruct_db.pyfiles in thenelsonjchen/bbl-tracker-public-dbrepository implement URL construction logic that follows this convention. - DuckDB can query these remote Parquet files directly using
read_parquet()with HTTP URLs, enabling efficient time-range analysis without local downloads.
Frequently Asked Questions
What does the HHMM component represent in the Parquet file naming convention?
The HHMM component represents the starting hour and minute of the 6-hour UTC time block contained in the file. Valid values are always 0000, 0600, 1200, or 1800, indicating midnight, 6 AM, noon, and 6 PM UTC respectively. This standardization allows the script.py utility to programmatically generate correct URLs for any given time range.
Why are the files partitioned into 6-hour intervals rather than hourly or daily?
According to the repository documentation, the 6-hour partition structure strikes a balance between data freshness and file management overhead. Hourly partitions would create too many small files, increasing HTTP request overhead when querying remote data. Daily partitions would delay data availability and create unnecessarily large files for incremental updates. The four daily partitions (0000, 0600, 1200, 1800) provide frequent enough updates while keeping the file count manageable.
How can I query data across multiple 6-hour partitions using DuckDB?
DuckDB's read_parquet function accepts an array of URLs, allowing you to query multiple 6-hour partitions simultaneously. In Python, construct a list of URLs following the YYYY-MM-DD-HHMM.parquet pattern for each 6-hour block you need, then pass this list to the SQL query. For example: SELECT * FROM read_parquet(['url1', 'url2', 'url3']). The reconstruct_db.py script demonstrates this pattern by automatically discovering and querying all recent partitions.
Where are the Parquet files hosted and how are the URLs constructed?
All Parquet files are hosted under the public base URL https://db-public.bbltracker.com/. The complete URL is constructed by appending the filename in the format YYYY-MM-DD-HHMM.parquet. Both script.py and reconstruct_db.py use Python f-strings to generate these URLs programmatically, rounding the desired timestamp down to the nearest 6-hour interval to ensure valid filenames.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →