How to Debug Network Errors and Handle Missing Parquet Files in BBL Tracker
Implement retry logic with exponential back-off for manifest downloads, validate shard URLs with HEAD requests before passing them to DuckDB, and filter the URL list to skip missing files while logging diagnostics for transient Cloudflare R2 failures.
The BBL Tracker public database distributes hourly Parquet shards via Cloudflare R2, consumed by Python scripts that rely on a central manifest. When debugging network errors and handling missing parquet files in the nelsonjchen/bbl-tracker-public-db repository, you must address both manifest retrieval failures and individual shard unavailability without crashing the ingestion pipeline.
Why Network Errors Occur in the BBL Tracker Pipeline
The dataset uses a manifest-first architecture where script.py and reconstruct_db.py fetch https://db-public.bbltracker.com/manifest.json to discover available shards. Network errors manifest when Cloudflare R2 returns non-200 status codes, rate limits requests, or when individual Parquet files are deleted or renamed after manifest publication. The current implementation in script.py (lines 13-20) uses bare urllib.request.urlopen calls that abort on any transient failure, while the DuckDB query construction at line 47 in script.py and lines 60-65 in reconstruct_db.py assumes all listed files exist.
Validating the Manifest Download with Retry Logic
The manifest fetch in script.py currently lacks resilience against transient network glitches. Wrap the urllib.request.urlopen call in an exponential back-off loop to handle temporary connectivity issues or rate limiting from Cloudflare R2.
import time, urllib.request, json
from urllib.error import URLError, HTTPError
MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"
def fetch_manifest(retries: int = 5, backoff: float = 1.0) -> dict:
for attempt in range(1, retries + 1):
try:
with urllib.request.urlopen(MANIFEST_URL) as resp:
if resp.status != 200:
raise HTTPError(MANIFEST_URL, resp.status, "Non‑200", resp.headers, None)
return json.loads(resp.read())
except (URLError, HTTPError) as exc:
print(f"[Attempt {attempt}] Failed to fetch manifest: {exc}")
if attempt == retries:
raise
time.sleep(backoff * attempt) # exponential back‑off
This pattern replaces the direct fetch at script.py lines 13-20, providing clear error messages when the manifest server is unreachable after multiple attempts.
Cross-Checking Shard URLs Before DuckDB Ingestion
DuckDB's read_parquet function raises exceptions if any URL in the list returns 404 or 403. Before constructing the SQL query at line 47 in script.py or lines 60-65 in reconstruct_db.py, send HTTP HEAD requests to validate each shard's existence without downloading the full file.
import urllib.request
from urllib.error import URLError, HTTPError
def url_exists(url: str) -> bool:
try:
req = urllib.request.Request(url, method="HEAD")
with urllib.request.urlopen(req) as resp:
return resp.status == 200
except (URLError, HTTPError):
return False
# After obtaining `all_files` from the manifest …
valid_urls = []
for fname in recent_files:
u = f"https://db-public.bbltracker.com/{fname}"
if url_exists(u):
valid_urls.append(u)
else:
print(f"[WARN] Shard missing: {fname}")
# Use `valid_urls` in DuckDB (see reconstruct_db.py lines 60‑65)
query = f"""
CREATE TABLE stock_history AS
SELECT * FROM read_parquet({valid_urls});
"""
This validation step catches deleted or renamed files early, preventing DuckDB from raising "File not found" errors during query execution.
Gracefully Skipping Missing Parquet Shards
Instead of passing the complete manifest list to DuckDB, build a filtered valid_urls list containing only accessible shards. This approach allows the pipeline to continue with available data even when specific hourly shards are missing due to upstream processing delays or deletions. Include detailed logging that captures the shard name, HTTP status code, and timestamp to diagnose patterns in Cloudflare rate-limiting or connectivity issues.
Implementing Fallback Strategies
When all recent shards fail validation, implement a fallback to previous day's data or exit with diagnostic information pointing users to the public bucket URL.
from datetime import datetime, timedelta
if not valid_urls:
# Try previous 6‑hour interval
prev_day = (datetime.utcnow() - timedelta(days=1)).strftime("%Y-%m-%d-1800.parquet")
fallback = f"https://db-public.bbltracker.com/{prev_day}"
if url_exists(fallback):
valid_urls = [fallback]
print(f"[INFO] Falling back to {prev_day}")
else:
raise RuntimeError("No accessible Parquet shards found. Check https://db-public.bbltracker.com")
This ensures the workflow remains resilient when the most recent 6-hour Parquet files (named deterministically as YYYY-MM-DD-HHMM.parquet per the README) are unavailable.
Summary
- Wrap manifest downloads in exponential back-off retry loops to handle transient R2 failures
- Validate individual Parquet shard URLs with HTTP HEAD requests before DuckDB ingestion
- Filter the URL list passed to
read_parquetto exclude missing files and prevent query failures - Log shard names and HTTP status codes to diagnose Cloudflare rate-limiting or connectivity issues
- Implement date-based fallbacks when recent shards are completely unavailable
Frequently Asked Questions
What causes "File not found" errors in DuckDB when querying BBL Tracker data?
DuckDB raises exceptions when read_parquet receives a URL list containing missing files. The manifest may reference Parquet shards that were deleted or renamed after publication, causing the query at lines 60-65 in reconstruct_db.py to fail when it attempts to read non-existent objects from https://db-public.bbltracker.com.
How should I handle rate limiting from Cloudflare R2 when downloading shards?
Implement exponential back-off in your urllib.request calls, waiting progressively longer between attempts (1 second, 2 seconds, 3 seconds) to allow transient rate limits to reset before aborting. This pattern handles temporary HTTP 429 responses without crashing the ingestion script.
Can I use the BBL Tracker scripts if some hourly Parquet files are permanently missing?
Yes. Filter the URL list to include only validated shards using the url_exists HEAD request pattern. DuckDB will successfully query the remaining available data, though your resulting dataset may have gaps in the time series for the missing hours.
Where does the BBL Tracker store its public database files?
All Parquet shards and the manifest.json live at https://db-public.bbltracker.com, hosted on Cloudflare R2. The files follow deterministic naming conventions (YYYY-MM-DD-HHMM.parquet) representing 6-hour intervals, as documented in README.md § File Structure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →