# How to Debug Network Errors and Handle Missing Parquet Files in BBL Tracker

> Debug BBL Tracker network errors and missing parquet files. Implement retry logic, validate URLs, and filter missing files to handle transient failures with diagnostic logging.

- Repository: [Nelson Chen/bbl-tracker-public-db](https://github.com/nelsonjchen/bbl-tracker-public-db)
- Tags: how-to-guide
- Published: 2026-03-08

---

**Implement retry logic with exponential back-off for manifest downloads, validate shard URLs with HEAD requests before passing them to DuckDB, and filter the URL list to skip missing files while logging diagnostics for transient Cloudflare R2 failures.**

The BBL Tracker public database distributes hourly Parquet shards via Cloudflare R2, consumed by Python scripts that rely on a central manifest. When debugging network errors and handling missing parquet files in the `nelsonjchen/bbl-tracker-public-db` repository, you must address both manifest retrieval failures and individual shard unavailability without crashing the ingestion pipeline.

## Why Network Errors Occur in the BBL Tracker Pipeline

The dataset uses a **manifest-first architecture** where [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) and [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py) fetch `https://db-public.bbltracker.com/manifest.json` to discover available shards. Network errors manifest when Cloudflare R2 returns non-200 status codes, rate limits requests, or when individual Parquet files are deleted or renamed after manifest publication. The current implementation in [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) (lines 13-20) uses bare `urllib.request.urlopen` calls that abort on any transient failure, while the DuckDB query construction at line 47 in [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) and lines 60-65 in [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py) assumes all listed files exist.

## Validating the Manifest Download with Retry Logic

The manifest fetch in [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) currently lacks resilience against transient network glitches. Wrap the `urllib.request.urlopen` call in an exponential back-off loop to handle temporary connectivity issues or rate limiting from Cloudflare R2.

```python
import time, urllib.request, json
from urllib.error import URLError, HTTPError

MANIFEST_URL = "https://db-public.bbltracker.com/manifest.json"

def fetch_manifest(retries: int = 5, backoff: float = 1.0) -> dict:
    for attempt in range(1, retries + 1):
        try:
            with urllib.request.urlopen(MANIFEST_URL) as resp:
                if resp.status != 200:
                    raise HTTPError(MANIFEST_URL, resp.status, "Non‑200", resp.headers, None)
                return json.loads(resp.read())
        except (URLError, HTTPError) as exc:
            print(f"[Attempt {attempt}] Failed to fetch manifest: {exc}")
            if attempt == retries:
                raise
            time.sleep(backoff * attempt)      # exponential back‑off

```

This pattern replaces the direct fetch at [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) lines 13-20, providing clear error messages when the manifest server is unreachable after multiple attempts.

## Cross-Checking Shard URLs Before DuckDB Ingestion

DuckDB's `read_parquet` function raises exceptions if any URL in the list returns 404 or 403. Before constructing the SQL query at line 47 in [`script.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/script.py) or lines 60-65 in [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py), send HTTP HEAD requests to validate each shard's existence without downloading the full file.

```python
import urllib.request
from urllib.error import URLError, HTTPError

def url_exists(url: str) -> bool:
    try:
        req = urllib.request.Request(url, method="HEAD")
        with urllib.request.urlopen(req) as resp:
            return resp.status == 200
    except (URLError, HTTPError):
        return False

# After obtaining `all_files` from the manifest …

valid_urls = []
for fname in recent_files:
    u = f"https://db-public.bbltracker.com/{fname}"
    if url_exists(u):
        valid_urls.append(u)
    else:
        print(f"[WARN] Shard missing: {fname}")

# Use `valid_urls` in DuckDB (see reconstruct_db.py lines 60‑65)

query = f"""
    CREATE TABLE stock_history AS
    SELECT * FROM read_parquet({valid_urls});
"""

```

This validation step catches deleted or renamed files early, preventing DuckDB from raising "File not found" errors during query execution.

## Gracefully Skipping Missing Parquet Shards

Instead of passing the complete manifest list to DuckDB, build a filtered `valid_urls` list containing only accessible shards. This approach allows the pipeline to continue with available data even when specific hourly shards are missing due to upstream processing delays or deletions. Include detailed logging that captures the shard name, HTTP status code, and timestamp to diagnose patterns in Cloudflare rate-limiting or connectivity issues.

## Implementing Fallback Strategies

When all recent shards fail validation, implement a fallback to previous day's data or exit with diagnostic information pointing users to the public bucket URL.

```python
from datetime import datetime, timedelta

if not valid_urls:
    # Try previous 6‑hour interval

    prev_day = (datetime.utcnow() - timedelta(days=1)).strftime("%Y-%m-%d-1800.parquet")
    fallback = f"https://db-public.bbltracker.com/{prev_day}"
    if url_exists(fallback):
        valid_urls = [fallback]
        print(f"[INFO] Falling back to {prev_day}")
    else:
        raise RuntimeError("No accessible Parquet shards found. Check https://db-public.bbltracker.com")

```

This ensures the workflow remains resilient when the most recent 6-hour Parquet files (named deterministically as `YYYY-MM-DD-HHMM.parquet` per the README) are unavailable.

## Summary

- Wrap manifest downloads in exponential back-off retry loops to handle transient R2 failures
- Validate individual Parquet shard URLs with HTTP HEAD requests before DuckDB ingestion
- Filter the URL list passed to `read_parquet` to exclude missing files and prevent query failures
- Log shard names and HTTP status codes to diagnose Cloudflare rate-limiting or connectivity issues
- Implement date-based fallbacks when recent shards are completely unavailable

## Frequently Asked Questions

### What causes "File not found" errors in DuckDB when querying BBL Tracker data?

DuckDB raises exceptions when `read_parquet` receives a URL list containing missing files. The manifest may reference Parquet shards that were deleted or renamed after publication, causing the query at lines 60-65 in [`reconstruct_db.py`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/reconstruct_db.py) to fail when it attempts to read non-existent objects from `https://db-public.bbltracker.com`.

### How should I handle rate limiting from Cloudflare R2 when downloading shards?

Implement exponential back-off in your `urllib.request` calls, waiting progressively longer between attempts (1 second, 2 seconds, 3 seconds) to allow transient rate limits to reset before aborting. This pattern handles temporary HTTP 429 responses without crashing the ingestion script.

### Can I use the BBL Tracker scripts if some hourly Parquet files are permanently missing?

Yes. Filter the URL list to include only validated shards using the `url_exists` HEAD request pattern. DuckDB will successfully query the remaining available data, though your resulting dataset may have gaps in the time series for the missing hours.

### Where does the BBL Tracker store its public database files?

All Parquet shards and the [`manifest.json`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/manifest.json) live at `https://db-public.bbltracker.com`, hosted on Cloudflare R2. The files follow deterministic naming conventions (`YYYY-MM-DD-HHMM.parquet`) representing 6-hour intervals, as documented in [`README.md`](https://github.com/nelsonjchen/bbl-tracker-public-db/blob/main/README.md) § File Structure.