# How Poly Data Handles Missing Markets: A Complete Technical Guide

> Learn how Poly Data handles missing markets using a supplemental CSV and Python functions to automatically detect, fetch, and persist market data for comprehensive analysis.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: how-to-guide
- Published: 2026-04-21

---

**Poly Data handles missing markets by maintaining a supplemental `missing_markets.csv` file that gets merged with the main `markets.csv` database, using two core functions in [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) to automatically detect, fetch, and persist markets not captured during bulk downloads.**

The **Poly Data** repository provides a robust mechanism for managing Polymarket prediction market data. A critical challenge in this domain is handling **missing markets**—market records that exist on-chain or in live trading data but were not captured during the initial bulk download process. This technical deep-dive explains exactly how Poly Data detects and handles missing markets through its dual-file architecture and utility functions.

## The Missing Market Problem in Poly Data

When working with decentralized prediction market data, gaps inevitably appear between what was originally scraped and what subsequently appears in live trading feeds. Poly Data solves this through a **two-tier storage system**:

1. **`markets.csv`** — The primary database of markets fetched during bulk updates
2. **`missing_markets.csv`** — A supplemental file for markets discovered later

According to the Poly Data source code, any market not present in `markets.csv` but encountered during live processing gets routed through this missing market pipeline.

## Core Functions for Missing Market Handling

The entire missing market workflow is implemented in **[`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py)** through two primary functions: `get_markets()` and `update_missing_tokens()`.

### Loading and Merging Market Data with `get_markets()`

The **`get_markets()`** function (lines 12–51) creates a unified view of all available markets by combining both CSV sources:

```python
from poly_utils.utils import get_markets

# Load markets.csv + missing_markets.csv (if present)

markets_df = get_markets()          # defaults to the two CSV names

print(f"Total markets available: {len(markets_df)}")

```

**What happens under the hood:**

1. **Detection** (lines 12–38): The function checks for `markets.csv` and optionally `missing_markets.csv`, reading each into a Polars DataFrame

2. **Merge and deduplication** (lines 43–48): DataFrames are concatenated, duplicate market IDs are removed (keeping first occurrence), and results are sorted by `createdAt`

3. **Unified output** (lines 50–51): Returns a single canonical Polars DataFrame for downstream processing

The deduplication logic ensures that if a market exists in both files, the `markets.csv` version takes precedence.

### Populating Missing Markets with `update_missing_tokens()`

When live trading data reveals token IDs without corresponding market records, **`update_missing_tokens()`** (lines 54–112) fetches and persists the missing market data:

```python
from poly_utils.utils import update_missing_tokens

# Suppose we discovered token IDs that are not in markets.csv

missing_token_ids = [
    "0x1234abcd...",   # token for market A

    "0xdeadbeef...",   # token for market B

]

# Fetch their market data and store in missing_markets.csv

update_missing_tokens(missing_token_ids)

```

**The function executes a six-step workflow:**

1. **Read existing missing markets** (lines 80–88): Loads `missing_markets.csv` if it exists, collecting already-processed market IDs to avoid redundant API calls

2. **Filter new tokens** (lines 90–93): Removes IDs already present in the missing markets file

3. **API querying** (lines 95–100): For each new token, queries the Polymarket API to retrieve complete market metadata

4. **Row construction** (lines 102–105): Builds market rows matching the exact column order of `markets.csv`

5. **Persistence** (lines 107–109): Appends new rows to `missing_markets.csv`, writing headers only when creating the file for the first time

6. **Return** (line 112): Returns the count of newly added markets

This incremental approach ensures that `missing_markets.csv` grows organically as new markets are discovered, without ever duplicating entries.

## Complete Workflow Example

The typical end-to-end usage pattern combines bulk updates, missing market detection, and unified access:

```python

# 1️⃣ Refresh the main market dump (runs periodically)

# from update_utils.update_markets import update_markets

# update_markets()

# 2️⃣ Identify token IDs that are missing (custom logic)

# missing_ids = ...

# 3️⃣ Back-fill those markets

update_missing_tokens(missing_ids)

# 4️⃣ Work with the complete market list

all_markets = get_markets()

# ... proceed with analytics, live-trade processing, etc.

```

This pattern ensures that downstream analytics in **[`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py)** always operate against a complete, deduplicated market dataset.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) | Core missing market logic: `get_markets()` and `update_missing_tokens()` |
| [`update_utils/update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_markets.py) | Bulk market fetching into `markets.csv` |
| [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) | Live trade enrichment using missing market utilities |

## Summary

- **Poly Data handles missing markets through a dual-file architecture**: `markets.csv` for primary data and `missing_markets.csv` for back-filled records

- **`get_markets()` in [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py)** automatically merges both sources, deduplicates by market ID, and returns a unified Polars DataFrame

- **`update_missing_tokens()` in [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py)** fetches missing market metadata from the Polymarket API, avoids duplicates through pre-checking, and persistently stores results in `missing_markets.csv`

- The incremental, idempotent design ensures that live trading pipelines always access complete market data without reprocessing or duplication

## Frequently Asked Questions

### What triggers the creation of missing_markets.csv?

The `missing_markets.csv` file is created on demand when `update_missing_tokens()` encounters token IDs that cannot be matched to existing markets. The function writes headers only when creating the file for the first time, then appends subsequent entries. No manual file creation is required.

### How does Poly Data prevent duplicate missing market entries?

Before querying the Polymarket API, `update_missing_tokens()` reads any existing `missing_markets.csv` into memory and extracts the market IDs (lines 80–88). These IDs are excluded from the API query batch. Additionally, `get_markets()` performs a final deduplication when merging files, keeping the first occurrence of any duplicate market ID.

### Can I use get_markets() without a missing_markets.csv file?

Yes. The `get_markets()` function is designed to work with `markets.csv` alone. The `missing_markets.csv` parameter defaults to `None`, and the function only attempts to load the supplemental file when explicitly provided or when the default file exists. This makes the function backward compatible with single-file deployments.

### What API does update_missing_tokens() use to fetch missing market data?

The function queries the Polymarket Gamma API endpoint. For each token ID, it constructs a request to retrieve the associated market metadata including condition ID, market slug, description, and resolution details. The response is parsed and normalized to match the exact column schema used in `markets.csv` before persistence.