How Does Poly Data Fetch Market Data from Polymarket?
Poly Data fetches market data from Polymarket's public Gamma API using two complementary Python routines—update_markets() for full chronological crawls and update_missing_tokens() for targeted retrieval by token ID—storing normalized results in CSV format for downstream Polars analysis.
Poly Data is an open-source toolkit (warproxxx/poly_data) designed to extract prediction market data from Polymarket. The system implements a robust pagination strategy that handles rate limits, parses JSON-encoded metadata, and maintains a stable local cache of market information in markets.csv.
Full Market Crawl via update_markets
The primary ingestion routine is update_markets(), located in update_utils/update_markets.py. This function performs a chronological crawl of Polymarket’s Gamma API endpoint (https://gamma-api.polymarket.com/markets), downloading every market in creation order.
Deterministic Pagination Strategy
The crawler implements resume capability by counting existing records before issuing requests. The helper count_csv_lines() determines the starting offset, allowing the script to append new data without duplication.
current_offset = count_csv_lines(csv_filename)
mode = 'a' if file_exists else 'w'
API requests use four query parameters to ensure deterministic ordering:
params = {
'order': 'createdAt',
'ascending': 'true',
'limit': batch_size,
'offset': current_offset
}
response = requests.get(base_url, params=params, timeout=30)
The script advances the offset by the actual number of markets written (not the requested batch size), stopping when a returned batch is smaller than batch_size—signaling exhaustion of the API dataset.
Robust Retry Handling
The implementation includes specific backoff strategies for HTTP error codes:
- HTTP 500: Wait 5 seconds and retry
- HTTP 429: Wait 10 seconds and retry
- Other non-200: Log error and retry after 3 seconds
This ensures reliable completion across large historical syncs without manual intervention.
Parsing and CSV Storage
Market objects from the Gamma API contain JSON-encoded strings for nested fields. The parser explicitly decodes outcomes and clobTokenIds using json.loads() before extracting individual values:
outcomes = json.loads(market.get('outcomes', '[]'))
clob_tokens = json.loads(market.get('clobTokenIds', '[]'))
Each row is written to the CSV with a fixed schema: createdAt, id, question, outcome1, outcome2, token_id_1, token_id_2, negativeRisk, ticker, closedTime. This standardized format allows downstream modules to consume data predictably.
Targeted Retrieval via update_missing_tokens
When specific markets are missing from the primary crawl (e.g., markets added after initial sync or skipped records), poly_utils/utils.py provides update_missing_tokens() to fill gaps individually.
Token-Based Filtering
Unlike the full crawl which uses offset pagination, this routine queries the same Gamma API endpoint using the clob_token_ids parameter to retrieve specific markets by their CLOB token ID:
response = requests.get(
'https://gamma-api.polymarket.com/markets',
params={'clob_token_ids': token_id},
timeout=30
)
This approach is essential for fetching markets that were not captured in the chronological crawl due to API hiccups or late additions.
Deduplication Logic
Before writing new records, the function reads existing IDs from the target CSV into a processed_market_ids set to prevent duplication:
with open(csv_filename, 'r', encoding='utf-8') as f:
reader = csv.DictReader(f)
for row in reader:
processed_market_ids.add(row['id'])
Records are only appended if their market ID is not already present, ensuring idempotent operations when run multiple times.
Data Consumption with Polars
Once markets.csv (and optionally missing_markets.csv) is populated, poly_utils/utils.py exposes get_markets() to load the data into a Polars DataFrame:
from poly_utils.utils import get_markets
df = get_markets() # Returns Polars DataFrame sorted by createdAt
This function reads both CSV files, deduplicates on the market id column, and returns a unified combined_df. The DataFrame is used throughout the repository—for example, process_live.py converts this wide format into a long format for time-series analysis.
Practical Implementation Examples
Running a Full Market Update
To fetch all markets in batches of 500 (default) and store them in markets.csv:
from update_utils.update_markets import update_markets
update_markets()
Alternatively, run the convenience wrapper:
from update_all import main
main() # Executes update_markets() and handles CSV initialization
Filling Specific Market Gaps
To retrieve markets by their CLOB token IDs:
from poly_utils.utils import update_missing_tokens
missing_ids = [
"0x1234567890abcdef...",
"0xfedcba0987654321..."
]
update_missing_tokens(missing_ids)
Loading Data for Analysis
To access the consolidated market table in your analysis pipeline:
from poly_utils.utils import get_markets
markets_df = get_markets()
print(markets_df.head())
Summary
- Poly Data fetches market data from the Polymarket Gamma API endpoint (
https://gamma-api.polymarket.com/markets). update_markets()inupdate_utils/update_markets.pyperforms chronological pagination usingcreatedAtordering, with robust retry logic for HTTP 500/429 errors.update_missing_tokens()inpoly_utils/utils.pyenables targeted fetching byclob_token_idsand implements CSV-level deduplication.- Both routines output to a standardized CSV schema consumed by
get_markets(), which returns a Polars DataFrame for downstream analysis. - The system supports resume capability via line-counting offsets, making it suitable for incremental updates of large datasets.
Frequently Asked Questions
What API endpoint does Poly Data use to fetch market data?
Poly Data queries Polymarket's public Gamma API at https://gamma-api.polymarket.com/markets. The full crawl uses query parameters for pagination (order=createdAt, ascending=true, limit, offset), while targeted fetches use the clob_token_ids filter parameter.
How does Poly Data handle API rate limits and errors?
The codebase implements specific retry logic with exponential backoff. When the API returns HTTP 429 (rate limit), the script waits 10 seconds before retrying. For HTTP 500 errors, it waits 5 seconds. All other non-200 responses trigger a 3-second wait and retry cycle, ensuring robust completion across unstable network conditions.
What is the difference between update_markets and update_missing_tokens?
update_markets() performs a full chronological crawl of all markets using offset pagination, storing every record in markets.csv. update_missing_tokens() is designed for targeted retrieval—accepting a list of specific CLOB token IDs, querying the API for just those markets, and appending them to the CSV only if not already present.
How is the fetched market data stored and accessed?
Data is persisted to CSV files (primarily markets.csv) with a fixed column schema including createdAt, id, question, outcome strings, token IDs, and metadata. The get_markets() function in poly_utils/utils.py reads these CSVs into a Polars DataFrame, deduplicates by market ID, and returns the data sorted by creation timestamp for analysis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →