Goldsky Subgraph Integration in Poly Data: Resumable GraphQL Scraping Explained
The Goldsky subgraph integration in Poly Data is a resumable GraphQL scraper that incrementally fetches orderFilled events from the Polymarket orderbook, paginates through them using timestamp and ID-based cursors, and persists the data to CSV while maintaining state in JSON to prevent duplicates and enable crash recovery.
Poly Data relies on this integration to capture every trade executed on the Polymarket orderbook. According to the warproxxx/poly_data source code, the implementation operates as a fault-tolerant pipeline that queries the Goldsky hosted subgraph, handles high-throughput pagination edge cases, and stores raw event data locally for downstream analytics.
How the Goldsky Subgraph Integration Works
The integration functions as a resumable GraphQL scraper that performs five core operations:
- Queries the Goldsky hosted subgraph (
orderbook-subgraph) via a GraphQL endpoint to retrieveorderFilledevents. - Pages through events using a combination of timestamp-based and "sticky" ID-based pagination to ensure no events are missed when many records share the same timestamp.
- Writes raw events to
goldsky/orderFilled.csvwith a fixed column order includingtimestamp,maker,makerAssetId,taker,takerAmountFilled, andtransactionHash. - Persists cursor state in
goldsky/cursor_state.jsonusing fieldslast_timestamp,last_id, andsticky_timestampso that subsequent runs resume exactly where the previous execution stopped. - Deduplicates and cleans up by removing events with duplicate
idvalues before appending to the CSV, and deletes the cursor file upon successful completion.
The "Sticky" Pagination Logic
A critical feature of the implementation is the sticky timestamp logic (lines 81-107 in update_utils/update_goldsky.py and lines 71-104 in parallel_sync.py). When a GraphQL batch returns the maximum first size and all events share an identical timestamp, the scraper remains on that timestamp and paginates by id until the timestamp advances. This guarantees complete, gap-free coverage even during high-throughput periods where thousands of orders fill simultaneously.
Both implementations share the same GraphQL query shape:
orderFilledEvents(
orderBy: timestamp,
orderDirection: asc,
first: <batch_size>,
where: { <timestamp or sticky clause> }
) {
id timestamp maker makerAmountFilled makerAssetId
taker takerAmountFilled takerAssetId transactionHash
}
Implementation Variants
The repository provides two concrete implementations of the Goldsky subgraph integration, each optimized for different operational scenarios.
High-Level Scraper: update_utils/update_goldsky.py
The update_utils/update_goldsky.py module serves as the user-friendly entry point. It exposes the update_goldsky() function, which uses the gql client library and pandas for CSV handling.
Key characteristics:
- Automatic CSV management: Handles headers, column flattening via
flatten_json, and pandas-based DataFrame appends. - Explicit cursor handling: Uses
save_cursor()andget_latest_cursor()to manage thegoldsky/cursor_state.jsonfile. - Simple invocation: Called directly or via the orchestrator in
update_all.py.
Parallel Backfill: parallel_sync.py
The parallel_sync.py module provides a low-level, performance-optimized variant designed for large historical backfills.
Key characteristics:
- Worker-based parallelism: Splits the missing time range across multiple worker threads (configurable via
--workersflag). - Manual HTTP handling: Uses plain
requestsrather than thegqlclient for lower overhead. - Segment merging: Stores temporary segment files per worker and merges them into
goldsky/orderFilled.csvafter all workers complete.
Running the Goldsky Subgraph Integration
Single-Threaded Update
To fetch the latest events using the high-level interface:
from update_utils.update_goldsky import update_goldsky
# Resumes from cursor state or CSV tail, fetches new events, appends to CSV
update_goldsky()
Source: [update_utils/update_goldsky.py](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py)
Full Pipeline Execution
The update_all.py orchestrator runs the complete data pipeline:
uv run python update_all.py
This executes three stages sequentially: market updates (update_markets()), Goldsky synchronization (update_goldsky()), and trade processing (process_live()).
Source: [update_all.py](https://github.com/warproxxx/poly_data/blob/main/update_all.py)
Parallel Historical Sync
For catching up large backlogs quickly:
python parallel_sync.py --workers 8
This command partitions the missing time range into eight segments, processes them concurrently, and merges the results.
Source: [parallel_sync.py](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)
Cursor State and Data Persistence
The integration maintains state in goldsky/cursor_state.json to enable crash recovery. Inspect the current position programmatically:
import json
import pathlib
cursor_path = pathlib.Path('goldsky/cursor_state.json')
if cursor_path.is_file():
state = json.loads(cursor_path.read_text())
print(f"Resume from timestamp {state['last_timestamp']} (sticky={state['sticky_timestamp']})")
else:
print("No cursor file – scraper will start from the beginning or CSV tail.")
The destination CSV at goldsky/orderFilled.csv stores every raw event with a fixed schema, serving as the authoritative source for downstream trade analysis.
Summary
- The Goldsky subgraph integration in Poly Data provides resumable, incremental ingestion of Polymarket
orderFilledevents via GraphQL. - Two implementations exist:
update_utils/update_goldsky.pyfor standard operations andparallel_sync.pyfor high-speed backfills. - "Sticky" pagination ensures complete data coverage during high-throughput periods by paginating on
idwhen timestamps collide. - Cursor persistence in
goldsky/cursor_state.jsonprevents duplicates and enables crash recovery. - The orchestrator
update_all.pyintegrates Goldsky fetching into the broader data pipeline alongside market and trade processing.
Frequently Asked Questions
How does the Goldsky subgraph integration prevent duplicate events?
The scraper deduplicates events on the id field before appending to goldsky/orderFilled.csv. Additionally, it maintains cursor state (last_timestamp, last_id, sticky_timestamp) in goldsky/cursor_state.json to resume exactly where it left off, ensuring that interrupted runs do not re-fetch already processed records.
What is the purpose of the "sticky" timestamp logic?
When multiple events share the same timestamp and a GraphQL batch reaches the first limit, the "sticky" logic keeps the scraper pinned to that timestamp while paginating by id until the timestamp changes. This prevents gaps that would otherwise occur if the scraper advanced to the next timestamp while events remained unprocessed at the current one.
When should I use parallel_sync.py instead of update_goldsky.py?
Use parallel_sync.py when performing large historical backfills that span significant time ranges, as it splits the workload across multiple worker threads for faster ingestion. Use update_utils/update_goldsky.py for regular incremental updates or when you need simple pandas-based CSV handling and automatic cursor management.
Where does the integration store the fetched event data?
Raw events are written to goldsky/orderFilled.csv with a fixed column schema including timestamp, maker, taker, makerAssetId, and transactionHash. The current pagination state is temporarily stored in goldsky/cursor_state.json during execution and removed upon successful completion.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →