Goldsky Subgraph Integration in Poly Data: Resumable GraphQL Scraping Explained

The Goldsky subgraph integration in Poly Data is a resumable GraphQL scraper that incrementally fetches orderFilled events from the Polymarket orderbook, paginates through them using timestamp and ID-based cursors, and persists the data to CSV while maintaining state in JSON to prevent duplicates and enable crash recovery.

Poly Data relies on this integration to capture every trade executed on the Polymarket orderbook. According to the warproxxx/poly_data source code, the implementation operates as a fault-tolerant pipeline that queries the Goldsky hosted subgraph, handles high-throughput pagination edge cases, and stores raw event data locally for downstream analytics.

How the Goldsky Subgraph Integration Works

The integration functions as a resumable GraphQL scraper that performs five core operations:

  • Queries the Goldsky hosted subgraph (orderbook-subgraph) via a GraphQL endpoint to retrieve orderFilled events.
  • Pages through events using a combination of timestamp-based and "sticky" ID-based pagination to ensure no events are missed when many records share the same timestamp.
  • Writes raw events to goldsky/orderFilled.csv with a fixed column order including timestamp, maker, makerAssetId, taker, takerAmountFilled, and transactionHash.
  • Persists cursor state in goldsky/cursor_state.json using fields last_timestamp, last_id, and sticky_timestamp so that subsequent runs resume exactly where the previous execution stopped.
  • Deduplicates and cleans up by removing events with duplicate id values before appending to the CSV, and deletes the cursor file upon successful completion.

The "Sticky" Pagination Logic

A critical feature of the implementation is the sticky timestamp logic (lines 81-107 in update_utils/update_goldsky.py and lines 71-104 in parallel_sync.py). When a GraphQL batch returns the maximum first size and all events share an identical timestamp, the scraper remains on that timestamp and paginates by id until the timestamp advances. This guarantees complete, gap-free coverage even during high-throughput periods where thousands of orders fill simultaneously.

Both implementations share the same GraphQL query shape:

orderFilledEvents(
    orderBy: timestamp,
    orderDirection: asc,
    first: <batch_size>,
    where: { <timestamp or sticky clause> }
) {
    id timestamp maker makerAmountFilled makerAssetId
    taker takerAmountFilled takerAssetId transactionHash
}

Implementation Variants

The repository provides two concrete implementations of the Goldsky subgraph integration, each optimized for different operational scenarios.

High-Level Scraper: update_utils/update_goldsky.py

The update_utils/update_goldsky.py module serves as the user-friendly entry point. It exposes the update_goldsky() function, which uses the gql client library and pandas for CSV handling.

Key characteristics:

  • Automatic CSV management: Handles headers, column flattening via flatten_json, and pandas-based DataFrame appends.
  • Explicit cursor handling: Uses save_cursor() and get_latest_cursor() to manage the goldsky/cursor_state.json file.
  • Simple invocation: Called directly or via the orchestrator in update_all.py.

Parallel Backfill: parallel_sync.py

The parallel_sync.py module provides a low-level, performance-optimized variant designed for large historical backfills.

Key characteristics:

  • Worker-based parallelism: Splits the missing time range across multiple worker threads (configurable via --workers flag).
  • Manual HTTP handling: Uses plain requests rather than the gql client for lower overhead.
  • Segment merging: Stores temporary segment files per worker and merges them into goldsky/orderFilled.csv after all workers complete.

Running the Goldsky Subgraph Integration

Single-Threaded Update

To fetch the latest events using the high-level interface:

from update_utils.update_goldsky import update_goldsky

# Resumes from cursor state or CSV tail, fetches new events, appends to CSV

update_goldsky()

Source: [update_utils/update_goldsky.py](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py)

Full Pipeline Execution

The update_all.py orchestrator runs the complete data pipeline:

uv run python update_all.py

This executes three stages sequentially: market updates (update_markets()), Goldsky synchronization (update_goldsky()), and trade processing (process_live()).

Source: [update_all.py](https://github.com/warproxxx/poly_data/blob/main/update_all.py)

Parallel Historical Sync

For catching up large backlogs quickly:

python parallel_sync.py --workers 8

This command partitions the missing time range into eight segments, processes them concurrently, and merges the results.

Source: [parallel_sync.py](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)

Cursor State and Data Persistence

The integration maintains state in goldsky/cursor_state.json to enable crash recovery. Inspect the current position programmatically:

import json
import pathlib

cursor_path = pathlib.Path('goldsky/cursor_state.json')
if cursor_path.is_file():
    state = json.loads(cursor_path.read_text())
    print(f"Resume from timestamp {state['last_timestamp']} (sticky={state['sticky_timestamp']})")
else:
    print("No cursor file – scraper will start from the beginning or CSV tail.")

The destination CSV at goldsky/orderFilled.csv stores every raw event with a fixed schema, serving as the authoritative source for downstream trade analysis.

Summary

  • The Goldsky subgraph integration in Poly Data provides resumable, incremental ingestion of Polymarket orderFilled events via GraphQL.
  • Two implementations exist: update_utils/update_goldsky.py for standard operations and parallel_sync.py for high-speed backfills.
  • "Sticky" pagination ensures complete data coverage during high-throughput periods by paginating on id when timestamps collide.
  • Cursor persistence in goldsky/cursor_state.json prevents duplicates and enables crash recovery.
  • The orchestrator update_all.py integrates Goldsky fetching into the broader data pipeline alongside market and trade processing.

Frequently Asked Questions

How does the Goldsky subgraph integration prevent duplicate events?

The scraper deduplicates events on the id field before appending to goldsky/orderFilled.csv. Additionally, it maintains cursor state (last_timestamp, last_id, sticky_timestamp) in goldsky/cursor_state.json to resume exactly where it left off, ensuring that interrupted runs do not re-fetch already processed records.

What is the purpose of the "sticky" timestamp logic?

When multiple events share the same timestamp and a GraphQL batch reaches the first limit, the "sticky" logic keeps the scraper pinned to that timestamp while paginating by id until the timestamp changes. This prevents gaps that would otherwise occur if the scraper advanced to the next timestamp while events remained unprocessed at the current one.

When should I use parallel_sync.py instead of update_goldsky.py?

Use parallel_sync.py when performing large historical backfills that span significant time ranges, as it splits the workload across multiple worker threads for faster ingestion. Use update_utils/update_goldsky.py for regular incremental updates or when you need simple pandas-based CSV handling and automatic cursor management.

Where does the integration store the fetched event data?

Raw events are written to goldsky/orderFilled.csv with a fixed column schema including timestamp, maker, taker, makerAssetId, and transactionHash. The current pagination state is temporarily stored in goldsky/cursor_state.json during execution and removed upon successful completion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →