# Goldsky Subgraph Integration in Poly Data: Resumable GraphQL Scraping Explained

> Explore the Goldsky subgraph integration in Poly Data. Learn how this resumable GraphQL scraper fetches orderFilled events, uses cursors, and stores data in CSVs for efficient, recoverable orderbook data.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: how-to-guide
- Published: 2026-04-21

---

**The Goldsky subgraph integration in Poly Data is a resumable GraphQL scraper that incrementally fetches `orderFilled` events from the Polymarket orderbook, paginates through them using timestamp and ID-based cursors, and persists the data to CSV while maintaining state in JSON to prevent duplicates and enable crash recovery.**

Poly Data relies on this integration to capture every trade executed on the Polymarket orderbook. According to the `warproxxx/poly_data` source code, the implementation operates as a fault-tolerant pipeline that queries the Goldsky hosted subgraph, handles high-throughput pagination edge cases, and stores raw event data locally for downstream analytics.

## How the Goldsky Subgraph Integration Works

The integration functions as a **resumable GraphQL scraper** that performs five core operations:

- **Queries the Goldsky hosted subgraph** (`orderbook-subgraph`) via a GraphQL endpoint to retrieve `orderFilled` events.
- **Pages through events** using a combination of timestamp-based and "sticky" ID-based pagination to ensure no events are missed when many records share the same timestamp.
- **Writes raw events** to `goldsky/orderFilled.csv` with a fixed column order including `timestamp`, `maker`, `makerAssetId`, `taker`, `takerAmountFilled`, and `transactionHash`.
- **Persists cursor state** in [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) using fields `last_timestamp`, `last_id`, and `sticky_timestamp` so that subsequent runs resume exactly where the previous execution stopped.
- **Deduplicates and cleans up** by removing events with duplicate `id` values before appending to the CSV, and deletes the cursor file upon successful completion.

### The "Sticky" Pagination Logic

A critical feature of the implementation is the **sticky timestamp logic** (lines 81-107 in [`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py) and lines 71-104 in [`parallel_sync.py`](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)). When a GraphQL batch returns the maximum `first` size and all events share an identical timestamp, the scraper remains on that timestamp and paginates by `id` until the timestamp advances. This guarantees **complete, gap-free coverage** even during high-throughput periods where thousands of orders fill simultaneously.

Both implementations share the same GraphQL query shape:

```graphql
orderFilledEvents(
    orderBy: timestamp,
    orderDirection: asc,
    first: <batch_size>,
    where: { <timestamp or sticky clause> }
) {
    id timestamp maker makerAmountFilled makerAssetId
    taker takerAmountFilled takerAssetId transactionHash
}

```

## Implementation Variants

The repository provides two concrete implementations of the Goldsky subgraph integration, each optimized for different operational scenarios.

### High-Level Scraper: update_utils/update_goldsky.py

The **[`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py)** module serves as the user-friendly entry point. It exposes the `update_goldsky()` function, which uses the `gql` client library and pandas for CSV handling.

Key characteristics:

- **Automatic CSV management**: Handles headers, column flattening via `flatten_json`, and pandas-based DataFrame appends.
- **Explicit cursor handling**: Uses `save_cursor()` and `get_latest_cursor()` to manage the [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) file.
- **Simple invocation**: Called directly or via the orchestrator in [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py).

### Parallel Backfill: parallel_sync.py

The **[`parallel_sync.py`](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)** module provides a low-level, performance-optimized variant designed for large historical backfills.

Key characteristics:

- **Worker-based parallelism**: Splits the missing time range across multiple worker threads (configurable via `--workers` flag).
- **Manual HTTP handling**: Uses plain `requests` rather than the `gql` client for lower overhead.
- **Segment merging**: Stores temporary segment files per worker and merges them into `goldsky/orderFilled.csv` after all workers complete.

## Running the Goldsky Subgraph Integration

### Single-Threaded Update

To fetch the latest events using the high-level interface:

```python
from update_utils.update_goldsky import update_goldsky

# Resumes from cursor state or CSV tail, fetches new events, appends to CSV

update_goldsky()

```

*Source*: [[`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py)](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py)

### Full Pipeline Execution

The [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py) orchestrator runs the complete data pipeline:

```bash
uv run python update_all.py

```

This executes three stages sequentially: market updates (`update_markets()`), Goldsky synchronization (`update_goldsky()`), and trade processing (`process_live()`).

*Source*: [[`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py)](https://github.com/warproxxx/poly_data/blob/main/update_all.py)

### Parallel Historical Sync

For catching up large backlogs quickly:

```bash
python parallel_sync.py --workers 8

```

This command partitions the missing time range into eight segments, processes them concurrently, and merges the results.

*Source*: [[`parallel_sync.py`](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)

## Cursor State and Data Persistence

The integration maintains state in [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) to enable crash recovery. Inspect the current position programmatically:

```python
import json
import pathlib

cursor_path = pathlib.Path('goldsky/cursor_state.json')
if cursor_path.is_file():
    state = json.loads(cursor_path.read_text())
    print(f"Resume from timestamp {state['last_timestamp']} (sticky={state['sticky_timestamp']})")
else:
    print("No cursor file – scraper will start from the beginning or CSV tail.")

```

The destination CSV at `goldsky/orderFilled.csv` stores every raw event with a fixed schema, serving as the authoritative source for downstream trade analysis.

## Summary

- The **Goldsky subgraph integration** in Poly Data provides resumable, incremental ingestion of Polymarket `orderFilled` events via GraphQL.
- **Two implementations** exist: [`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py) for standard operations and [`parallel_sync.py`](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py) for high-speed backfills.
- **"Sticky" pagination** ensures complete data coverage during high-throughput periods by paginating on `id` when timestamps collide.
- **Cursor persistence** in [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) prevents duplicates and enables crash recovery.
- The **orchestrator** [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py) integrates Goldsky fetching into the broader data pipeline alongside market and trade processing.

## Frequently Asked Questions

### How does the Goldsky subgraph integration prevent duplicate events?

The scraper deduplicates events on the `id` field before appending to `goldsky/orderFilled.csv`. Additionally, it maintains cursor state (`last_timestamp`, `last_id`, `sticky_timestamp`) in [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) to resume exactly where it left off, ensuring that interrupted runs do not re-fetch already processed records.

### What is the purpose of the "sticky" timestamp logic?

When multiple events share the same timestamp and a GraphQL batch reaches the `first` limit, the "sticky" logic keeps the scraper pinned to that timestamp while paginating by `id` until the timestamp changes. This prevents gaps that would otherwise occur if the scraper advanced to the next timestamp while events remained unprocessed at the current one.

### When should I use parallel_sync.py instead of update_goldsky.py?

Use **[`parallel_sync.py`](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)** when performing large historical backfills that span significant time ranges, as it splits the workload across multiple worker threads for faster ingestion. Use **[`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py)** for regular incremental updates or when you need simple pandas-based CSV handling and automatic cursor management.

### Where does the integration store the fetched event data?

Raw events are written to `goldsky/orderFilled.csv` with a fixed column schema including `timestamp`, `maker`, `taker`, `makerAssetId`, and `transactionHash`. The current pagination state is temporarily stored in [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) during execution and removed upon successful completion.