# How to Resume Goldsky Data Scraping in Poly Data: Complete Guide

> Easily resume Goldsky data scraping in Poly Data using the JSON cursor state file. Pick up where you left off without re-downloading historical data. Learn how to implement this essential feature.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: how-to-guide
- Published: 2026-04-21

---

**Poly Data uses a JSON cursor state file to remember the last timestamp and event ID, allowing the scraper to resume exactly where it left off without re-downloading historical data.**

The `warproxxx/poly_data` repository provides a robust Goldsky subgraph scraper that handles interruptions gracefully. Instead of fetching the entire dataset on every run, the system maintains a persistent cursor that tracks progress through the `orderFilled` events, ensuring efficient incremental updates.

## How the Resume Mechanism Works

The scraper implements a stateful pagination system that persists progress to disk after every successful batch. This design protects against network failures, script interruptions, or system crashes.

### Step 1: Initialize the Directory

Before fetching data, the script ensures the `goldsky` directory exists to store both the output CSV and the cursor state file.

In [`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py) (lines 18‑20), the scraper checks:

```python
if not os.path.isdir('goldsky'):
    os.makedirs('goldsky', exist_ok=True)

```

### Step 2: Load the Cursor State

The `get_latest_cursor()` function (lines 33‑44) attempts to read [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json). If present, it returns the stored `last_timestamp`, `last_id`, and optional `sticky_timestamp`.

If no cursor file exists, the function falls back to reading the last line of `goldsky/orderFilled.csv` using the Unix `tail` command for speed (lines 65‑81), with a pandas fallback for compatibility:

```python
result = subprocess.run(['tail', '-n', '1', csv_path], capture_output=True, text=True)

```

This timestamp becomes the starting point for the next query.

### Step 3: Handle Sticky Timestamps

When many events share the same timestamp (common in high-frequency trading data), standard timestamp-based pagination would skip records. The scraper detects this condition (lines 80‑86) and enters "sticky" mode:

```python
if batch_first_timestamp == batch_last_timestamp:
    sticky_timestamp = batch_first_timestamp

```

In this mode, the cursor keeps the timestamp fixed and paginates by `id` instead, ensuring no events are missed within the same second.

### Step 4: Execute the Query Loop

Each iteration builds a GraphQL `where` clause based on the current cursor state (lines 20‑26). Normal mode uses `timestamp_gt`, while sticky mode uses both `timestamp` and `id_gt`:

```python
if sticky_timestamp is not None:
    where_clause = f'{{timestamp: "{sticky_timestamp}", id_gt: "{last_id}"}}'
else:
    where_clause = f'{{timestamp_gt: "{last_timestamp}"}}'

```

### Step 5: Persist State After Each Batch

After writing a batch to CSV, `save_cursor()` (lines 21‑23) immediately writes the latest position to disk:

```python
save_cursor(last_timestamp, last_id, sticky_timestamp)

```

This atomic write ensures that an interrupted run can resume precisely from the last committed record.

### Step 6: Clean Up on Completion

When the scraper reaches the end of available data, it removes the cursor file (lines 27‑30), signaling a fresh start for the next full run:

```python
if os.path.isfile(CURSOR_FILE):
    os.remove(CURSOR_FILE)

```

## Core Files and Responsibilities

Understanding the module structure helps with debugging and custom integration.

- **[`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py)** – Contains the main `scrape()` function, cursor logic (`get_latest_cursor()`, `save_cursor()`), and the GraphQL query loop.
- **[`parallel_sync.py`](https://github.com/warproxxx/poly_data/blob/main/parallel_sync.py)** – Provides the low‑level `goldsky_query()` helper function used by the scraper to execute GraphQL requests.
- **[`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py)** – The convenience entry point that orchestrates market updates, Goldsky scraping via `update_goldsky()`, and live data processing.
- **[`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py)** – Demonstrates how downstream modules consume the generated `orderFilled.csv` file.

## Practical Code Examples

### Run a Full Update (Auto-Resume)

Execute the main entry point to continue from the existing cursor:

```bash
python update_all.py

```

The script prints "Updating goldsky" and invokes `update_goldsky()`, which automatically detects and loads any existing cursor state.

### Resume Only the Goldsky Scraper

For targeted updates without running the full pipeline:

```bash
python -c "from update_utils.update_goldsky import update_goldsky; update_goldsky()"

```

If [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) exists, scraping continues from that point; otherwise, it starts from the end of the CSV or from timestamp zero.

### Inspect or Reset the Cursor Manually

Force a full re-scrape by removing the cursor file, or check current progress:

```python
import json
import os

CURSOR = "goldsky/cursor_state.json"

# Display current cursor state

if os.path.isfile(CURSOR):
    with open(CURSOR, 'r') as f:
        print(json.load(f))
else:
    print("No cursor file found. Scraper will start from CSV tail or timestamp 0.")

# Uncomment to force restart:

# os.remove(CURSOR)

```

### Verify Sticky Mode Status

Check if the scraper is currently paginating within a single timestamp:

```python
from update_utils.update_goldsky import get_latest_cursor

timestamp, last_id, sticky = get_latest_cursor()
print(f"Timestamp: {timestamp}, LastID: {last_id}, Sticky: {sticky}")

```

When `sticky` is not `None`, the next query will constrain both `timestamp` and `id_gt` to avoid skipping events.

## Summary

- **Poly Data** avoids redundant downloads by maintaining a [`cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/cursor_state.json) file that records the last successful fetch position.
- The **resume logic** lives in [`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py), specifically within `get_latest_cursor()` and `save_cursor()`.
- **Sticky timestamp handling** prevents data loss when multiple events occur in the same second by switching from timestamp-based to ID-based pagination.
- The system falls back to reading the CSV tail if no cursor exists, ensuring continuity even if the JSON state is manually deleted.
- State is persisted **before** each network request, making the scraper resilient to crashes and interruptions.

## Frequently Asked Questions

### What happens if the cursor file becomes corrupted?

If [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) contains invalid JSON or is unreadable, `get_latest_cursor()` catches the exception and falls back to reading the last line of `goldsky/orderFilled.csv`. If the CSV is also empty or missing, the scraper starts from timestamp zero, effectively performing a full refresh.

### How does the sticky timestamp feature prevent duplicate or missed records?

When a batch of fetched records all share the same timestamp, simply incrementing the timestamp would skip the remaining events from that second. The scraper detects this condition and sets `sticky_timestamp`, keeping the timestamp fixed while paginating by `id_gt`. This ensures every event is captured even when thousands occur within the same block timestamp.

### Can I run the Goldsky scraper independently of the main update script?

Yes. While [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py) provides the standard orchestration, you can import and run `update_goldsky()` directly from [`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py). This is useful for debugging, testing specific date ranges, or running the scraper on a different schedule than the market updates.

### Where does the scraper store the actual event data?

Raw order-filled events are appended to `goldsky/orderFilled.csv` (created at runtime). The cursor state is stored separately in [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json). This separation allows you to archive or move the CSV without affecting the resume capability, provided you preserve the JSON cursor file or accept a fallback to the CSV tail on the next run.