How Poly Data Achieves Idempotent Data Collection: CSV Offsets and Cursor State

Poly Data guarantees idempotent data collection by combining offset‑based resumption in market updates, JSON cursor files for event streaming, and explicit deduplication logic that eliminates duplicate records across reruns.

The warproxxx/poly_data repository provides a robust data ingestion pipeline designed to be safely re‑executable without creating duplicates. Its idempotency strategies rely on persistent state tracking—either by counting existing CSV rows or by maintaining cursor files—ensuring that interrupted runs resume exactly where they left off while completed runs start fresh.

Offset‑Based Resume in Market Data Updates

The update_markets utility implements idempotency through line‑counting logic that calculates the next API offset based on data already present in the CSV file.

Counting Existing Records with count_csv_lines

In update_utils/update_markets.py, the script first determines how many records exist locally before making any API requests. The count_csv_lines() function returns the number of data rows already stored, which initializes the current_offset variable.


# From update_utils/update_markets.py

current_offset = count_csv_lines(csv_filename)

When the markets CSV already exists, the script opens it in append mode and sets the API request parameter to skip previously fetched rows: params['offset'] = current_offset.

Appending Only New Rows

After fetching a batch, the script increments the offset by the number of rows actually processed (current_offset += batch_count), ensuring subsequent requests always start after the last successfully written record. This guarantees that reruns never re‑process earlier rows, even if the script is executed multiple times consecutively.

Cursor‑Driven Pagination for Goldsky Events

The update_goldsky module handles order‑filled event data from Goldsky using a sophisticated cursor mechanism that tracks precise temporal positions.

Persisting State with cursor_state.json

Rather than counting rows, this scraper maintains a dedicated JSON file at goldsky/cursor_state.json that records the last_timestamp, last_id, and a sticky_timestamp. The get_latest_cursor() function reads this state to build the GraphQL where clause, or falls back to the last CSV line if the cursor file is missing.


# From update_utils/update_goldsky.py

last_timestamp, last_id, sticky_timestamp = get_latest_cursor()

After processing each batch, save_cursor(last_timestamp, last_id, sticky_timestamp) persists the exact position to disk, enabling seamless resumption across interrupted executions.

Handling Sticky Timestamps

The sticky timestamp logic solves pagination challenges when multiple events share identical timestamps. When a batch is full and all rows share the same timestamp, the scraper continues paginating on that timestamp using id ordering until no rows remain. This prevents gaps or overlaps that could occur with simple timestamp‑based cursors.

Deduplication Before Writing

As a final safety mechanism, the dataframe undergoes drop_duplicates(subset=['id']) before any write operation. This guards against race conditions where overlapping batches might contain identical records, ensuring the output CSV remains clean regardless of execution anomalies.

Cursor Cleanup and Full Reset Behavior

Upon successful completion of a full scrape, update_goldsky removes the cursor file via os.remove(CURSOR_FILE). This deletion guarantees that subsequent runs start from the beginning of time only when the previous execution truly finished. If the cursor file persists, the next run interprets this as an incomplete previous execution and resumes from the saved position rather than starting fresh.

Summary

  • Offset calculation: update_markets counts existing CSV lines to set the API offset, appending only new records on each run.
  • Cursor persistence: update_goldsky uses goldsky/cursor_state.json to track the last processed timestamp and ID, enabling precise resumption.
  • Sticky timestamp handling: When batches contain identical timestamps, pagination continues by id until exhaustion, preventing data gaps.
  • Deduplication: Both utilities apply drop_duplicates(subset=['id']) as a safety net against overlapping writes.
  • Clean‑up protocol: Removing the cursor file after successful completion ensures the next full scrape starts from zero, while interruptions trigger automatic resume.

Frequently Asked Questions

What happens if the Goldsky scraper is interrupted mid‑run?

If interrupted, the goldsky/cursor_state.json file persists the last successfully processed timestamp and ID. On the next execution, get_latest_cursor() reads this file and resumes fetching from that exact position, preventing duplicate or missing records.

How does Poly Data prevent duplicate market records on rerun?

The update_markets script calculates current_offset by counting existing CSV rows via count_csv_lines(), then requests data starting from that offset. It appends only new batches and increments the offset accordingly, ensuring previously fetched records are never requested again.

What is the purpose of the sticky timestamp in Goldsky updates?

The sticky timestamp handles scenarios where a single timestamp contains more events than one API batch can return. It forces the scraper to remain on that timestamp and paginate by id until all records for that moment are exhausted, preventing the cursor from advancing prematurely and leaving gaps.

When does Poly Data start a full scrape instead of resuming?

A full scrape begins only when update_goldsky completes successfully and removes goldsky/cursor_state.json using os.remove(CURSOR_FILE). If the cursor file exists, the script assumes the previous run was interrupted and resumes from the saved cursor position rather than starting from the beginning of time.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →