Sticky Timestamp Mechanism for Order Events in warproxxx/poly_data
The sticky timestamp mechanism ensures complete, lossless pagination of blockchain order events by temporarily switching from timestamp-based to ID-based cursor pagination when multiple events share identical timestamps, preventing data gaps in GraphQL queries.
When scraping orderFilledEvents from the Goldsky GraphQL endpoint, the warproxxx/poly_data repository addresses a critical edge case in high-frequency trading data: hundreds or thousands of events can share the exact same timestamp. Standard pagination using timestamp_gt risks skipping events that fall on boundary timestamps, creating incomplete datasets. The sticky timestamp mechanism solves this by detecting uniform timestamps in full batches and switching to a hybrid filter that keeps the timestamp constant while paginating by event ID.
The Problem: Why Standard Pagination Fails for Dense Timestamps
In blockchain trading data, multiple order events frequently execute within the same millisecond. A conventional scraper advances through datasets using a timestamp_gt filter to fetch events after the last processed time. However, if a batch contains 1000 events all timestamped 2024-01-15T10:30:00.000Z, and the scraper advances last_timestamp to that value, the subsequent query requesting timestamp_gt: "2024-01-15T10:30:00.000Z" will skip any remaining events at that exact timestamp. This creates data loss during high-density trading periods.
How the Sticky Timestamp Mechanism Works
The implementation in update_utils/update_goldsky.py detects when a batch consists entirely of events sharing a single timestamp. Instead of advancing the time cursor, it enters sticky mode and重构 the query filter to maintain the timestamp constant while paginating by the event's unique ID.
Triggering Sticky Mode
When a batch returns the maximum number of records (at_once), the code evaluates timestamp uniformity. According to lines 80-92 in update_goldsky.py:
if len(df) >= at_once:
# Batch is full – check timestamp uniformity
if batch_first_timestamp == batch_last_timestamp:
# All events share the same timestamp → stay sticky
sticky_timestamp = batch_last_timestamp
last_id = batch_last_id
else:
# Mixed timestamps → stay sticky at the last timestamp
sticky_timestamp = batch_last_timestamp
last_id = batch_last_id
Whether the batch contains uniform timestamps or mixed values, the system captures the final timestamp and ID, preparing for sticky pagination.
Query Construction in Sticky Mode
Once sticky_timestamp is populated, the GraphQL query builder switches from timestamp_gt to a combined filter. Lines 121-125 construct the where clause:
if sticky_timestamp is not None:
# We're in sticky mode: stay at this timestamp and paginate by id
where_clause = f'timestamp: "{sticky_timestamp}", id_gt: "{last_id}"'
else:
# Normal mode: advance by timestamp
where_clause = f'timestamp_gt: "{last_timestamp}"'
This ensures that while the timestamp remains fixed, the query advances through records using the unique id field, guaranteeing sequential retrieval without gaps.
Exiting Sticky Mode and Handling Empty Results
The scraper exits sticky mode when it detects the batch is no longer full, indicating all events for that timestamp have been processed. Lines 94-100 handle the transition:
else:
# Batch not full – we have all events, can advance normally
if sticky_timestamp is not None:
# We were in sticky mode, now exhausted – advance past this timestamp
last_timestamp = sticky_timestamp
sticky_timestamp = None
last_id = None
Additionally, if a query returns an empty result set while a sticky timestamp is active, the code interprets this as having exhausted all events at that timestamp (lines 160-164), preventing infinite loops and allowing progression to subsequent timestamps.
Implementation Details in update_goldsky.py
The core logic resides in update_utils/update_goldsky.py, which orchestrates the Goldsky scraper execution. Key architectural components include:
- Batch uniformity detection: Comparing
batch_first_timestampagainstbatch_last_timestampto identify single-timestamp windows - Cursor state management: Tracking
sticky_timestamp,last_timestamp, andlast_idacross pagination cycles - Dynamic query parameterization: Building GraphQL where clauses based on the current pagination mode
The scraper utilizes helper functions from poly_utils/utils.py (including flatten and save_cursor) for data transformation and cursor persistence, while update_all.py orchestrates the overall execution pipeline.
Summary
- The sticky timestamp mechanism prevents data loss when paginating order events that share identical timestamps in the Goldsky GraphQL dataset.
- When a full batch contains uniform timestamps, the system enters sticky mode and switches from
timestamp_gttotimestamp: "<value>", id_gt: "<last_id>"filtering. - This approach guarantees complete retrieval of high-density event clusters by paginating via unique IDs while keeping the timestamp constant.
- The implementation automatically exits sticky mode when batches are no longer full, advancing the timestamp cursor and resetting state variables to resume normal pagination.
Frequently Asked Questions
What triggers the sticky timestamp mechanism to activate?
The mechanism activates when the scraper retrieves a full batch of events (reaching the at_once limit) and detects that the first and last events in that batch share the same timestamp. This condition indicates potential additional events exist at that timestamp that would be skipped by standard time-based pagination.
How does the GraphQL query change when entering sticky mode?
In normal operation, the query uses timestamp_gt: "<last_timestamp>" to fetch newer events. In sticky mode, it switches to timestamp: "<sticky_timestamp>", id_gt: "<last_id>", keeping the timestamp fixed while using the event ID as the pagination cursor to traverse events within that specific timestamp.
What happens if no results are returned while in sticky mode?
If a query returns an empty result set while sticky_timestamp is active, the code interprets this as having exhausted all events at that timestamp (lines 160-164). It then advances last_timestamp to the sticky timestamp value and clears the sticky state, allowing the next query to proceed to subsequent timestamps.
Where is the sticky timestamp logic implemented in the repository?
The primary implementation is in update_utils/update_goldsky.py, specifically within the pagination loop that constructs GraphQL queries and manages cursor state. Supporting utilities for cursor persistence are located in poly_utils/utils.py, and the scraper is invoked through update_all.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →