How to Parse Cursor IDE Telemetry from SQLite state.vscdb

The Uber ADR Sensor provides a CursorParser class that connects to the Cursor IDE's SQLite state.vscdb database, extracts conversation bubbles and composer metadata, and converts them into standardized AgentEvent objects for downstream analysis.

The uber/ADR repository contains an open-source telemetry sensor that enables you to parse Cursor IDE telemetry from SQLite state.vscdb and normalize it into a common schema. This article provides a deep dive into the CursorParser implementation, including specific file paths, method signatures, and executable code samples derived directly from the source.

Database Architecture and Location

File System Paths for state.vscdb

Cursor persists interaction history in a SQLite database named state.vscdb. On macOS, this file resides at ~/Library/Application Support/Cursor/User/globalStorage/state.vscdb, while Linux systems store it at ~/.config/Cursor/User/globalStorage/state.vscdb. The CursorParser constructor automatically selects the appropriate platform-specific path during initialization according to the path resolution logic found in cursor_parser.py#L30-L36.

Internal Schema Structure

The database contains two critical data categories: composerData entries (metadata including creation and update timestamps) and bubbleId entries (individual conversation messages). The parser queries keys matching composerData:% for session metadata and bubbleId:% for message content to reconstruct complete conversation threads.

Step-by-Step Parsing Workflow

Database Connection Initialization

When instantiated, CursorParser establishes a read-only SQLite connection using Python's built-in sqlite3 module. The connection is created via sqlite3.connect(self.db_path) as implemented in cursor_parser.py#L58-L61, and a cursor object is obtained for batched reads to ensure minimal memory footprint when processing large histories.

Metadata Extraction and Timestamp Normalization

The private method get_composer_metadata executes batched SQL queries to retrieve all rows where the key starts with composerData:. Each blob is decoded from JSON into a dictionary mapping composer_id to its createdAt and lastUpdatedAt timestamps, enabling chronological filtering. This logic resides in cursor_parser.py#L44-L65.

Conversation Filtering and Bubble Grouping

To maintain performance, the parser filters conversations older than a configurable window (defaulting to 14 days) before processing content. The _iter_cursor_batches generator yields rows in chunks, limiting resident memory when iterating over thousands of bubbles. Each bubble is JSON-decoded and grouped by its conv_id for sequential processing as shown in cursor_parser.py#L89-L98.

Event Construction and Schema Compliance

For each conversation, bubbles are sorted by their internal _bubble_id. The parser extracts plain text using extract_text_from_bubble and identifies tool invocations via extract_tools_from_bubble. These components populate ChatMessage and ToolUsage objects defined in agent_event_schema.py. A complete AgentEvent is assembled with a unified timestamp derived from composer metadata and appended to the result list cursor_parser.py#L86-L123.

Implementation Examples

from Sensor.adr_sensor.parsers.cursor_parser import CursorParser

# Initialize with default 14-day window

parser = CursorParser()
events = parser.parse_all()  # Returns List[AgentEvent]

# Inspect first parsed event

if events:
    first = events[0]
    print(f"Session: {first.session_id}")
    print(f"Timestamp: {first.timestamp}")
    for msg in first.chat_history:
        print(f"{msg.role.upper()}: {msg.content}")
        for tool in msg.tools:
            print(f"  → Tool: {tool.tool_name}({tool.arguments})")

Customizing the retention window:


# Parse only the last 7 days of telemetry

parser = CursorParser(max_age_days=7)
recent_events = parser.parse_all()

Command-line execution:

python -m Sensor.adr_sensor.cli cursor

# Outputs JSON lines of AgentEvent objects

Key Source Files Reference

Component Description Source
cursor_parser.py Core implementation reading state.vscdb, extracting metadata, batching bubble rows, and building AgentEvent objects. Sensor/adr_sensor/parsers/cursor_parser.py
base_parser.py Abstract base class defining the parse_all() interface inherited by CursorParser. Sensor/adr_sensor/parsers/base_parser.py
agent_event_schema.py Data classes (AgentEvent, ChatMessage, ToolUsage) representing normalized telemetry. Sensor/adr_sensor/schemas/agent_event_schema.py
timestamp_utils.py Helper functions for normalizing timestamp formats from Cursor's metadata. Sensor/adr_sensor/utils/timestamp_utils.py
string_utils.py Utilities for truncating large strings while preserving semantic fragments. Sensor/adr_sensor/utils/string_utils.py
cli.py Command-line entry point exposing the cursor parser as a sub-command. Sensor/adr_sensor/cli.py

Summary

  • The CursorParser class in uber/ADR provides a production-ready solution to parse Cursor IDE telemetry from SQLite state.vscdb.
  • It implements automatic path resolution for macOS and Linux, batched SQL queries for memory efficiency, and configurable time-based filtering via the max_age_days parameter.
  • Raw database entries are transformed into standardized AgentEvent objects containing ChatMessage and ToolUsage data structures.
  • All functionality is backed by concrete implementations in cursor_parser.py, agent_event_schema.py, and supporting utility modules.

Frequently Asked Questions

Where does Cursor IDE store its telemetry database?

Cursor IDE stores its telemetry in a SQLite database named state.vscdb. On macOS, the file is located at ~/Library/Application Support/Cursor/User/globalStorage/state.vscdb, while on Linux it resides at ~/.config/Cursor/User/globalStorage/state.vscdb. The CursorParser automatically resolves these paths based on the host operating system.

What data structure does the CursorParser return?

The parse_all() method returns a list of AgentEvent objects as defined in agent_event_schema.py. Each AgentEvent contains a session_id, unified timestamp, and a chat_history list comprising ChatMessage instances that may include associated ToolUsage records.

How does the parser handle large conversation histories?

The parser implements batched reading via _iter_cursor_batches to process SQLite rows in chunks rather than loading the entire database into memory. Additionally, it filters conversations older than the configured max_age_days (default 14 days) before processing individual bubbles, significantly reducing computational overhead.

Can I customize the time window for parsing events?

Yes. The CursorParser constructor accepts an optional max_age_days parameter. Instantiating CursorParser(max_age_days=7) restricts processing to conversations updated within the last seven days, allowing targeted analysis of recent activity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →