# How to Parse Cursor IDE Telemetry from SQLite state.vscdb

> Learn to parse Cursor IDE telemetry from state.vscdb using Uber's CursorParser. Extract conversation bubbles and composer metadata into standardized AgentEvent objects for analysis.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-07

---

**The Uber ADR Sensor provides a `CursorParser` class that connects to the Cursor IDE's SQLite `state.vscdb` database, extracts conversation bubbles and composer metadata, and converts them into standardized `AgentEvent` objects for downstream analysis.**

The **uber/ADR** repository contains an open-source telemetry sensor that enables you to **parse Cursor IDE telemetry from SQLite state.vscdb** and normalize it into a common schema. This article provides a deep dive into the `CursorParser` implementation, including specific file paths, method signatures, and executable code samples derived directly from the source.

## Database Architecture and Location

### File System Paths for state.vscdb

Cursor persists interaction history in a SQLite database named `state.vscdb`. On macOS, this file resides at `~/Library/Application Support/Cursor/User/globalStorage/state.vscdb`, while Linux systems store it at `~/.config/Cursor/User/globalStorage/state.vscdb`. The `CursorParser` constructor automatically selects the appropriate platform-specific path during initialization according to the path resolution logic found in [cursor_parser.py#L30-L36](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/cursor_parser.py#L30-L36).

### Internal Schema Structure

The database contains two critical data categories: `composerData` entries (metadata including creation and update timestamps) and `bubbleId` entries (individual conversation messages). The parser queries keys matching `composerData:%` for session metadata and `bubbleId:%` for message content to reconstruct complete conversation threads.

## Step-by-Step Parsing Workflow

### Database Connection Initialization

When instantiated, `CursorParser` establishes a read-only SQLite connection using Python's built-in `sqlite3` module. The connection is created via `sqlite3.connect(self.db_path)` as implemented in [cursor_parser.py#L58-L61](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/cursor_parser.py#L58-L61), and a cursor object is obtained for batched reads to ensure minimal memory footprint when processing large histories.

### Metadata Extraction and Timestamp Normalization

The private method `get_composer_metadata` executes batched SQL queries to retrieve all rows where the key starts with `composerData:`. Each blob is decoded from JSON into a dictionary mapping `composer_id` to its `createdAt` and `lastUpdatedAt` timestamps, enabling chronological filtering. This logic resides in [cursor_parser.py#L44-L65](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/cursor_parser.py#L44-L65).

### Conversation Filtering and Bubble Grouping

To maintain performance, the parser filters conversations older than a configurable window (defaulting to 14 days) before processing content. The `_iter_cursor_batches` generator yields rows in chunks, limiting resident memory when iterating over thousands of bubbles. Each bubble is JSON-decoded and grouped by its `conv_id` for sequential processing as shown in [cursor_parser.py#L89-L98](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/cursor_parser.py#L89-L98).

### Event Construction and Schema Compliance

For each conversation, bubbles are sorted by their internal `_bubble_id`. The parser extracts plain text using `extract_text_from_bubble` and identifies tool invocations via `extract_tools_from_bubble`. These components populate `ChatMessage` and `ToolUsage` objects defined in [`agent_event_schema.py`](https://github.com/uber/ADR/blob/main/agent_event_schema.py). A complete `AgentEvent` is assembled with a unified timestamp derived from composer metadata and appended to the result list [cursor_parser.py#L86-L123](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/cursor_parser.py#L86-L123).

## Implementation Examples

```python
from Sensor.adr_sensor.parsers.cursor_parser import CursorParser

# Initialize with default 14-day window

parser = CursorParser()
events = parser.parse_all()  # Returns List[AgentEvent]

# Inspect first parsed event

if events:
    first = events[0]
    print(f"Session: {first.session_id}")
    print(f"Timestamp: {first.timestamp}")
    for msg in first.chat_history:
        print(f"{msg.role.upper()}: {msg.content}")
        for tool in msg.tools:
            print(f"  → Tool: {tool.tool_name}({tool.arguments})")

```

**Customizing the retention window:**

```python

# Parse only the last 7 days of telemetry

parser = CursorParser(max_age_days=7)
recent_events = parser.parse_all()

```

**Command-line execution:**

```bash
python -m Sensor.adr_sensor.cli cursor

# Outputs JSON lines of AgentEvent objects

```

## Key Source Files Reference

| Component | Description | Source |
|-----------|-------------|--------|
| [`cursor_parser.py`](https://github.com/uber/ADR/blob/main/cursor_parser.py) | Core implementation reading `state.vscdb`, extracting metadata, batching bubble rows, and building `AgentEvent` objects. | [Sensor/adr_sensor/parsers/cursor_parser.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/cursor_parser.py) |
| [`base_parser.py`](https://github.com/uber/ADR/blob/main/base_parser.py) | Abstract base class defining the `parse_all()` interface inherited by `CursorParser`. | [Sensor/adr_sensor/parsers/base_parser.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/base_parser.py) |
| [`agent_event_schema.py`](https://github.com/uber/ADR/blob/main/agent_event_schema.py) | Data classes (`AgentEvent`, `ChatMessage`, `ToolUsage`) representing normalized telemetry. | [Sensor/adr_sensor/schemas/agent_event_schema.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py) |
| [`timestamp_utils.py`](https://github.com/uber/ADR/blob/main/timestamp_utils.py) | Helper functions for normalizing timestamp formats from Cursor's metadata. | [Sensor/adr_sensor/utils/timestamp_utils.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/timestamp_utils.py) |
| [`string_utils.py`](https://github.com/uber/ADR/blob/main/string_utils.py) | Utilities for truncating large strings while preserving semantic fragments. | [Sensor/adr_sensor/utils/string_utils.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/string_utils.py) |
| [`cli.py`](https://github.com/uber/ADR/blob/main/cli.py) | Command-line entry point exposing the cursor parser as a sub-command. | [Sensor/adr_sensor/cli.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/cli.py) |

## Summary

- The **CursorParser** class in uber/ADR provides a production-ready solution to **parse Cursor IDE telemetry from SQLite state.vscdb**.
- It implements automatic path resolution for macOS and Linux, batched SQL queries for memory efficiency, and configurable time-based filtering via the `max_age_days` parameter.
- Raw database entries are transformed into standardized `AgentEvent` objects containing `ChatMessage` and `ToolUsage` data structures.
- All functionality is backed by concrete implementations in [`cursor_parser.py`](https://github.com/uber/ADR/blob/main/cursor_parser.py), [`agent_event_schema.py`](https://github.com/uber/ADR/blob/main/agent_event_schema.py), and supporting utility modules.

## Frequently Asked Questions

### Where does Cursor IDE store its telemetry database?

Cursor IDE stores its telemetry in a SQLite database named `state.vscdb`. On macOS, the file is located at `~/Library/Application Support/Cursor/User/globalStorage/state.vscdb`, while on Linux it resides at `~/.config/Cursor/User/globalStorage/state.vscdb`. The `CursorParser` automatically resolves these paths based on the host operating system.

### What data structure does the CursorParser return?

The `parse_all()` method returns a list of `AgentEvent` objects as defined in [`agent_event_schema.py`](https://github.com/uber/ADR/blob/main/agent_event_schema.py). Each `AgentEvent` contains a `session_id`, unified `timestamp`, and a `chat_history` list comprising `ChatMessage` instances that may include associated `ToolUsage` records.

### How does the parser handle large conversation histories?

The parser implements batched reading via `_iter_cursor_batches` to process SQLite rows in chunks rather than loading the entire database into memory. Additionally, it filters conversations older than the configured `max_age_days` (default 14 days) before processing individual bubbles, significantly reducing computational overhead.

### Can I customize the time window for parsing events?

Yes. The `CursorParser` constructor accepts an optional `max_age_days` parameter. Instantiating `CursorParser(max_age_days=7)` restricts processing to conversations updated within the last seven days, allowing targeted analysis of recent activity.