# How to Use the Research API with Auto-Import for Web and Drive Sources in notebooklm-py

> Learn to use the Research API with auto-import for web and Drive sources in notebooklm-py. Start, poll, and import sources efficiently with simple Python code.

- Repository: [Teng Lin/notebooklm-py](https://github.com/teng-lin/notebooklm-py)
- Tags: how-to-guide
- Published: 2026-03-09

---

**Call `client.research.start()` with `source="web"` or `source="drive"`, poll until completion with `client.research.poll()`, then auto-import discovered sources using `client.research.import_sources()`.**

The `notebooklm-py` library provides an async Python interface to Google NotebookLM's undocumented Research API. This guide explains how to programmatically research topics from web or Google Drive sources and automatically import discovered documents into your notebooks using the three-layer architecture implemented in the repository.

## Understanding the Research API Architecture

The Research workflow is built on three distinct layers that handle authentication, transport, and business logic.

**NotebookLMClient** ([`src/notebooklm/client.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/client.py), lines 87-95) provides the high-level entry point. It initializes a persistent HTTP session and exposes the `research` sub-API via `client.research`.

**Core/RPC Layer** ([`src/notebooklm/rpc/types.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/rpc/types.py), lines 53-57) manages HTTP transport and translates method calls into Google's `batchexecute` RPC protocol. This layer defines the method identifiers: `START_FAST_RESEARCH`, `START_DEEP_RESEARCH`, `POLL_RESEARCH`, and `IMPORT_RESEARCH`.

**ResearchAPI** ([`src/notebooklm/_research.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/_research.py)) implements the three core primitives: `start()`, `poll()`, and `import_sources()`. These methods wrap RPC calls and normalize response formats for the caller.

## Starting a Research Task

The `start()` method (lines 45-61 in [`_research.py`](https://github.com/teng-lin/notebooklm-py/blob/main/_research.py)) initiates background research for a given notebook.

The method signature accepts four parameters:

- **`notebook_id`** – The target notebook identifier.
- **`query`** – The research query string.
- **`source`** – Either `"web"` (default) or `"drive"`.
- **`mode`** – Either `"fast"` (default) or `"deep"` (web-only).

**Validation logic** at lines 75-80 enforces that Drive sources only support `"fast"` mode. Attempting to use `"deep"` with `"drive"` raises a `ValueError` immediately.

The method returns a dictionary containing `task_id`, `report_id`, and request metadata. The client automatically retries on auth expiration via `refresh_auth` (lines 44-86 in [`client.py`](https://github.com/teng-lin/notebooklm-py/blob/main/client.py)).

## Polling for Research Completion

After starting a task, you must poll until the background research finishes. The `poll()` method (lines 110-130 in [`_research.py`](https://github.com/teng-lin/notebooklm-py/blob/main/_research.py)) checks the status of active research for a given notebook.

The method returns a status string and discovered sources:

- **`"completed"`** – Research finished successfully; the response includes the `sources` list.
- **`"no_research"`** – The task vanished or was cancelled.
- **`"in_progress"`** – The task is still running.

Each source in the returned list follows the format `{"url": "...", "title": "..."}`.

## Auto-Importing Discovered Sources

Once polling returns `"completed"`, use `import_sources()` (lines 195-210 in [`_research.py`](https://github.com/teng-lin/notebooklm-py/blob/main/_research.py)) to add discovered documents to your notebook.

This method accepts:

- **`notebook_id`** – The target notebook.
- **`task_id`** – The research task identifier from `start()`.
- **`sources`** – A list of source dictionaries (typically a subset of the poll response).

The method returns a list of imported source objects, each containing `id` and `title` fields.

## Complete Implementation Example

The following async workflow demonstrates the full research-to-import pipeline for both web and Drive sources.

```python
import asyncio
from notebooklm import NotebookLMClient

async def research_and_import(
    notebook_id: str,
    query: str,
    source: str = "web",   # "web" or "drive"

    mode: str = "fast",    # "fast" (default) or "deep" (web-only)

    max_import: int = 5,   # how many top sources to import automatically

):
    # ------------------------------------------------------------------

    # 1️⃣ Initialise the client (uses stored Playwright state)

    # ------------------------------------------------------------------

    async with await NotebookLMClient.from_storage() as client:
        # ------------------------------------------------------------------

        # 2️⃣ Start the research task

        # ------------------------------------------------------------------

        start_res = await client.research.start(
            notebook_id, query, source=source, mode=mode
        )
        if not start_res:
            raise RuntimeError("Failed to start research")
        task_id = start_res["task_id"]
        print(f"Research started – task_id={task_id}")

        # ------------------------------------------------------------------

        # 3️⃣ Poll until the task is completed

        # ------------------------------------------------------------------

        while True:
            poll_res = await client.research.poll(notebook_id)
            if poll_res.get("status") == "completed":
                print("Research completed")
                break
            elif poll_res.get("status") == "no_research":
                raise RuntimeError("Research task vanished")
            else:
                # In-progress – wait a bit before the next poll

                await asyncio.sleep(5)

        # ------------------------------------------------------------------

        # 4️⃣ Auto-import the top *max_import* sources

        # ------------------------------------------------------------------

        sources_to_import = poll_res["sources"][:max_import]
        imported = await client.research.import_sources(
            notebook_id, task_id, sources_to_import
        )
        print(f"Imported {len(imported)} sources:")
        for src in imported:
            print(f"  • {src['id']} – {src['title']}")
        return imported

# ----------------------------------------------------------------------

# Example usage

# ----------------------------------------------------------------------

if __name__ == "__main__":
    # Replace with a real notebook ID you own

    NB_ID = "abc123def456"
    asyncio.run(
        research_and_import(
            notebook_id=NB_ID,
            query="latest advances in quantum computing",
            source="web",   # or "drive"

            mode="fast",    # "deep" works for web only

            max_import=5,
        )
    )

```

For Drive sources, simply set `source="drive"` and keep `mode="fast"` to avoid validation errors.

## Summary

- **Instantiation**: Use `NotebookLMClient.from_storage()` to create an authenticated client with persistent sessions.
- **Research Initiation**: Call `client.research.start()` with `source="web"` or `"drive"` and `mode="fast"` or `"deep"`.
- **Status Monitoring**: Poll `client.research.poll()` in a loop until `status` equals `"completed"`.
- **Auto-Import**: Pass the discovered sources list to `client.research.import_sources()` to add them to your notebook.
- **Async-Only**: All methods are coroutines and must be awaited; the client handles token refresh automatically.

## Frequently Asked Questions

### What is the difference between "fast" and "deep" research modes?

**"Fast" mode** performs a quick search and returns results rapidly, while **"deep" mode** conducts more thorough research but only works with web sources. According to the validation logic in [`src/notebooklm/_research.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/_research.py) lines 75-80, setting `source="drive"` with `mode="deep"` raises a `ValueError` because the underlying Google API does not support deep research for Drive sources.

### Can I import all discovered sources or only a subset?

You can import any subset of the sources returned by the polling operation. The `import_sources()` method accepts a list of source dictionaries, allowing you to slice the results (e.g., `poll_res["sources"][:5]`) or filter them programmatically before import. Each source dictionary must contain `url` and `title` keys.

### How does the client handle authentication retries during research?

The `NotebookLMClient` automatically refreshes expired authentication tokens via the `refresh_auth` method (lines 44-86 in [`src/notebooklm/client.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/client.py)). If an RPC call fails due to auth expiration, the client renews the session and retries the request transparently without raising exceptions to the caller.

### Why does the Research API use RPC instead of REST?

The underlying Google NotebookLM service uses an undocumented `batchexecute` RPC protocol. The `notebooklm-py` library abstracts this complexity by mapping Python method calls to RPC method IDs defined in [`src/notebooklm/rpc/types.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/rpc/types.py) (lines 53-57), including `START_FAST_RESEARCH` and `IMPORT_RESEARCH`, while exposing a clean async Python interface.