How to Use the Research API with Auto-Import for Web and Drive Sources in notebooklm-py

Call client.research.start() with source="web" or source="drive", poll until completion with client.research.poll(), then auto-import discovered sources using client.research.import_sources().

The notebooklm-py library provides an async Python interface to Google NotebookLM's undocumented Research API. This guide explains how to programmatically research topics from web or Google Drive sources and automatically import discovered documents into your notebooks using the three-layer architecture implemented in the repository.

Understanding the Research API Architecture

The Research workflow is built on three distinct layers that handle authentication, transport, and business logic.

NotebookLMClient (src/notebooklm/client.py, lines 87-95) provides the high-level entry point. It initializes a persistent HTTP session and exposes the research sub-API via client.research.

Core/RPC Layer (src/notebooklm/rpc/types.py, lines 53-57) manages HTTP transport and translates method calls into Google's batchexecute RPC protocol. This layer defines the method identifiers: START_FAST_RESEARCH, START_DEEP_RESEARCH, POLL_RESEARCH, and IMPORT_RESEARCH.

ResearchAPI (src/notebooklm/_research.py) implements the three core primitives: start(), poll(), and import_sources(). These methods wrap RPC calls and normalize response formats for the caller.

Starting a Research Task

The start() method (lines 45-61 in _research.py) initiates background research for a given notebook.

The method signature accepts four parameters:

  • notebook_id – The target notebook identifier.
  • query – The research query string.
  • source – Either "web" (default) or "drive".
  • mode – Either "fast" (default) or "deep" (web-only).

Validation logic at lines 75-80 enforces that Drive sources only support "fast" mode. Attempting to use "deep" with "drive" raises a ValueError immediately.

The method returns a dictionary containing task_id, report_id, and request metadata. The client automatically retries on auth expiration via refresh_auth (lines 44-86 in client.py).

Polling for Research Completion

After starting a task, you must poll until the background research finishes. The poll() method (lines 110-130 in _research.py) checks the status of active research for a given notebook.

The method returns a status string and discovered sources:

  • "completed" – Research finished successfully; the response includes the sources list.
  • "no_research" – The task vanished or was cancelled.
  • "in_progress" – The task is still running.

Each source in the returned list follows the format {"url": "...", "title": "..."}.

Auto-Importing Discovered Sources

Once polling returns "completed", use import_sources() (lines 195-210 in _research.py) to add discovered documents to your notebook.

This method accepts:

  • notebook_id – The target notebook.
  • task_id – The research task identifier from start().
  • sources – A list of source dictionaries (typically a subset of the poll response).

The method returns a list of imported source objects, each containing id and title fields.

Complete Implementation Example

The following async workflow demonstrates the full research-to-import pipeline for both web and Drive sources.

import asyncio
from notebooklm import NotebookLMClient

async def research_and_import(
    notebook_id: str,
    query: str,
    source: str = "web",   # "web" or "drive"

    mode: str = "fast",    # "fast" (default) or "deep" (web-only)

    max_import: int = 5,   # how many top sources to import automatically

):
    # ------------------------------------------------------------------

    # 1️⃣ Initialise the client (uses stored Playwright state)

    # ------------------------------------------------------------------

    async with await NotebookLMClient.from_storage() as client:
        # ------------------------------------------------------------------

        # 2️⃣ Start the research task

        # ------------------------------------------------------------------

        start_res = await client.research.start(
            notebook_id, query, source=source, mode=mode
        )
        if not start_res:
            raise RuntimeError("Failed to start research")
        task_id = start_res["task_id"]
        print(f"Research started – task_id={task_id}")

        # ------------------------------------------------------------------

        # 3️⃣ Poll until the task is completed

        # ------------------------------------------------------------------

        while True:
            poll_res = await client.research.poll(notebook_id)
            if poll_res.get("status") == "completed":
                print("Research completed")
                break
            elif poll_res.get("status") == "no_research":
                raise RuntimeError("Research task vanished")
            else:
                # In-progress – wait a bit before the next poll

                await asyncio.sleep(5)

        # ------------------------------------------------------------------

        # 4️⃣ Auto-import the top *max_import* sources

        # ------------------------------------------------------------------

        sources_to_import = poll_res["sources"][:max_import]
        imported = await client.research.import_sources(
            notebook_id, task_id, sources_to_import
        )
        print(f"Imported {len(imported)} sources:")
        for src in imported:
            print(f"  • {src['id']} – {src['title']}")
        return imported

# ----------------------------------------------------------------------

# Example usage

# ----------------------------------------------------------------------

if __name__ == "__main__":
    # Replace with a real notebook ID you own

    NB_ID = "abc123def456"
    asyncio.run(
        research_and_import(
            notebook_id=NB_ID,
            query="latest advances in quantum computing",
            source="web",   # or "drive"

            mode="fast",    # "deep" works for web only

            max_import=5,
        )
    )

For Drive sources, simply set source="drive" and keep mode="fast" to avoid validation errors.

Summary

  • Instantiation: Use NotebookLMClient.from_storage() to create an authenticated client with persistent sessions.
  • Research Initiation: Call client.research.start() with source="web" or "drive" and mode="fast" or "deep".
  • Status Monitoring: Poll client.research.poll() in a loop until status equals "completed".
  • Auto-Import: Pass the discovered sources list to client.research.import_sources() to add them to your notebook.
  • Async-Only: All methods are coroutines and must be awaited; the client handles token refresh automatically.

Frequently Asked Questions

What is the difference between "fast" and "deep" research modes?

"Fast" mode performs a quick search and returns results rapidly, while "deep" mode conducts more thorough research but only works with web sources. According to the validation logic in src/notebooklm/_research.py lines 75-80, setting source="drive" with mode="deep" raises a ValueError because the underlying Google API does not support deep research for Drive sources.

Can I import all discovered sources or only a subset?

You can import any subset of the sources returned by the polling operation. The import_sources() method accepts a list of source dictionaries, allowing you to slice the results (e.g., poll_res["sources"][:5]) or filter them programmatically before import. Each source dictionary must contain url and title keys.

How does the client handle authentication retries during research?

The NotebookLMClient automatically refreshes expired authentication tokens via the refresh_auth method (lines 44-86 in src/notebooklm/client.py). If an RPC call fails due to auth expiration, the client renews the session and retries the request transparently without raising exceptions to the caller.

Why does the Research API use RPC instead of REST?

The underlying Google NotebookLM service uses an undocumented batchexecute RPC protocol. The notebooklm-py library abstracts this complexity by mapping Python method calls to RPC method IDs defined in src/notebooklm/rpc/types.py (lines 53-57), including START_FAST_RESEARCH and IMPORT_RESEARCH, while exposing a clean async Python interface.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →