How to Use the Research API with Auto-Import for Web and Drive Sources in notebooklm-py
Call client.research.start() with source="web" or source="drive", poll until completion with client.research.poll(), then auto-import discovered sources using client.research.import_sources().
The notebooklm-py library provides an async Python interface to Google NotebookLM's undocumented Research API. This guide explains how to programmatically research topics from web or Google Drive sources and automatically import discovered documents into your notebooks using the three-layer architecture implemented in the repository.
Understanding the Research API Architecture
The Research workflow is built on three distinct layers that handle authentication, transport, and business logic.
NotebookLMClient (src/notebooklm/client.py, lines 87-95) provides the high-level entry point. It initializes a persistent HTTP session and exposes the research sub-API via client.research.
Core/RPC Layer (src/notebooklm/rpc/types.py, lines 53-57) manages HTTP transport and translates method calls into Google's batchexecute RPC protocol. This layer defines the method identifiers: START_FAST_RESEARCH, START_DEEP_RESEARCH, POLL_RESEARCH, and IMPORT_RESEARCH.
ResearchAPI (src/notebooklm/_research.py) implements the three core primitives: start(), poll(), and import_sources(). These methods wrap RPC calls and normalize response formats for the caller.
Starting a Research Task
The start() method (lines 45-61 in _research.py) initiates background research for a given notebook.
The method signature accepts four parameters:
notebook_id– The target notebook identifier.query– The research query string.source– Either"web"(default) or"drive".mode– Either"fast"(default) or"deep"(web-only).
Validation logic at lines 75-80 enforces that Drive sources only support "fast" mode. Attempting to use "deep" with "drive" raises a ValueError immediately.
The method returns a dictionary containing task_id, report_id, and request metadata. The client automatically retries on auth expiration via refresh_auth (lines 44-86 in client.py).
Polling for Research Completion
After starting a task, you must poll until the background research finishes. The poll() method (lines 110-130 in _research.py) checks the status of active research for a given notebook.
The method returns a status string and discovered sources:
"completed"– Research finished successfully; the response includes thesourceslist."no_research"– The task vanished or was cancelled."in_progress"– The task is still running.
Each source in the returned list follows the format {"url": "...", "title": "..."}.
Auto-Importing Discovered Sources
Once polling returns "completed", use import_sources() (lines 195-210 in _research.py) to add discovered documents to your notebook.
This method accepts:
notebook_id– The target notebook.task_id– The research task identifier fromstart().sources– A list of source dictionaries (typically a subset of the poll response).
The method returns a list of imported source objects, each containing id and title fields.
Complete Implementation Example
The following async workflow demonstrates the full research-to-import pipeline for both web and Drive sources.
import asyncio
from notebooklm import NotebookLMClient
async def research_and_import(
notebook_id: str,
query: str,
source: str = "web", # "web" or "drive"
mode: str = "fast", # "fast" (default) or "deep" (web-only)
max_import: int = 5, # how many top sources to import automatically
):
# ------------------------------------------------------------------
# 1️⃣ Initialise the client (uses stored Playwright state)
# ------------------------------------------------------------------
async with await NotebookLMClient.from_storage() as client:
# ------------------------------------------------------------------
# 2️⃣ Start the research task
# ------------------------------------------------------------------
start_res = await client.research.start(
notebook_id, query, source=source, mode=mode
)
if not start_res:
raise RuntimeError("Failed to start research")
task_id = start_res["task_id"]
print(f"Research started – task_id={task_id}")
# ------------------------------------------------------------------
# 3️⃣ Poll until the task is completed
# ------------------------------------------------------------------
while True:
poll_res = await client.research.poll(notebook_id)
if poll_res.get("status") == "completed":
print("Research completed")
break
elif poll_res.get("status") == "no_research":
raise RuntimeError("Research task vanished")
else:
# In-progress – wait a bit before the next poll
await asyncio.sleep(5)
# ------------------------------------------------------------------
# 4️⃣ Auto-import the top *max_import* sources
# ------------------------------------------------------------------
sources_to_import = poll_res["sources"][:max_import]
imported = await client.research.import_sources(
notebook_id, task_id, sources_to_import
)
print(f"Imported {len(imported)} sources:")
for src in imported:
print(f" • {src['id']} – {src['title']}")
return imported
# ----------------------------------------------------------------------
# Example usage
# ----------------------------------------------------------------------
if __name__ == "__main__":
# Replace with a real notebook ID you own
NB_ID = "abc123def456"
asyncio.run(
research_and_import(
notebook_id=NB_ID,
query="latest advances in quantum computing",
source="web", # or "drive"
mode="fast", # "deep" works for web only
max_import=5,
)
)
For Drive sources, simply set source="drive" and keep mode="fast" to avoid validation errors.
Summary
- Instantiation: Use
NotebookLMClient.from_storage()to create an authenticated client with persistent sessions. - Research Initiation: Call
client.research.start()withsource="web"or"drive"andmode="fast"or"deep". - Status Monitoring: Poll
client.research.poll()in a loop untilstatusequals"completed". - Auto-Import: Pass the discovered sources list to
client.research.import_sources()to add them to your notebook. - Async-Only: All methods are coroutines and must be awaited; the client handles token refresh automatically.
Frequently Asked Questions
What is the difference between "fast" and "deep" research modes?
"Fast" mode performs a quick search and returns results rapidly, while "deep" mode conducts more thorough research but only works with web sources. According to the validation logic in src/notebooklm/_research.py lines 75-80, setting source="drive" with mode="deep" raises a ValueError because the underlying Google API does not support deep research for Drive sources.
Can I import all discovered sources or only a subset?
You can import any subset of the sources returned by the polling operation. The import_sources() method accepts a list of source dictionaries, allowing you to slice the results (e.g., poll_res["sources"][:5]) or filter them programmatically before import. Each source dictionary must contain url and title keys.
How does the client handle authentication retries during research?
The NotebookLMClient automatically refreshes expired authentication tokens via the refresh_auth method (lines 44-86 in src/notebooklm/client.py). If an RPC call fails due to auth expiration, the client renews the session and retries the request transparently without raising exceptions to the caller.
Why does the Research API use RPC instead of REST?
The underlying Google NotebookLM service uses an undocumented batchexecute RPC protocol. The notebooklm-py library abstracts this complexity by mapping Python method calls to RPC method IDs defined in src/notebooklm/rpc/types.py (lines 53-57), including START_FAST_RESEARCH and IMPORT_RESEARCH, while exposing a clean async Python interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →