How to Extract Fulltext Content from Indexed Sources in NotebookLM

The NotebookLM Python client exposes indexed source content through client.sources.get_fulltext(), which returns a SourceFulltext dataclass containing the complete extracted text along with metadata such as title, URL, and character count.

The teng-lin/notebooklm-py library provides a programmatic interface to Google's NotebookLM service, enabling developers to extract and analyze raw text from any indexed source. Whether processing web pages, PDFs, or YouTube transcripts that have been added to a notebook, the client handles the underlying RPC communication and data parsing automatically. This guide demonstrates how to retrieve complete source text using both the asynchronous Python API and the command-line interface.

How Fulltext Extraction Works

When you call get_fulltext(), the client sends a GET_SOURCE RPC request to the NotebookLM backend. According to the source code in src/notebooklm/_sources.py, the method constructs a specific parameter array [[source_id], [2], [2]] and transmits it via RPCMethod.GET_SOURCE to the /notebook/<notebook_id> endpoint.

The RPC response contains nested arrays representing the source's title, type code, optional URL, and content blocks. The internal helper _extract_all_text() recursively walks these nested structures in src/notebooklm/_sources.py and concatenates the strings into a single content field. The result is then wrapped in the SourceFulltext dataclass defined in src/notebooklm/types.py, which provides convenient properties like .kind (translating internal type codes to the SourceType enum) and .char_count for the extracted text length.

Extracting Fulltext with the Python Client

The recommended approach uses the asynchronous NotebookLMClient to retrieve source content. The SourcesAPI.get_fulltext() coroutine handles authentication, request serialization, and response parsing automatically.

import asyncio
from notebooklm import NotebookLMClient

async def show_fulltext(notebook_id: str, source_id: str):
    # Initialise the client (loads stored credentials)

    async with await NotebookLMClient.from_storage() as client:
        # Retrieve the full‑text object

        fulltext = await client.sources.get_fulltext(notebook_id, source_id)

        print(f"Source ID   : {fulltext.source_id}")
        print(f"Title       : {fulltext.title}")
        print(f"URL         : {fulltext.url or '(no URL)'}")
        print(f"Type        : {fulltext.kind}")          # e.g. SourceType.WEB_PAGE

        print(f"Chars       : {fulltext.char_count}")
        print("\n--- Full Text ---------------------------------------------------\n")
        print(fulltext.content)

# Example usage

asyncio.run(show_fulltext("nb_123abc", "src_456def"))

This implementation hits the GET_SOURCE RPC path and returns a fully populated SourceFulltext object containing the complete indexed text together with its metadata.

Retrieving Fulltext via the Command Line

For scripting and automation workflows, the CLI provides direct access to the same functionality without writing Python code.


# Print the full text to stdout

notebooklm source fulltext src_456def -n nb_123abc

# Save the result as JSON (useful for downstream processing)

notebooklm source fulltext src_456def -n nb_123abc --json > src_456def_fulltext.json

The CLI command internally invokes the same SourcesAPI.get_fulltext() method, ensuring the output format exactly matches the SourceFulltext dataclass schema.

Processing SourceFulltext Metadata and Citations

The SourceFulltext dataclass in src/notebooklm/types.py provides more than just raw text. It exposes source_id, title, url, kind, and char_count fields for immediate programmatic access. Additionally, the class includes a find_citation_context() helper method that extracts surrounding text for a given citation string.


# Assume we already have a SourceFulltext object `ft`

matches = ft.find_citation_context("important claim", context_chars=150)
for context, pos in matches:
    print(f"Match at {pos}: {context}\n")

This utility, implemented in src/notebooklm/types.py, slices the content string around each occurrence of the search term, returning a list of (context, start_position) tuples. This functionality powers chat reference systems that need to ground responses in specific source locations.

Handling Empty or Missing Content

Sources that are still processing or contain no extractable text return an empty string for the content field. The client automatically logs a warning when this occurs, as implemented in src/notebooklm/_sources.py.

fulltext = await client.sources.get_fulltext(nb_id, src_id)
if not fulltext.content:
    print("⚠️ No indexed text – source may still be processing or empty.")
else:
    # Normal processing

    process_text(fulltext.content)

Always validate the content field before processing to handle indexing delays or malformed sources gracefully.

Summary

  • Primary method: Use client.sources.get_fulltext(notebook_id, source_id) to retrieve complete indexed text via the GET_SOURCE RPC call with parameters [[source_id], [2], [2]].
  • Return structure: The method returns a SourceFulltext dataclass containing content, title, url, kind, and char_count fields.
  • CLI availability: The same functionality is accessible via notebooklm source fulltext <source-id> -n <notebook-id>.
  • Citation support: Use find_citation_context() to locate specific passages within the extracted text.
  • Edge cases: Empty content triggers a warning log and returns an empty string for the content field.

Frequently Asked Questions

What data structure does get_fulltext() return?

The method returns a SourceFulltext dataclass defined in src/notebooklm/types.py. This object contains the complete extracted text in the content field, along with metadata including source_id, title, url, kind (as a SourceType enum), and char_count indicating the text length.

How does the client handle sources that haven't finished indexing?

If a source has no indexed content available, the internal logic in src/notebooklm/_sources.py logs a warning message and returns a SourceFulltext object with an empty content string. Your code should check if not fulltext.content: to detect this state and handle it appropriately, such as by retrying later or skipping the source.

Can I search for specific text within the extracted fulltext?

Yes. The SourceFulltext object provides a find_citation_context() method that searches the content string for specific passages. It returns surrounding context windows as (context, position) tuples, allowing you to locate citations or specific claims within the source material without manual string parsing.

What is the underlying RPC method used for extraction?

The client uses the GET_SOURCE RPC method (RPCMethod.GET_SOURCE) sent to the /notebook/<notebook_id> endpoint. The request payload contains the specific parameter array [[source_id], [2], [2]] as implemented in src/notebooklm/_sources.py. The response is parsed by the _extract_all_text() helper to reconstruct the complete text from nested array structures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →