# How to Extract Fulltext Content from Indexed Sources in NotebookLM

> Easily extract fulltext content from indexed sources in NotebookLM using the Python client. Get complete text, title, URL, and character count with client.sources.get_fulltext().

- Repository: [Teng Lin/notebooklm-py](https://github.com/teng-lin/notebooklm-py)
- Tags: how-to-guide
- Published: 2026-03-09

---

**The NotebookLM Python client exposes indexed source content through `client.sources.get_fulltext()`, which returns a `SourceFulltext` dataclass containing the complete extracted text along with metadata such as title, URL, and character count.**

The `teng-lin/notebooklm-py` library provides a programmatic interface to Google's NotebookLM service, enabling developers to extract and analyze raw text from any indexed source. Whether processing web pages, PDFs, or YouTube transcripts that have been added to a notebook, the client handles the underlying RPC communication and data parsing automatically. This guide demonstrates how to retrieve complete source text using both the asynchronous Python API and the command-line interface.

## How Fulltext Extraction Works

When you call `get_fulltext()`, the client sends a `GET_SOURCE` RPC request to the NotebookLM backend. According to the source code in [`src/notebooklm/_sources.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/_sources.py), the method constructs a specific parameter array `[[source_id], [2], [2]]` and transmits it via `RPCMethod.GET_SOURCE` to the `/notebook/<notebook_id>` endpoint.

The RPC response contains nested arrays representing the source's title, type code, optional URL, and content blocks. The internal helper `_extract_all_text()` recursively walks these nested structures in [`src/notebooklm/_sources.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/_sources.py) and concatenates the strings into a single `content` field. The result is then wrapped in the `SourceFulltext` dataclass defined in [`src/notebooklm/types.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/types.py), which provides convenient properties like `.kind` (translating internal type codes to the `SourceType` enum) and `.char_count` for the extracted text length.

## Extracting Fulltext with the Python Client

The recommended approach uses the asynchronous `NotebookLMClient` to retrieve source content. The `SourcesAPI.get_fulltext()` coroutine handles authentication, request serialization, and response parsing automatically.

```python
import asyncio
from notebooklm import NotebookLMClient

async def show_fulltext(notebook_id: str, source_id: str):
    # Initialise the client (loads stored credentials)

    async with await NotebookLMClient.from_storage() as client:
        # Retrieve the full‑text object

        fulltext = await client.sources.get_fulltext(notebook_id, source_id)

        print(f"Source ID   : {fulltext.source_id}")
        print(f"Title       : {fulltext.title}")
        print(f"URL         : {fulltext.url or '(no URL)'}")
        print(f"Type        : {fulltext.kind}")          # e.g. SourceType.WEB_PAGE

        print(f"Chars       : {fulltext.char_count}")
        print("\n--- Full Text ---------------------------------------------------\n")
        print(fulltext.content)

# Example usage

asyncio.run(show_fulltext("nb_123abc", "src_456def"))

```

This implementation hits the `GET_SOURCE` RPC path and returns a fully populated **SourceFulltext** object containing the complete indexed text together with its metadata.

## Retrieving Fulltext via the Command Line

For scripting and automation workflows, the CLI provides direct access to the same functionality without writing Python code.

```bash

# Print the full text to stdout

notebooklm source fulltext src_456def -n nb_123abc

# Save the result as JSON (useful for downstream processing)

notebooklm source fulltext src_456def -n nb_123abc --json > src_456def_fulltext.json

```

The CLI command internally invokes the same `SourcesAPI.get_fulltext()` method, ensuring the output format exactly matches the `SourceFulltext` dataclass schema.

## Processing SourceFulltext Metadata and Citations

The **SourceFulltext** dataclass in [`src/notebooklm/types.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/types.py) provides more than just raw text. It exposes `source_id`, `title`, `url`, `kind`, and `char_count` fields for immediate programmatic access. Additionally, the class includes a `find_citation_context()` helper method that extracts surrounding text for a given citation string.

```python

# Assume we already have a SourceFulltext object `ft`

matches = ft.find_citation_context("important claim", context_chars=150)
for context, pos in matches:
    print(f"Match at {pos}: {context}\n")

```

This utility, implemented in [`src/notebooklm/types.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/types.py), slices the `content` string around each occurrence of the search term, returning a list of `(context, start_position)` tuples. This functionality powers chat reference systems that need to ground responses in specific source locations.

## Handling Empty or Missing Content

Sources that are still processing or contain no extractable text return an empty string for the content field. The client automatically logs a warning when this occurs, as implemented in [`src/notebooklm/_sources.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/_sources.py).

```python
fulltext = await client.sources.get_fulltext(nb_id, src_id)
if not fulltext.content:
    print("⚠️ No indexed text – source may still be processing or empty.")
else:
    # Normal processing

    process_text(fulltext.content)

```

Always validate the `content` field before processing to handle indexing delays or malformed sources gracefully.

## Summary

- **Primary method**: Use `client.sources.get_fulltext(notebook_id, source_id)` to retrieve complete indexed text via the `GET_SOURCE` RPC call with parameters `[[source_id], [2], [2]]`.
- **Return structure**: The method returns a `SourceFulltext` dataclass containing `content`, `title`, `url`, `kind`, and `char_count` fields.
- **CLI availability**: The same functionality is accessible via `notebooklm source fulltext <source-id> -n <notebook-id>`.
- **Citation support**: Use `find_citation_context()` to locate specific passages within the extracted text.
- **Edge cases**: Empty content triggers a warning log and returns an empty string for the content field.

## Frequently Asked Questions

### What data structure does `get_fulltext()` return?

The method returns a **SourceFulltext** dataclass defined in [`src/notebooklm/types.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/types.py). This object contains the complete extracted text in the `content` field, along with metadata including `source_id`, `title`, `url`, `kind` (as a `SourceType` enum), and `char_count` indicating the text length.

### How does the client handle sources that haven't finished indexing?

If a source has no indexed content available, the internal logic in [`src/notebooklm/_sources.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/_sources.py) logs a warning message and returns a **SourceFulltext** object with an empty `content` string. Your code should check `if not fulltext.content:` to detect this state and handle it appropriately, such as by retrying later or skipping the source.

### Can I search for specific text within the extracted fulltext?

Yes. The **SourceFulltext** object provides a `find_citation_context()` method that searches the `content` string for specific passages. It returns surrounding context windows as `(context, position)` tuples, allowing you to locate citations or specific claims within the source material without manual string parsing.

### What is the underlying RPC method used for extraction?

The client uses the `GET_SOURCE` RPC method (`RPCMethod.GET_SOURCE`) sent to the `/notebook/<notebook_id>` endpoint. The request payload contains the specific parameter array `[[source_id], [2], [2]]` as implemented in [`src/notebooklm/_sources.py`](https://github.com/teng-lin/notebooklm-py/blob/main/src/notebooklm/_sources.py). The response is parsed by the `_extract_all_text()` helper to reconstruct the complete text from nested array structures.