How to Download Artifacts in Batch from NotebookLM (MP3, MP4, PDF, PNG, CSV)
Use the CLI command notebooklm download <type> --all <output-dir> or the Python method client.artifacts._download_urls_batch() to fetch all completed artifacts of a specific type in parallel.
The notebooklm-py library provides two robust methods to batch download AI-generated studio artifacts from Google NotebookLM. Whether you need audio briefings, video summaries, slide decks, infographics, or data tables, you can retrieve them programmatically via the command line or Python API.
CLI Batch Download Method
The fastest way to download artifacts in bulk is through the CLI interface implemented in src/notebooklm/cli/download.py. The _download_artifacts_generic function orchestrates the entire workflow, handling authentication, artifact enumeration, and parallel downloads.
Downloading All Artifacts by Type
Each artifact type maps to a specific command and file extension. Use the --all flag to fetch every completed artifact of that type:
# Audio (MP3/M4A) → ./audio/
notebooklm download audio --all ./audio/
# Video (MP4) → ./video/
notebooklm download video --all ./video/
# Slide decks (PDF) → ./slides/
notebooklm download slide-deck --all ./slides/
# Infographics (PNG) → ./infographics/
notebooklm download infographic --all ./infographics/
# Data tables (CSV) → ./data-tables/
notebooklm download data-table --all ./data-tables/
The CLI resolves the notebook context, validates authentication tokens, and delegates to the internal _download_urls_batch helper for the actual file transfer.
Filtering and Selection Options
Instead of downloading everything, you can target specific artifacts using flags processed by the CLI logic:
--all– Downloads every completed artifact of the specified type--latest– Fetches only the most recently generated artifact--earliest– Fetches the oldest artifact--name <pattern>– Filters artifacts by title substring match--artifact <id>– Downloads a specific artifact by its unique ID
The selection logic filters the list returned by _get_completed_artifacts_as_dicts before building the download queue.
Python API Batch Download Method
For programmatic workflows, the NotebookLMClient exposes the artifact namespace through src/notebooklm/_artifacts.py. The private helper _download_urls_batch (lines 1976–2014) handles concurrent downloads using httpx.AsyncClient with proper cookie handling and HTTPS validation against Google domains.
Listing and Preparing Artifacts
First, retrieve the artifact list and filter for completed items:
import asyncio
from pathlib import Path
from notebooklm import NotebookLMClient, ArtifactType
async def batch_download(notebook_id: str, out_dir: Path) -> None:
async with await NotebookLMClient.from_storage() as client:
# Define artifact types and their extensions
cfg = {
ArtifactType.AUDIO: ".mp3",
ArtifactType.VIDEO: ".mp4",
ArtifactType.SLIDE_DECK: ".pdf",
ArtifactType.INFOGRAPHIC: ".png",
ArtifactType.DATA_TABLE: ".csv",
}
for art_type, ext in cfg.items():
# List artifacts and filter for completed ones
artifacts = await client.artifacts.list(notebook_id, art_type)
completed = [a for a in artifacts if a.is_completed]
# Build (url, path) tuples
urls_paths = []
for art in completed:
url = art.download_url # Resolved from internal metadata
filename = f"{art.title or art.id}{ext}"
output_path = out_dir / art_type.name.lower() / filename
urls_paths.append((url, str(output_path)))
if urls_paths:
downloaded = await client.artifacts._download_urls_batch(urls_paths)
print(f"{art_type.name}: downloaded {len(downloaded)} files")
The download_url property extracts the correct field from the raw RPC payload (e.g., metadata[6][5] for audio, metadata[8] for video) without requiring manual parsing.
Executing Concurrent Downloads
The _download_urls_batch method accepts a list of tuples containing the download URL and local file path. It streams each file concurrently, validates the URL against a whitelist of Google domains, and writes files atomically to prevent corruption. This ensures identical behavior whether called from the CLI or your custom script.
Customizing Output Formats
For slide decks, you can specify alternative formats beyond the default PDF:
# Download as PowerPoint instead of PDF
notebooklm download slide-deck --format pptx --all ./slides/
The CLI adjusts the file extension to .pptx before calling client.artifacts.download_slide_deck, as implemented in lines 71–78 of src/notebooklm/cli/download.py.
Summary
- Two interfaces: Use the CLI for quick one-off downloads or the Python API for automated pipelines
- Five supported types: Audio (MP3/M4A), Video (MP4), Slide Decks (PDF/PPTX), Infographics (PNG), and Data Tables (CSV)
- Parallel processing: Both methods use
_download_urls_batchinsrc/notebooklm/_artifacts.pywithhttpx.AsyncClientfor concurrent downloads - Secure handling: URLs are validated against Google domains and cookies are managed securely throughout the session
- Flexible filtering: CLI supports
--all,--latest,--earliest,--name, and--artifactselection criteria
Frequently Asked Questions
What artifact types support batch downloading in NotebookLM?
The library supports batch downloads for five artifact types defined in src/notebooklm/rpc/types.py: AUDIO (MP3/M4A), VIDEO (MP4), SLIDE_DECK (PDF/PPTX), INFOGRAPHIC (PNG), and DATA_TABLE (CSV). Each type uses the same underlying _download_urls_batch helper but extracts download URLs from different metadata fields in the RPC response.
Is the batch download process concurrent?
Yes. Both the CLI and Python API use the _download_urls_batch method which initializes an httpx.AsyncClient to stream multiple files simultaneously. This parallel approach significantly reduces download time when retrieving dozens of artifacts compared to sequential requests.
Can I download artifacts from multiple notebooks in a single command?
No, the current implementation in notebooklm-py requires a specific notebook context for each batch operation. You must run separate commands or API calls for each notebook ID. However, you can script this by iterating over multiple notebook IDs in a shell loop or Python async gather operation.
How does authentication work for batch artifact downloads?
The CLI uses load_auth_from_storage and fetch_tokens to retrieve OAuth credentials, while the Python API requires NotebookLMClient.from_storage() to establish an authenticated session. Both methods store and refresh tokens automatically, passing the necessary cookies to the httpx.AsyncClient to authenticate each download request against Google's servers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →