How EPUB Extraction Works Using ebooklib in main.py: A Complete Guide

The extract_epub_to_txt function in main.py converts EPUB files to plain text by leveraging ebooklib to parse book structure and BeautifulSoup to strip HTML markup from chapter content.

The qiaomu-anything-to-notebooklm repository automates content ingestion for NotebookLM analysis across multiple file formats. For EPUB e-books, the tool implements a dedicated extraction pipeline that isolates readable text from the HTML-based chapter documents found within the EPUB archive. This implementation resides in the extract_epub_to_txt helper function, located at lines 50-70 of main.py.

The EPUB Extraction Pipeline

The extraction process follows a five-stage pipeline that transforms structured EPUB documents into clean, analysis-ready text files. Each stage handles specific aspects of the EPUB specification while maintaining compatibility with the NotebookLM upload workflow.

Step 1: Import Required Libraries

The function begins by importing the core dependencies needed for EPUB parsing and HTML processing:

import ebooklib
from ebooklib import epub
from bs4 import BeautifulSoup

These imports provide the epub module for reading book archives and BeautifulSoup for HTML-to-text conversion. Both dependencies are listed in requirements.txt and must be installed via pip.

Step 2: Load the EPUB Book Structure

The function uses epub.read_epub() to parse the EPUB file into an EpubBook object containing all items, metadata, and navigation data:

book = epub.read_epub(str(epub_path))

This method returns a structured representation where individual chapters, images, stylesheets, and fonts are accessible as discrete items within the archive.

Step 3: Filter for Document Items

EPUB archives contain multiple item types, but only HTML/XHTML documents hold readable text. The code iterates through all items and filters for ITEM_DOCUMENT types:

for item in book.get_items():
    if item.get_type() == ebooklib.ITEM_DOCUMENT:
        # Process chapter...

This check ensures the extractor processes only actual chapter content while ignoring images, audio files, CSS stylesheets, and font resources that would corrupt text output.

Step 4: Convert HTML to Plain Text

For each document item, the function extracts raw HTML bytes and passes them to BeautifulSoup for parsing:

soup = BeautifulSoup(item.get_content(), 'html.parser')
content.append(soup.get_text())

The get_text() method recursively extracts all text nodes while stripping HTML tags, navigation scripts, and inline styles, returning only human-readable content.

Step 5: Write to Temporary File

Finally, the concatenated chapter texts are written to a unique temporary file with UTF-8 encoding:

txt_path = tempfile.mktemp(suffix='.txt', prefix='epub_')
with open(txt_path, 'w', encoding='utf-8') as f:
    f.write('\n\n'.join(content))

The double line break separator (\n\n) visually distinguishes between chapters in the output file, while tempfile.mktemp() ensures no filename collisions occur during batch processing.

Complete Implementation Reference

Here is the complete extract_epub_to_txt function as implemented in main.py:

def extract_epub_to_txt(epub_path):
    import tempfile
    
    book = epub.read_epub(str(epub_path))
    content = []
    
    for item in book.get_items():
        if item.get_type() == ebooklib.ITEM_DOCUMENT:
            soup = BeautifulSoup(item.get_content(), 'html.parser')
            content.append(soup.get_text())
    
    txt_path = tempfile.mktemp(suffix='.txt', prefix='epub_')
    with open(txt_path, 'w', encoding='utf-8') as f:
        f.write('\n\n'.join(content))
    
    return txt_path

This function returns the absolute path to the generated text file, which the main script then passes to the NotebookLM upload logic when input_type == 'epub' is detected in the CLI flow.

Practical Usage Examples

Standalone Extraction

You can import and use the function directly for one-off conversions without triggering the full NotebookLM pipeline:

from pathlib import Path
from main import extract_epub_to_txt

epub_file = Path('sample_book.epub')
txt_file = extract_epub_to_txt(epub_file)

print(f'Generated plain-text file: {txt_file}')

# Output: /tmp/epub_xxx123.txt

Integrated CLI Workflow

When invoked through the command-line interface, the extraction happens automatically before upload:

python3 main.py my_book.epub --deep-analysis

The script performs the following sequence: detects the EPUB format via detect_input_type, calls extract_epub_to_txt to generate the temporary text file, uploads the content to NotebookLM, and initiates the three-round deep analysis questionnaire on the extracted text.

Custom NLP Pipeline Integration

For projects requiring raw text processing, reuse the extraction function to feed content into NLP workflows:

text_path = extract_epub_to_txt('/path/to/another_book.epub')
with open(text_path, encoding='utf-8') as f:
    raw_text = f.read()
    

# Process raw_text with spaCy, NLTK, or transformers

This pattern allows the ebooklib-based extraction to serve as a preprocessing step for custom text analysis pipelines while maintaining the temporary file cleanup responsibilities within your application logic.

Summary

  • extract_epub_to_txt in main.py handles all EPUB-to-text conversion using ebooklib for archive parsing and BeautifulSoup for HTML stripping.
  • The function filters for ebooklib.ITEM_DOCUMENT types to process only HTML chapters while ignoring images, fonts, and stylesheets.
  • Text extraction relies on BeautifulSoup.get_text() to remove HTML markup and scripts, returning clean plain text suitable for language models.
  • Output files are created using tempfile.mktemp() with UTF-8 encoding and double line breaks between chapters to preserve document structure.
  • The temporary file path is returned to the caller for subsequent NotebookLM upload or integration with custom text processing workflows.

Frequently Asked Questions

What Python dependencies are required for EPUB extraction?

The implementation requires ebooklib for EPUB archive parsing and beautifulsoup4 for HTML-to-text conversion. According to the requirements.txt in the repository, these must be installed via pip before running the extraction. The tempfile module is part of the Python standard library and requires no additional installation.

How does the extractor handle images and multimedia in EPUB files?

The extract_epub_to_txt function specifically checks for ebooklib.ITEM_DOCUMENT types, which filters out images (ITEM_IMAGE), audio files (ITEM_AUDIO), and stylesheets (ITEM_STYLE). Only HTML/XHTML chapter documents are processed, ensuring the output contains only readable text content without binary media data or CSS markup.

Can I modify the extraction to preserve HTML formatting instead of plain text?

Yes, you can modify the function to preserve HTML by removing the .get_text() call and instead using str(soup) or item.get_content().decode('utf-8') directly. However, the current implementation is optimized for NotebookLM consumption, which requires plain text input. Preserving HTML would require updates to the downstream processing logic as well, as NotebookLM performs better on clean text without markup.

Where are the extracted text files stored on the system?

The function uses tempfile.mktemp(suffix='.txt', prefix='epub_') to generate unique filenames in the system's default temporary directory. On Linux systems, this typically resolves to /tmp/epub_xxx123.txt, while macOS uses /var/folders/... and Windows uses %TEMP%. The files persist until explicitly deleted or until the operating system cleans temporary files according to its retention policy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →