# How EPUB Extraction Works Using ebooklib in main.py: A Complete Guide

> Learn how to extract EPUB to TXT using ebooklib in main.py. This guide details parsing book structure and stripping HTML for plain text conversion.

- Repository: [向阳乔木/qiaomu-anything-to-notebooklm](https://github.com/joeseesun/qiaomu-anything-to-notebooklm)
- Tags: how-to-guide
- Published: 2026-05-16

---

**The `extract_epub_to_txt` function in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py) converts EPUB files to plain text by leveraging `ebooklib` to parse book structure and BeautifulSoup to strip HTML markup from chapter content.**

The `qiaomu-anything-to-notebooklm` repository automates content ingestion for NotebookLM analysis across multiple file formats. For EPUB e-books, the tool implements a dedicated extraction pipeline that isolates readable text from the HTML-based chapter documents found within the EPUB archive. This implementation resides in the `extract_epub_to_txt` helper function, located at lines 50-70 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py).

## The EPUB Extraction Pipeline

The extraction process follows a five-stage pipeline that transforms structured EPUB documents into clean, analysis-ready text files. Each stage handles specific aspects of the EPUB specification while maintaining compatibility with the NotebookLM upload workflow.

### Step 1: Import Required Libraries

The function begins by importing the core dependencies needed for EPUB parsing and HTML processing:

```python
import ebooklib
from ebooklib import epub
from bs4 import BeautifulSoup

```

These imports provide the `epub` module for reading book archives and `BeautifulSoup` for HTML-to-text conversion. Both dependencies are listed in [`requirements.txt`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/requirements.txt) and must be installed via pip.

### Step 2: Load the EPUB Book Structure

The function uses `epub.read_epub()` to parse the EPUB file into an `EpubBook` object containing all items, metadata, and navigation data:

```python
book = epub.read_epub(str(epub_path))

```

This method returns a structured representation where individual chapters, images, stylesheets, and fonts are accessible as discrete items within the archive.

### Step 3: Filter for Document Items

EPUB archives contain multiple item types, but only HTML/XHTML documents hold readable text. The code iterates through all items and filters for `ITEM_DOCUMENT` types:

```python
for item in book.get_items():
    if item.get_type() == ebooklib.ITEM_DOCUMENT:
        # Process chapter...

```

This check ensures the extractor processes only actual chapter content while ignoring images, audio files, CSS stylesheets, and font resources that would corrupt text output.

### Step 4: Convert HTML to Plain Text

For each document item, the function extracts raw HTML bytes and passes them to BeautifulSoup for parsing:

```python
soup = BeautifulSoup(item.get_content(), 'html.parser')
content.append(soup.get_text())

```

The `get_text()` method recursively extracts all text nodes while stripping HTML tags, navigation scripts, and inline styles, returning only human-readable content.

### Step 5: Write to Temporary File

Finally, the concatenated chapter texts are written to a unique temporary file with UTF-8 encoding:

```python
txt_path = tempfile.mktemp(suffix='.txt', prefix='epub_')
with open(txt_path, 'w', encoding='utf-8') as f:
    f.write('\n\n'.join(content))

```

The double line break separator (`\n\n`) visually distinguishes between chapters in the output file, while `tempfile.mktemp()` ensures no filename collisions occur during batch processing.

## Complete Implementation Reference

Here is the complete `extract_epub_to_txt` function as implemented in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py):

```python
def extract_epub_to_txt(epub_path):
    import tempfile
    
    book = epub.read_epub(str(epub_path))
    content = []
    
    for item in book.get_items():
        if item.get_type() == ebooklib.ITEM_DOCUMENT:
            soup = BeautifulSoup(item.get_content(), 'html.parser')
            content.append(soup.get_text())
    
    txt_path = tempfile.mktemp(suffix='.txt', prefix='epub_')
    with open(txt_path, 'w', encoding='utf-8') as f:
        f.write('\n\n'.join(content))
    
    return txt_path

```

This function returns the absolute path to the generated text file, which the main script then passes to the NotebookLM upload logic when `input_type == 'epub'` is detected in the CLI flow.

## Practical Usage Examples

### Standalone Extraction

You can import and use the function directly for one-off conversions without triggering the full NotebookLM pipeline:

```python
from pathlib import Path
from main import extract_epub_to_txt

epub_file = Path('sample_book.epub')
txt_file = extract_epub_to_txt(epub_file)

print(f'Generated plain-text file: {txt_file}')

# Output: /tmp/epub_xxx123.txt

```

### Integrated CLI Workflow

When invoked through the command-line interface, the extraction happens automatically before upload:

```bash
python3 main.py my_book.epub --deep-analysis

```

The script performs the following sequence: detects the EPUB format via `detect_input_type`, calls `extract_epub_to_txt` to generate the temporary text file, uploads the content to NotebookLM, and initiates the three-round deep analysis questionnaire on the extracted text.

### Custom NLP Pipeline Integration

For projects requiring raw text processing, reuse the extraction function to feed content into NLP workflows:

```python
text_path = extract_epub_to_txt('/path/to/another_book.epub')
with open(text_path, encoding='utf-8') as f:
    raw_text = f.read()
    

# Process raw_text with spaCy, NLTK, or transformers

```

This pattern allows the `ebooklib`-based extraction to serve as a preprocessing step for custom text analysis pipelines while maintaining the temporary file cleanup responsibilities within your application logic.

## Summary

- **`extract_epub_to_txt`** in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py) handles all EPUB-to-text conversion using `ebooklib` for archive parsing and BeautifulSoup for HTML stripping.
- The function filters for `ebooklib.ITEM_DOCUMENT` types to process only HTML chapters while ignoring images, fonts, and stylesheets.
- Text extraction relies on `BeautifulSoup.get_text()` to remove HTML markup and scripts, returning clean plain text suitable for language models.
- Output files are created using `tempfile.mktemp()` with UTF-8 encoding and double line breaks between chapters to preserve document structure.
- The temporary file path is returned to the caller for subsequent NotebookLM upload or integration with custom text processing workflows.

## Frequently Asked Questions

### What Python dependencies are required for EPUB extraction?

The implementation requires `ebooklib` for EPUB archive parsing and `beautifulsoup4` for HTML-to-text conversion. According to the [`requirements.txt`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/requirements.txt) in the repository, these must be installed via pip before running the extraction. The `tempfile` module is part of the Python standard library and requires no additional installation.

### How does the extractor handle images and multimedia in EPUB files?

The `extract_epub_to_txt` function specifically checks for `ebooklib.ITEM_DOCUMENT` types, which filters out images (`ITEM_IMAGE`), audio files (`ITEM_AUDIO`), and stylesheets (`ITEM_STYLE`). Only HTML/XHTML chapter documents are processed, ensuring the output contains only readable text content without binary media data or CSS markup.

### Can I modify the extraction to preserve HTML formatting instead of plain text?

Yes, you can modify the function to preserve HTML by removing the `.get_text()` call and instead using `str(soup)` or `item.get_content().decode('utf-8')` directly. However, the current implementation is optimized for NotebookLM consumption, which requires plain text input. Preserving HTML would require updates to the downstream processing logic as well, as NotebookLM performs better on clean text without markup.

### Where are the extracted text files stored on the system?

The function uses `tempfile.mktemp(suffix='.txt', prefix='epub_')` to generate unique filenames in the system's default temporary directory. On Linux systems, this typically resolves to [`/tmp/epub_xxx123.txt`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main//tmp/epub_xxx123.txt), while macOS uses `/var/folders/...` and Windows uses `%TEMP%`. The files persist until explicitly deleted or until the operating system cleans temporary files according to its retention policy.