EPUB Extraction Fallbacks in Book-to-Skill: A Complete Technical Guide

Book-to-Skill uses a multi-layered fallback system for EPUB extraction that prioritizes spine-based parsing, falls back to manifest traversal, and applies encoding detection and HTML sanitization to recover readable text from malformed or non-standard EPUB files.

The Book-to-Skill repository provides a robust pipeline for converting book content into structured skill data. When processing EPUB files—the standard open e-book format—the tool implements several fallback mechanisms to handle real-world variability in publisher implementations, corrupted archives, and non-compliant markup.

Primary Extraction Strategy: Spine-Based Parsing

The canonical path for EPUB content lives in book_to_skill/parsers/epub.py [source]. The parser's first-line approach leverages the EPUB spine, a mandatory element in the content.opf package document that declares the intended reading order.

The extraction flow proceeds as follows:

  • ZIP archive access via Python's zipfile module (wrapped by utilities in book_to_skill/dependencies.py) opens the .epub in-memory without disk extraction
  • OPF parsing locates the spine element and builds an ordered list of manifest item IDs
  • Chapter instantiation creates Chapter objects (defined in book_to_skill/utils.py) with metadata extracted from XHTML <title> tags or inferred from heading structures

This spine-first approach respects author intent and guarantees sequence correctness.

First Fallback: Manifest Traversal for Missing Spine References

When the spine is incomplete, empty, or references non-existent files, the parser falls back to manifest enumeration:

  • The parser walks all manifest items with media-type="application/xhtml+xml" or text/html
  • Items are processed in document order (the order they appear in the OPF manifest)
  • Duplicate detection prevents double-processing items referenced multiple times in a fragmented spine

This fallback ensures extraction continues even when publishers supply malformed package documents.

Second Fallback: Raw Archive Enumeration

For severely corrupted EPUBs where content.opf cannot be parsed or located, the parser implements a raw ZIP fallback:


# Conceptual fallback flow in epub.py

import zipfile
from pathlib import Path

def extract_with_fallbacks(epub_path: Path):
    """
    Demonstrates the fallback hierarchy Book-to-Skill employs.
    """
    with zipfile.ZipFile(epub_path, 'r') as zf:
        # Fallback 3: Raw archive scanning

        xhtml_candidates = [
            name for name in zf.namelist()
            if name.endswith(('.xhtml', '.html', '.htm'))
            and not name.startswith('__MACOSX')  # macOS metadata skip

        ]
        # Sort by path depth to approximate reading order

        xhtml_candidates.sort(key=lambda p: (p.count('/'), p))
        return xhtml_candidates

The raw enumeration filters out common noise (__MACOSX directories, Thumbs.db, metadata files) and sorts by path depth as a heuristic for chapter ordering.

Encoding Fallbacks: From Declared to Detected

EPUB content files declare encoding in XML prologs or <meta> tags. When declarations are missing or incorrect, the parser cascades through:

Priority Method Implementation Location
1 XML/HTML declared encoding book_to_skill/parsers/html.py
2 HTTP-EQUIV meta charset html.py extraction regex
3 UTF-8 assumption with BOM strip epub.py byte preprocessing
4 chardet heuristics dependencies.py wrapper

The chardet integration (optional dependency) provides statistical encoding detection for legacy EPUBs using non-UTF encodings like Windows-1252 or ISO-8859-1.

HTML Parsing Fallbacks: Structured to Tolerant

Extracted XHTML/HTML passes through book_to_skill/parsers/html.py, which implements its own fallback ladder:

  • BeautifulSoup with lxml parser (fast, strict) for well-formed documents
  • BeautifulSoup with html.parser (lenient) for malformed markup
  • Regex-based tag stripping as last resort for severely broken HTML

The sanitizer in book_to_skill/sanitize.py then applies:

  • Unicode control character removal (bidirectional overrides, private use areas)
  • Whitespace normalization
  • Boilerplate heuristic detection (publisher footers, copyright blocks)

Asset Handling Fallbacks and Missing Image Reporting

EPUBs embed images referenced via relative paths. The parser in epub.py tracks asset resolution with degrading strategies:


# Asset resolution with fallbacks

image_path = chapter_relative_path  # Primary: spine-declared base

if image_path not in zip_namelist:
    image_path = image_path.lstrip('/')  # Fallback: absolute-relative mismatch

if image_path not in zip_namelist:
    image_path = image_path.replace('%20', ' ')  # URL encoding fix

# Final: report unresolvable references

missing_images.append(original_reference)

The test suite tests/test_epub_image_reporting.py validates that missing images are surfaced to callers rather than silently dropped, enabling downstream consumers to insert placeholders or warnings.

CLI Integration and Parser Selection

The command-line interface in book_to_skill/cli.py routes EPUB files through the parser registry (parsers/__init__.py). The detection logic:


# Extension-based parser selection

$ book-to-skill extract ./manual.epub --output ./skills/

# Internally dispatches to EpubParser via registry mapping

The registry pattern allows runtime inspection of available parsers and extension-based fallback chains (e.g., treating .epub3 identically to .epub).

Usage Examples

Programmatic extraction with error tolerance:

from book_to_skill.parsers.epub import EpubParser
from pathlib import Path

epub = Path("legacy_book.epub")
parser = EpubParser(epub)

chapters = list(parser.iter_chapters())
print(f"Extracted {len(chapters)} chapters via fallback chain")

# Inspect what failed

if parser.missing_images:
    print(f"Warning: {len(parser.missing_images)} unresolvable images")

CLI extraction with verbose logging:

$ book-to-skill extract ./corrupted.epub --output ./output --verbose

# Logs indicate: "Spine incomplete, falling back to manifest enumeration"

# Logs indicate: "Chapter 5: declared encoding invalid, using chardet"

Summary

  • Spine-first extraction is the primary strategy in book_to_skill/parsers/epub.py
  • Manifest traversal handles incomplete or broken spines
  • Raw ZIP enumeration recovers content from severely corrupted EPUBs
  • Encoding detection cascades from declared → UTF-8 → chardet heuristics
  • HTML parsing degrades from lxmlhtml.parser → regex stripping
  • Missing asset reporting ensures transparency via test_epub_image_reporting.py validation
  • All fallbacks operate in-memory without temporary file extraction

Frequently Asked Questions

What happens if an EPUB has no content.opf file?

The parser falls back to raw ZIP archive enumeration, scanning for .xhtml, .html, and .htm files while filtering metadata directories like __MACOSX. Files are sorted by path depth to approximate reading order. This handles non-standard EPUBs or renamed ZIP archives encountered in the wild.

How does Book-to-Skill handle EPUBs with incorrect character encodings?

The parser first attempts the declared encoding from XML prologs or meta tags. If decoding fails, it strips any byte-order mark and assumes UTF-8. As a final fallback, it invokes chardet statistical detection for legacy encodings like Windows-1252 or ISO-8859-1.

Can the tool extract text from DRM-protected EPUBs?

No. Book-to-Skill operates on standard ZIP-based EPUB archives and does not implement DRM removal. Attempting to process encrypted EPUBs will trigger a ZIP decryption error early in the pipeline, surfacing a clear failure message rather than silent corruption.

Where is the test coverage for fallback behavior?

The tests/test_epub_image_reporting.py module validates asset resolution and missing image reporting. Additional integration tests exercise malformed spine and encoding fallback paths, ensuring the extraction pipeline degrades gracefully across edge cases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →