How Distilly Parses SRT and VTT Subtitle Files: A Complete Guide

Distilly parses SRT and VTT subtitle files through a lightweight, regex-based pipeline that strips headers, removes timestamps and HTML markup, deduplicates lines, and formats the output into readable paragraphs.

Distilly is an open-source tool developed by titanwings that converts raw subtitle files into clean, readable transcripts. The parsing logic resides in tools/research/srt_to_transcript.py and handles both SubRip (.srt) and WebVTT (.vtt) formats through a unified, format-agnostic approach.

Overview of the Subtitle Parsing Pipeline

The conversion process operates in six distinct stages, implemented primarily in the clean_subtitle_text function. This design intentionally avoids external subtitle parsing libraries, relying instead on regular expressions and string operations to process files line by line.

  1. Header stripping (VTT-specific)
  2. Line filtering (empty lines, indices, timestamps)
  3. Markup cleanup (HTML tags, style cues)
  4. Deduplication (consecutive identical lines)
  5. Paragraph formation (sentence boundaries or character limits)
  6. File output (via convert_file)

Step-by-Step Parsing Logic

Header Stripping for VTT Files

When processing WebVTT files, Distilly first removes the mandatory WEBVTT header and any NOTE blocks using the _strip_vtt_headers function. This regex-based operation ensures that metadata and comments do not appear in the final transcript.

According to the source code in tools/research/srt_to_transcript.py (lines 16-20), the implementation uses pattern matching to identify and discard these structural elements before processing the actual subtitle content.

Line Filtering and Timestamp Removal

The parser examines each line individually, discarding three specific categories:

  • Empty lines that provide no semantic content
  • Numeric indices such as "1", "2", "3" that SRT files use to sequence subtitle entries
  • Timestamp lines matching the TIMESTAMP_PATTERN (e.g., 00:00:01.000 --> 00:00:03.000)

As implemented in lines 31-35 of tools/research/srt_to_transcript.py, the TIMESTAMP_PATTERN identifies both standard SRT timestamps and WebVTT variants, ensuring cross-format compatibility.

Markup Cleanup and Style Cues

Remaining lines undergo sanitization to remove presentation markup. The parser strips HTML-like tags (<i>, <b>, <u>, etc.) and WebVTT-specific style cues such as align:start or position:0%. Inline whitespace is collapsed to single spaces to normalize formatting.

This cleanup occurs in lines 36-39 of the core implementation file, ensuring that only the actual spoken dialogue remains.

Deduplication and Paragraph Formation

Distilly applies two final text-processing stages to improve readability:

Deduplication: Consecutive identical lines are collapsed into a single occurrence. This prevents repetition when subtitle files contain duplicate entries for emphasis or timing purposes (lines 42-46).

Paragraph formation: Cleaned lines accumulate into paragraphs that emit when either:

  • The paragraph reaches approximately 240 characters, or
  • The line ends with sentence-terminating punctuation (., !, ?, or CJK equivalents)

The final buffer flushes as a terminal paragraph, with paragraphs separated by blank lines (lines 48-56).

Key Implementation Files and Functions

The parsing functionality resides in a minimal, self-contained module:

  • tools/research/srt_to_transcript.py: Contains the core clean_subtitle_text function for string processing and convert_file for file I/O operations
  • tests/test_research_tools.py: Validates that timestamps, HTML tags, and duplicates are correctly removed (lines 19-36)
  • bin/distilly.mjs: CLI entry point that forwards arguments to the Python implementation

Usage Examples

Programmatic Usage in Python

Import the functions directly to process subtitles in memory or convert files:

from pathlib import Path
from research.srt_to_transcript import clean_subtitle_text, convert_file

# Process subtitle text directly

raw_content = Path("lecture.vtt").read_text(encoding="utf-8")
transcript = clean_subtitle_text(raw_content)
print(transcript)

# Convert file to transcript

convert_file(Path("lecture.srt"), Path("lecture_transcript.txt"))

Command-Line Interface

Invoke the parser via the Distilly CLI or Python module:


# Using the Distilly CLI

./bin/distilly.mjs tools/research/srt_to_transcript.py input.vtt output.txt

# Direct Python execution

python -m research.srt_to_transcript input.srt

# Creates input_transcript.txt in the same directory

Summary

  • Distilly parses SRT and VTT subtitle files using a format-agnostic regex pipeline that requires no external parsing libraries
  • The process strips WEBVTT headers, timestamps, HTML markup, and WebVTT style cues while collapsing duplicate lines
  • Output formats into readable paragraphs based on sentence boundaries or a 240-character threshold
  • Core functions clean_subtitle_text and convert_file reside in tools/research/srt_to_transcript.py
  • Unit tests in tests/test_research_tools.py verify parsing accuracy and file conversion

Frequently Asked Questions

Does Distilly require external libraries to parse subtitle files?

No. According to the titanwings/distilly source code, the implementation relies solely on Python's standard library regular expressions and string operations. The clean_subtitle_text function processes files through pattern matching rather than importing specialized subtitle parsing dependencies.

How does Distilly handle HTML tags in subtitle files?

Distilly removes HTML-like markup such as <i>, <b>, and <u> tags during the markup cleanup phase. This occurs in tools/research/srt_to_transcript.py alongside the removal of WebVTT-specific style cues like align:start and position:0%, ensuring only plain text dialogue remains in the output.

What is the difference between clean_subtitle_text and convert_file?

The clean_subtitle_text function accepts a string containing raw subtitle content and returns a cleaned transcript string. The convert_file function wraps this logic with file I/O operations—it reads the source subtitle file, calls clean_subtitle_text, and writes the result to a destination path, typically named <original>_transcript.txt.

Can Distilly parse subtitle formats other than SRT and VTT?

The current implementation specifically targets SRT and WebVTT formats through timestamp pattern matching and VTT header detection. While the regex-based approach might partially handle similar text-based formats, the tool does not officially support ASS, SSA, or other advanced subtitle formats.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →