# How Distilly Parses SRT and VTT Subtitle Files: A Complete Guide

> Distilly parses SRT and VTT subtitle files with a regex pipeline. Learn how it strips headers, removes timestamps, deduplicates lines, and formats readable paragraphs.

- Repository: [Tianyi Zhou/distilly](https://github.com/titanwings/distilly)
- Tags: how-to-guide
- Published: 2026-09-10

---

**Distilly parses SRT and VTT subtitle files through a lightweight, regex-based pipeline that strips headers, removes timestamps and HTML markup, deduplicates lines, and formats the output into readable paragraphs.**

Distilly is an open-source tool developed by titanwings that converts raw subtitle files into clean, readable transcripts. The parsing logic resides in [`tools/research/srt_to_transcript.py`](https://github.com/titanwings/distilly/blob/main/tools/research/srt_to_transcript.py) and handles both SubRip (`.srt`) and WebVTT (`.vtt`) formats through a unified, format-agnostic approach.

## Overview of the Subtitle Parsing Pipeline

The conversion process operates in six distinct stages, implemented primarily in the `clean_subtitle_text` function. This design intentionally avoids external subtitle parsing libraries, relying instead on regular expressions and string operations to process files line by line.

1. **Header stripping** (VTT-specific)
2. **Line filtering** (empty lines, indices, timestamps)
3. **Markup cleanup** (HTML tags, style cues)
4. **Deduplication** (consecutive identical lines)
5. **Paragraph formation** (sentence boundaries or character limits)
6. **File output** (via `convert_file`)

## Step-by-Step Parsing Logic

### Header Stripping for VTT Files

When processing WebVTT files, Distilly first removes the mandatory `WEBVTT` header and any `NOTE` blocks using the `_strip_vtt_headers` function. This regex-based operation ensures that metadata and comments do not appear in the final transcript.

According to the source code in [`tools/research/srt_to_transcript.py`](https://github.com/titanwings/distilly/blob/main/tools/research/srt_to_transcript.py) (lines 16-20), the implementation uses pattern matching to identify and discard these structural elements before processing the actual subtitle content.

### Line Filtering and Timestamp Removal

The parser examines each line individually, discarding three specific categories:

- **Empty lines** that provide no semantic content
- **Numeric indices** such as "1", "2", "3" that SRT files use to sequence subtitle entries
- **Timestamp lines** matching the `TIMESTAMP_PATTERN` (e.g., `00:00:01.000 --> 00:00:03.000`)

As implemented in lines 31-35 of [`tools/research/srt_to_transcript.py`](https://github.com/titanwings/distilly/blob/main/tools/research/srt_to_transcript.py), the `TIMESTAMP_PATTERN` identifies both standard SRT timestamps and WebVTT variants, ensuring cross-format compatibility.

### Markup Cleanup and Style Cues

Remaining lines undergo sanitization to remove presentation markup. The parser strips HTML-like tags (`<i>`, `<b>`, `<u>`, etc.) and WebVTT-specific style cues such as `align:start` or `position:0%`. Inline whitespace is collapsed to single spaces to normalize formatting.

This cleanup occurs in lines 36-39 of the core implementation file, ensuring that only the actual spoken dialogue remains.

### Deduplication and Paragraph Formation

Distilly applies two final text-processing stages to improve readability:

**Deduplication**: Consecutive identical lines are collapsed into a single occurrence. This prevents repetition when subtitle files contain duplicate entries for emphasis or timing purposes (lines 42-46).

**Paragraph formation**: Cleaned lines accumulate into paragraphs that emit when either:
- The paragraph reaches approximately 240 characters, **or**
- The line ends with sentence-terminating punctuation (`.`, `!`, `?`, or CJK equivalents)

The final buffer flushes as a terminal paragraph, with paragraphs separated by blank lines (lines 48-56).

## Key Implementation Files and Functions

The parsing functionality resides in a minimal, self-contained module:

- **[`tools/research/srt_to_transcript.py`](https://github.com/titanwings/distilly/blob/main/tools/research/srt_to_transcript.py)**: Contains the core `clean_subtitle_text` function for string processing and `convert_file` for file I/O operations
- **[`tests/test_research_tools.py`](https://github.com/titanwings/distilly/blob/main/tests/test_research_tools.py)**: Validates that timestamps, HTML tags, and duplicates are correctly removed (lines 19-36)
- **`bin/distilly.mjs`**: CLI entry point that forwards arguments to the Python implementation

## Usage Examples

### Programmatic Usage in Python

Import the functions directly to process subtitles in memory or convert files:

```python
from pathlib import Path
from research.srt_to_transcript import clean_subtitle_text, convert_file

# Process subtitle text directly

raw_content = Path("lecture.vtt").read_text(encoding="utf-8")
transcript = clean_subtitle_text(raw_content)
print(transcript)

# Convert file to transcript

convert_file(Path("lecture.srt"), Path("lecture_transcript.txt"))

```

### Command-Line Interface

Invoke the parser via the Distilly CLI or Python module:

```bash

# Using the Distilly CLI

./bin/distilly.mjs tools/research/srt_to_transcript.py input.vtt output.txt

# Direct Python execution

python -m research.srt_to_transcript input.srt

# Creates input_transcript.txt in the same directory

```

## Summary

- Distilly parses SRT and VTT subtitle files using a **format-agnostic regex pipeline** that requires no external parsing libraries
- The process strips **WEBVTT headers**, **timestamps**, **HTML markup**, and **WebVTT style cues** while collapsing duplicate lines
- Output formats into **readable paragraphs** based on sentence boundaries or a 240-character threshold
- Core functions `clean_subtitle_text` and `convert_file` reside in [`tools/research/srt_to_transcript.py`](https://github.com/titanwings/distilly/blob/main/tools/research/srt_to_transcript.py)
- Unit tests in [`tests/test_research_tools.py`](https://github.com/titanwings/distilly/blob/main/tests/test_research_tools.py) verify parsing accuracy and file conversion

## Frequently Asked Questions

### Does Distilly require external libraries to parse subtitle files?

No. According to the titanwings/distilly source code, the implementation relies solely on Python's standard library regular expressions and string operations. The `clean_subtitle_text` function processes files through pattern matching rather than importing specialized subtitle parsing dependencies.

### How does Distilly handle HTML tags in subtitle files?

Distilly removes HTML-like markup such as `<i>`, `<b>`, and `<u>` tags during the markup cleanup phase. This occurs in [`tools/research/srt_to_transcript.py`](https://github.com/titanwings/distilly/blob/main/tools/research/srt_to_transcript.py) alongside the removal of WebVTT-specific style cues like `align:start` and `position:0%`, ensuring only plain text dialogue remains in the output.

### What is the difference between clean_subtitle_text and convert_file?

The `clean_subtitle_text` function accepts a string containing raw subtitle content and returns a cleaned transcript string. The `convert_file` function wraps this logic with file I/O operations—it reads the source subtitle file, calls `clean_subtitle_text`, and writes the result to a destination path, typically named `<original>_transcript.txt`.

### Can Distilly parse subtitle formats other than SRT and VTT?

The current implementation specifically targets SRT and WebVTT formats through timestamp pattern matching and VTT header detection. While the regex-based approach might partially handle similar text-based formats, the tool does not officially support ASS, SSA, or other advanced subtitle formats.