How to Create Custom youtube-dl Post-Processors: A Complete Guide
youtube-dl post-processors are objects that transform downloaded files after completion by implementing the run() method in the PostProcessor base class and registering them via YoutubeDL.add_post_processor().
youtube-dl is a command-line program to download videos from YouTube and other sites. After a video file finishes downloading, youtube-dl post-processors handle transformations like audio extraction, format conversion, metadata embedding, and thumbnail insertion. Understanding how to create custom post-processors allows you to extend youtube-dl with specialized file-level processing logic tailored to your workflow.
What Are youtube-dl Post-Processors?
The PostProcessor Base Class in common.py
The foundation of the post-processing system lives in youtube_dl/postprocessor/common.py. This file defines the abstract base class PostProcessor that all post-processors must inherit from.
The critical method to implement is run(self, information):
- The
informationargument is a dictionary containing download metadata, including afilepathkey pointing to the downloaded file. - The method must return a tuple
(files_to_delete, updated_information). - Raise
PostProcessingErroron failure.
The default implementation in PostProcessor.run() returns an empty delete list and the untouched information dictionary, effectively performing no operation.
# https://github.com/ytdl-org/youtube-dl/blob/master/youtube_dl/postprocessor/common.py
class PostProcessor(object):
def run(self, information):
"""
The "information" argument is a dict like those produced by InfoExtractors.
It contains a ``filepath`` entry that points to the downloaded file.
Return a tuple (files_to_delete, updated_information) or raise
PostProcessingError on failure.
"""
return [], information
How Post-Processors Integrate with the Downloader
The downloader class in youtube_dl/YoutubeDL.py manages the post-processor chain through two key mechanisms:
-
Registration: The
add_post_processor(self, pp)method appends aPostProcessorinstance toself._ppsand sets the downloader reference on the PP. -
Execution: The
post_process(self, filename, ie_info)method constructs the full processing chain by combining:- Extractor-specific post-processors from
ie_info['__postprocessors'] - Global post-processors from
self._pps
- Extractor-specific post-processors from
The execution loop calls pp.run(info) for each processor. If a post-processor returns files in its delete list and the user has not specified -k/--keep-video, youtube-dl automatically removes those files.
# https://github.com/ytdl-org/youtube-dl/blob/master/youtube_dl/YoutubeDL.py
def add_post_processor(self, pp):
"""Add a PostProcessor object to the end of the chain."""
self._pps.append(pp)
pp.set_downloader(self)
def post_process(self, filename, ie_info):
"""Run all the postprocessors on the given file."""
info = dict(ie_info)
info['filepath'] = filename
pps_chain = []
if ie_info.get('__postprocessors') is not None:
pps_chain.extend(ie_info['__postprocessors'])
pps_chain.extend(self._pps)
for pp in pps_chain:
files_to_delete, info = pp.run(info)
if files_to_delete and not self.params.get('keepvideo', False):
for old_filename in files_to_delete:
os.remove(encodeFilename(old_filename))
Built-in youtube-dl Post-Processors
The youtube_dl/postprocessor/ directory contains numerous concrete implementations. Key examples include:
| File | Class | Typical Role |
|---|---|---|
ffmpeg.py |
FFmpegExtractAudioPP |
Extract audio with FFmpeg |
embedthumbnail.py |
EmbedThumbnailPP |
Embed thumbnail into MP4/MKV |
metadatafromtitle.py |
MetadataFromTitlePP |
Parse title-based metadata |
execafterdownload.py |
ExecAfterDownloadPP |
Run an external command after download |
These classes inherit from PostProcessor (or from FFmpegPostProcessor, which itself derives from PostProcessor) and override run() with the actual processing logic.
Creating a Custom youtube-dl Post-Processor
Building a custom post-processor requires three steps: subclassing the base class, implementing the run() method, and registering the instance with the downloader.
Step 1: Subclass PostProcessor
Create a new class that inherits from PostProcessor (or FFmpegPostProcessor if you need FFmpeg integration):
from youtube_dl.postprocessor.common import PostProcessor
class MyCustomPP(PostProcessor):
def run(self, information):
# Implementation here
return [], information
Step 2: Implement the run() Method
The run() method receives the information dictionary containing filepath and other metadata. Process the file as needed, then return the tuple structure.
Here is a complete example creating a preview clip using FFmpeg:
from youtube_dl.postprocessor.ffmpeg import FFmpegPostProcessor, FFmpegPostProcessorError
import os
class PreviewClipPP(FFmpegPostProcessor):
def run(self, info):
in_file = info['filepath']
out_file = os.path.splitext(in_file)[0] + '-preview.mp4'
cmd = [
self.get_command(),
'-ss', '00:00:00', '-t', '00:00:05',
'-i', in_file,
'-c', 'copy',
out_file
]
self.run_ffmpeg_multiple(cmd)
info['preview_file'] = out_file
return [], info
For simpler file operations without FFmpeg, use standard Python libraries:
import shutil
import os
from youtube_dl.postprocessor.common import PostProcessor
class CopyRenamePP(PostProcessor):
def run(self, info):
src = info['filepath']
dst = src + '.copy'
shutil.copy2(src, dst)
info['copied_path'] = dst
return [], info
Step 3: Register Your Post-Processor
Register the post-processor with the YoutubeDL instance before downloading:
from youtube_dl import YoutubeDL
from my_pp_module import PreviewClipPP
ydl = YoutubeDL({'format': 'best'})
ydl.add_post_processor(PreviewClipPP())
ydl.download(['https://youtu.be/dQw4w9WgXcQ'])
The add_post_processor() method appends your PP to the internal chain and sets the downloader reference, allowing your PP to access downloader parameters via self._downloader.
Testing Your Custom Post-Processor
The official test suite in test/test_YoutubeDL.py provides patterns for validating post-processors. The SimplePP class demonstrates the minimal skeleton:
class SimplePP(PostProcessor):
def run(self, info):
# Create a dummy side-file
with open('audio_file.txt', 'w') as f:
f.write('EXAMPLE')
# Delete the original video unless keep-video is set
return [info['filepath']], info
When testing:
- Create a temporary video file
- Instantiate your PP and call
ydl.add_post_processor() - Run
ydl.post_process()with the temporary file - Assert that side files exist and original files are removed when
keepvideo=False
This pattern ensures your custom logic integrates correctly with youtube-dl's file lifecycle management.
Summary
- youtube-dl post-processors are objects that transform downloaded files after completion, defined in
youtube_dl/postprocessor/common.py. - The base
PostProcessorclass requires implementingrun(self, information)which returns(files_to_delete, updated_info). - Registration happens via
YoutubeDL.add_post_processor()inyoutube_dl/YoutubeDL.py, which stores PPs inself._pps. - Built-in processors in
youtube_dl/postprocessor/handle audio extraction, thumbnail embedding, and metadata parsing. - Custom PPs can subclass
PostProcessorfor simple operations orFFmpegPostProcessorfor FFmpeg integration, then register programmatically with the downloader instance.
Frequently Asked Questions
What is the difference between PostProcessor and FFmpegPostProcessor?
PostProcessor is the abstract base class in youtube_dl/postprocessor/common.py that defines the run() interface for all post-processors. FFmpegPostProcessor is a specialized subclass in youtube_dl/postprocessor/ffmpeg.py that provides helper methods like get_command() and run_ffmpeg_multiple() for executing FFmpeg operations. Use FFmpegPostProcessor when your custom processor needs to invoke FFmpeg, and the base PostProcessor for pure Python file operations.
How do I prevent youtube-dl from deleting the original file after post-processing?
Set the keepvideo parameter to True when instantiating YoutubeDL, or pass the -k or --keep-video flag from the command line. When keepvideo is enabled, youtube-dl ignores the files_to_delete list returned by post-processors and retains the original downloaded file. This logic is handled in the deletion loop inside post_process() in youtube_dl/YoutubeDL.py.
Can I use multiple post-processors on the same download?
Yes. The YoutubeDL class maintains a list self._pps that stores all registered post-processors. When you call add_post_processor(), your PP is appended to this list. During post_process(), youtube-dl executes each processor sequentially in the order they were added, passing the updated information dictionary from one PP to the next. You can also combine extractor-specific PPs from ie_info['__postprocessors'] with global PPs.
Where should I place my custom post-processor code for youtube-dl to find it?
For programmatic usage, place your custom post-processor module anywhere in your Python path and import it into your script that instantiates YoutubeDL. If you are extending youtube-dl itself, place your new post-processor file in the youtube_dl/postprocessor/ directory and import your class in youtube_dl/postprocessor/__init__.py to make it available for CLI integration. The test suite in test/test_YoutubeDL.py demonstrates how to test PPs without modifying the core library.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →