ReClip Architecture Explained: A Flask-Based YouTube Downloader Service
ReClip is a lightweight Flask web service that wraps yt-dlp to provide a REST API for video metadata extraction, playlist parsing, and media downloads through four architectural layers: request handling, business logic, process management, and storage delivery.
The ReClip architecture implemented in averygan/reclip offers a thin, container-friendly alternative to heavyweight download solutions. By combining Flask's HTTP capabilities with yt-dlp's extraction power, it delivers a stateless API suitable for ephemeral deployments without external database dependencies.
Four-Layer ReClip Architecture
The codebase organizes functionality into distinct layers, each with clear responsibilities and source file locations.
Request Handling Layer
Flask serves as the HTTP entry point, accepting JSON-encoded POST and GET requests.
- Flask application initialization:
app = Flask(__name__)inapp.pylines 7–15 - Route decorators:
@app.routedefines three main endpoints:/api/info,/api/playlist, and/api/download
This layer deserializes incoming JSON, validates required fields (url, optional format), and dispatches to business logic functions.
Business Logic Layer
Core operations reside in app.py lines 97–140, implementing three primary functions:
| Function | Purpose | yt-dlp Command Equivalent |
|---|---|---|
get_info |
Extract video metadata | yt-dlp -j <url> |
get_playlist_info |
List playlist entries | yt-dlp --flat-playlist -J <url> |
run_download |
Execute media download | yt-dlp -f <format> -o <template> <url> |
The parse_ytdlp_json helper processes yt-dlp's JSON output, selecting highest-bitrate formats per resolution for the formats array returned to clients.
Process Management Layer
ReClip handles long-running downloads through subprocess execution and background threading:
# From app.py lines 16-33 and 77-84
subprocess.run(
cmd,
capture_output=True,
text=True,
timeout=300
)
Key implementation details:
- Job tracking: Global
jobsdictionary stores in-memory state using UUID-basedjob_idkeys - Threading:
threading.Thread(daemon=True)spawns downloads without blocking the Flask worker - Status values:
downloading,done, orerrorwith descriptive messages
The timeout mechanism (300 seconds default) prevents hung processes from consuming resources indefinitely.
Storage and Delivery Layer
Downloaded media flows through a controlled lifecycle in app.py lines 34–75 and 99–105:
- Destination:
DOWNLOAD_DIRconstant (default:downloads/) - Filename sanitization: Special characters removed to prevent path traversal
- Cleanup: Post-download removal of yt-dlp's intermediate
.partfiles - Delivery: Flask's
send_filestreams the final artifact with original filename preserved
No persistent storage beyond the filesystem—aligning with container-native deployment patterns.
ReClip Request Flow: Step by Step
Understanding how requests traverse the architecture clarifies design decisions.
Metadata Extraction Flow
POST /api/info → get_info() → subprocess.run(yt-dlp -j)
→ parse_ytdlp_json() → JSON response with title, thumbnail, formats
The parse_ytdlp_json function specifically handles yt-dlp's verbose format objects, extracting height, vbr (video bitrate), abr (audio bitrate), and ext fields for client-side format selection.
Download Execution Flow
POST /api/download → start_download() → jobs[job_id] = {"status": "downloading"}
→ threading.Thread(target=run_download).start() → 202 Accepted with job_id
Clients then poll GET /api/status/<job_id> until status transitions to done, finally retrieving via GET /api/file/<job_id>.
Complete ReClip API Usage Example
import requests
import time
BASE_URL = "http://localhost:8899"
# 1. Inspect available formats
info_resp = requests.post(
f"{BASE_URL}/api/info",
json={"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"}
)
info = info_resp.json()
print(f"Title: {info['title']}")
print(f"Available formats: {len(info['formats'])}")
# 2. Initiate download (video+audio merged)
download_resp = requests.post(
f"{BASE_URL}/api/download",
json={
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"format": "best" # or specific format_id from info['formats']
}
)
job_id = download_resp.json()["job_id"]
# 3. Poll for completion
while True:
status_resp = requests.get(f"{BASE_URL}/api/status/{job_id}")
status = status_resp.json()
if status["status"] == "done":
break
elif status["status"] == "error":
raise RuntimeError(f"Download failed: {status.get('error')}")
time.sleep(1)
# 4. Retrieve final file
file_resp = requests.get(f"{BASE_URL}/api/file/{job_id}")
with open(status["filename"], "wb") as f:
f.write(file_resp.content)
Key ReClip Source Files
| File | Lines of Code | Architectural Role |
|---|---|---|
app.py |
~140 | Complete server implementation: routing, logic, subprocess management |
templates/index.html |
~50 | Browser-based UI demonstrating API consumption |
requirements.txt |
2 | Flask dependency declaration |
Dockerfile |
12 | Multi-stage build with yt-dlp binary installation |
docker-compose.yml |
8 | Port 8899 exposure and volume mount |
The entire service fits in a single Python file with no external service dependencies—a deliberate simplicity choice for the ReClip architecture.
Summary
-
ReClip implements a four-layer architecture: Flask request handling, yt-dlp business logic, threaded subprocess management, and filesystem storage delivery.
-
All state lives in the
jobsdictionary—no database required, enabling horizontal scaling with load balancer sticky sessions or external job stores if needed. -
Source code in
app.pydemonstrates production patterns: timeout handling, input sanitization, daemon threads, and streaming file responses. -
The API design follows REST conventions with asynchronous job polling for long-running downloads, returning HTTP 202 on initiation and 200 on completion.
Frequently Asked Questions
What dependencies does ReClip require?
ReClip requires Flask (Python web framework) and yt-dlp (standalone binary). The requirements.txt lists only Flask; yt-dlp is installed via curl in the Dockerfile or must exist on the host PATH. No database, message queue, or cache systems are needed.
How does ReClip handle concurrent downloads?
Each download spawns an independent threading.Thread marked as daemon=True. Threads execute subprocess.run calls to yt-dlp, with status tracked in the global jobs dictionary. For production scale, deploy multiple ReClip containers behind a load balancer rather than increasing thread count.
Can ReClip download entire playlists?
Yes. The /api/playlist endpoint accepts playlist URLs and returns flat video entries using yt-dlp's --flat-playlist -J flags. Individual downloads must then be initiated separately via /api/download for each video URL in the playlist response.
Is ReClip suitable for production use?
The ReClip architecture targets containerized, short-lived deployments. In-memory job state means progress is lost on restart. For production, consider: externalizing jobs to Redis, adding authentication middleware, implementing rate limiting, and mounting DOWNLOAD_DIR to persistent or object storage rather than local filesystem.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →