# How ReClip Manages Download Jobs: In-Memory Queue Architecture

> Discover how ReClip manages download jobs with an in-memory queue and daemon threads. Learn how it avoids blocking Flask requests and tracks job progress asynchronously.

- Repository: [Avery Gan/reclip](https://github.com/averygan/reclip)
- Tags: architecture
- Published: 2026-09-05

---

**ReClip handles video download jobs using an in-memory dictionary registry and daemon threads that execute yt-dlp commands asynchronously, preventing Flask request blocking while tracking job states from "downloading" to "done" or "error".**

ReClip is a lightweight Flask-based video downloader that processes media requests without blocking the main web thread. Understanding how ReClip manages download jobs reveals a thread-based architecture that balances simplicity with performance for personal or small-scale deployments. The entire job management system is contained within the single-file Flask application in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py).

## Job Registry and State Storage

### Global In-Memory Dictionary

At the core of ReClip's job management is a global `jobs` dictionary defined at line 13 in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py):

```python
jobs = {}

```

This registry stores every active and completed download job using randomly generated UUIDs as keys. Each job entry maintains a record containing the `status`, source `url`, video `title`, output `file` path, and sanitized `filename` for delivery.

### Job Lifecycle States

Jobs transition through a simple finite state machine. When created, a job receives an initial status of `"downloading"`. Upon completion of the background worker, the status updates to either `"done"` (success) or `"error"` (failure), with the latter populating an `error` key containing the exception or subprocess details.

## Initiating Download Jobs

### The Download Endpoint

Clients initiate downloads by POSTing to `/api/download`. According to the ReClip source code at lines 77-79, this endpoint:

1. Generates a cryptographically random `job_id`
2. Inserts a new entry into the `jobs` dictionary with the target URL and format preference
3. Immediately returns the `job_id` to the client for polling

This non-blocking response ensures the HTTP connection closes within milliseconds, even for multi-gigabyte downloads.

### Background Worker Threads

For each new job, ReClip spawns a dedicated daemon thread. Lines 80-82 in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) implement this pattern:

```python
thread = threading.Thread(target=run_download,
                          args=(job_id, url, format_choice, format_id))
thread.daemon = True
thread.start()

```

Marking the thread as `daemon = True` ensures that if the main Flask process terminates, all active downloads are automatically cleaned up without leaving zombie processes. The thread executes the `run_download` function, passing the job context and format selection arguments.

## Download Execution and Error Handling

### yt-dlp Command Construction

Inside `run_download` (lines 36-44), ReClip builds a yt-dlp command list based on the requested format:

- **Audio format**: Downloads best audio and converts to MP3
- **Video format**: Downloads best quality MP4 available

The function constructs the final command array and executes it via `subprocess.run` with a strict 5-minute timeout (line 48).

### File Processing and Cleanup

After subprocess completion, lines 54-66 scan the `downloads` folder to locate the generated file. The logic differentiates between audio (`*.mp3`) and video (`*.mp4`) outputs, removes extraneous files generated during the download process, and sanitizes the filename for safe HTTP delivery.

### State Transitions

Successful completions update the job record at lines 74-84:

```python
job["status"] = "done"
job["file"] = chosen
job["filename"] = sanitized_name

```

Error handling (lines 50-53 and 85-89) captures non-zero exit codes, subprocess timeouts, and unexpected exceptions, setting `job["status"] = "error"` and populating `job["error"]` with diagnostic information for client visibility.

## Status Retrieval and File Delivery

### Polling Endpoint

Clients monitor progress via GET requests to `/api/status/<job_id>`. As implemented at lines 87-96, this endpoint returns a JSON object containing the current `status`, any `error` message, and the final `filename` once the job reaches the `"done"` state.

### File Streaming

Once a job reports completion, the actual media file is retrieved via `/api/file/<job_id>`. Lines 99-104 use Flask's `send_file` to stream the stored file as an attachment, ensuring the browser prompts for download using the sanitized filename stored in the job registry.

## Practical Implementation Examples

### Starting a Download Job

```python
import requests

resp = requests.post(
    "http://localhost:8899/api/download",
    json={"url": "https://www.youtube.com/watch?v=abc123",
          "format": "video"}  # or "audio"

)
job_id = resp.json()["job_id"]

```

### Polling for Completion

```python
import time
import requests

while True:
    status = requests.get(f"http://localhost:8899/api/status/{job_id}").json()
    print(status["status"])
    if status["status"] in ("done", "error"):
        break
    time.sleep(2)

```

### Retrieving the Finished File

```python
if status["status"] == "done":
    r = requests.get(f"http://localhost:8899/api/file/{job_id}", stream=True)
    with open(status["filename"], "wb") as f:
        for chunk in r.iter_content(8192):
            f.write(chunk)

```

## Summary

- **In-Memory Registry**: ReClip uses a global `jobs` dictionary in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) (line 13) to track all download states without requiring external databases or message queues.
- **Thread-Based Concurrency**: Each download spawns a daemon thread via `threading.Thread` (lines 80-82), keeping the Flask main thread responsive while yt-dlp processes media in the background.
- **Synchronous Subprocess Execution**: The `run_download` function uses `subprocess.run` with a 300-second timeout to wrap yt-dlp, handling format conversion and file cleanup before updating job status.
- **Client Polling Architecture**: Stateless HTTP endpoints (`/api/status/<job_id>` and `/api/file/<job_id>`) allow clients to poll for completion and stream results without persistent WebSocket connections.

## Frequently Asked Questions

### How does ReClip track the state of active download jobs?

ReClip tracks active downloads using the global `jobs` dictionary defined in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py). Each entry maps a unique job ID to a dictionary containing `status`, `url`, `title`, `file`, and `filename` keys. The status field transitions from `"downloading"` to `"done"` or `"error"` as the background thread progresses, allowing the `/api/status/<job_id>` endpoint to report real-time state without database queries.

### What happens if a download exceeds the 5-minute timeout?

If `subprocess.run` (line 48) exceeds the 300-second timeout, it raises a `TimeoutExpired` exception. The error handler at lines 85-89 catches this exception, updates the job status to `"error"`, and stores the timeout details in the `job["error"]` field. The client polling the status endpoint will receive the error state and can handle the failure accordingly.

### Is ReClip's job manager suitable for high-traffic production workloads?

No. ReClip's architecture uses an in-memory dictionary (`jobs = {}`) and spawns unlimited daemon threads per request, which limits scalability to the available RAM and thread count of the host machine. For production workloads requiring high concurrency, persistent storage, or crash recovery, the codebase would need migration to a proper task queue like Celery or RQ with Redis backing.

### How are downloaded files cleaned up after retrieval?

ReClip performs cleanup during the `run_download` execution (lines 54-66) by removing extraneous files from the downloads folder, keeping only the target MP3 or MP4 file. However, the code does not automatically delete files after client download via `/api/file/<job_id>`—the file remains on disk until manually removed or overwritten by subsequent downloads with the same sanitized filename.