How Modly Handles Resumable Downloads for AI Models: Architecture & Implementation
Modly implements resumable downloads for AI models by coordinating HTTP Range requests in its Python FastAPI backend with partial file (.part) management, while the Electron frontend monitors progress via SSE streams and enforces stall timeouts to ensure interruption recovery.
Large AI models from HuggingFace can reach several gigabytes, making network interruptions costly without resume capability. The lightningpixel/modly repository solves this through a coordinated architecture between the Electron main process and a Python FastAPI backend that supports partial content streaming and persistent temporary files.
Architectural Overview
Modly's resumable download system spans four distinct layers that communicate via HTTP and Server-Sent Events (SSE):
| Layer | Responsibility | Key Implementation |
|---|---|---|
| Frontend UI | Triggers downloads and receives progress events via SSE | downloadModelFromHF in electron/main/model-downloader.ts |
| Electron Main Process | Opens HTTP connections, parses SSE messages, forwards callbacks, enforces stall timeouts | SSE handling and timeout logic in downloadModelFromHF |
| Python FastAPI Backend | Streams files, supports HTTP Range requests, writes to .part files, emits progress |
_download_file_streamed and _download_status in api/routers/model.py |
| File System | Stores partial downloads as *.part files, preserves incomplete state for resume |
Temporary file handling in _download_file_streamed |
Backend Resume Logic
The Python backend in api/routers/model.py implements the core resumable download logic through HTTP Range request handling and atomic file operations.
Partial File Detection
Before initiating a download, the system checks for existing temporary files:
# Conceptual flow based on _download_file_streamed lines 18-24
part_file = f"{target_path}.part"
existing_bytes = os.path.getsize(part_file) if os.path.exists(part_file) else 0
headers = {}
if existing_bytes > 0:
headers["Range"] = f"bytes={existing_bytes}-"
If a .part file exists, its size determines the byte offset for the Range header, allowing the server to resume transmission from the interruption point.
Range Header and Resume Flag
When resuming, the backend sets a boolean flag based on the server response:
# Lines 27-36 in _download_file_streamed
resumed = False
if existing_bytes > 0:
response = requests.get(url, headers=headers, stream=True)
if response.status_code == 206: # Partial Content
resumed = True
The resumed variable propagates to _download_status, which prefixes UI messages with "Resuming..." to indicate recovery mode.
File Appending and Progress Reporting
The backend opens files in append-binary mode ("ab") when resuming, or write-binary mode ("wb") for fresh downloads:
mode = "ab" if resumed else "wb"
with open(part_file, mode) as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
progress_cb(downloaded_bytes, total_bytes, resumed)
Progress callbacks emit SSE events containing the current byte count, total size (extracted from Content-Range or Content-Length headers), and resume status.
Cancellation Handling
On explicit cancellation, only .part files are removed while completed downloads remain intact. This ensures that interrupted downloads can resume on the next attempt, while finished models stay available in the library.
Electron Client Implementation
The Electron main process in electron/main/model-downloader.ts does not implement resume logic directly but ensures continuity through connection management and state verification.
SSE Parsing and Progress Forwarding
The client establishes an SSE connection to the backend endpoint and parses progress events:
// Conceptual implementation based on lines 65-76
response.on('data', (chunk) => {
const lines = chunk.toString().split('\n');
for (const line of lines) {
if (line.startsWith('data:')) {
const data = JSON.parse(line.slice(5));
if (data.percent) onProgress(data);
if (data.error) throw new Error(data.error);
}
}
});
Stall Detection and Timeout Recovery
To handle frozen connections without explicit errors, the client implements a stall timeout:
// Lines 24-52 in downloadModelFromHF
const STALL_TIMEOUT_MS = 120_000; // 2 minutes
let lastDataTime = Date.now();
const checkStall = setInterval(() => {
if (Date.now() - lastDataTime > STALL_TIMEOUT_MS) {
clearInterval(checkStall);
abortController.abort();
reject(new Error('Download stalled'));
}
}, 5000);
If no SSE data arrives within two minutes, the connection aborts and the UI can retry, triggering the backend's resume logic to continue from the last saved byte.
Local State Verification
Two helper functions prevent redundant downloads:
isModelDownloaded(modelsDir, modelId): Checks final file existence without network overheadlistDownloadedModels(modelsDir): Returns arrays of downloaded models with size metadata for UI display
End-to-End Resumable Flow
- User initiates download → UI calls
downloadModelFromHF(repoId, modelId, onProgress) - Electron creates SSE request → Connects to
http://127.0.0.1:8765/model/hf-download - Backend detects partial file → Calculates
existing_bytesand sendsRange: bytes=<offset>-header - Server responds with 206 → Backend sets
resumed=Trueand opens.partfile in"ab"mode - Progress streaming → SSE events emit
"percent"and"status"(prefixed with "Resuming..." when applicable) - Stall handling → If the connection freezes, Electron aborts after 120 seconds; retry resumes from saved state
- Completion → Temporary
.partfile is renamed to final model name; subsequent calls toisModelDownloadedreturntrue
Implementation Examples
Triggering a Resumable Download
import { downloadModelFromHF } from '@/electron/main/model-downloader';
function handleProgress(progress) {
console.log(
`Downloading ${progress.file}: ${progress.percent}%`,
progress.status
);
}
// Automatically resumes if partial download exists
downloadModelFromHF(
'stabilityai/stable-diffusion',
'sd-v1-4',
handleProgress,
/* skipPrefixes */ ['samples/'],
/* includePrefixes */ undefined
).catch(err => {
console.error('Download failed:', err);
});
Checking Download Status
import { isModelDownloaded } from '@/electron/main/model-downloader';
const modelsDir = '/Users/me/.modly/models';
const modelId = 'sd-v1-4';
if (isModelDownloaded(modelsDir, modelId)) {
console.log('Model already downloaded – skipping fetch.');
}
Listing Downloaded Models
import { listDownloadedModels } from '@/electron/main/model-downloader';
const models = listDownloadedModels('/Users/me/.modly/models');
models.forEach(m => {
console.log(`${m.id} – ${m.name}: ${m.size_gb} GB`);
});
Summary
- HTTP Range Requests: The backend uses
Range: bytes=<offset>-headers and handles 206 Partial Content responses to resume interrupted transfers. - Temporary File Management: Partial downloads persist as
.partfiles in append-binary mode, preserving exact byte positions across application restarts. - SSE Progress Streaming: Real-time progress updates flow from Python backend to Electron frontend via Server-Sent Events, including explicit "Resuming..." status indicators.
- Stall Protection: A 120-second timeout in the Electron layer detects frozen connections and triggers retry logic that leverages the backend's resume capability.
- State Verification: Local filesystem checks via
isModelDownloadedandlistDownloadedModelsprevent redundant network requests for completed models.
Frequently Asked Questions
How does Modly detect partially downloaded files?
Modly checks for existing .part files in the target directory before initiating a download. In api/routers/model.py, the _download_file_streamed function calculates existing_bytes from the temporary file's size and includes this value in the HTTP Range header sent to the HuggingFace server.
What happens if the internet connection drops during download?
The Electron frontend monitors data flow through a 120-second stall timeout (STALL_TIMEOUT_MS). If no SSE progress events arrive within this window, the connection aborts and throws an error. When the user retries, the backend detects the surviving .part file and resumes from the last received byte using HTTP 206 Partial Content responses.
Are partial files cleaned up automatically?
No, .part files persist intentionally to enable resume functionality. They are only deleted upon explicit cancellation via the UI or when a download completes successfully and the file is renamed to its final destination. This design ensures interruption recovery even across application restarts.
Can users configure the stall timeout duration?
The stall timeout is hardcoded to 120 seconds (120_000 milliseconds) in electron/main/model-downloader.ts as a constant. Modly does not expose this value through user settings in the current implementation, though the architecture would support making it configurable through the settings-store.ts module if needed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →