How Modly Manages Model Downloads and Caching: A Complete Technical Guide
Modly stores all model weights in a user-controlled MODELS_DIR directory (default: ~/.modly/models), uses lazy download-on-first-use via huggingface_hub.snapshot_download, and provides a streaming API with pause/cancel support for fine-grained progress control.
This guide examines how the lightningpixel/modly repository implements its model download and caching system. Understanding Modly's approach to model downloads and caching helps developers integrate custom models, optimize storage, and build responsive UIs that track download progress.
How Modly Organizes Model Storage
Modly centralizes all model artifacts under a single configurable directory. The GeneratorRegistry in api/services/generator_registry.py maintains this path and orchestrates model lifecycle operations.
The MODELS_DIR Configuration
- Default location:
~/.modly/models - Runtime relocation: Supported via settings API
- Structure: Each model gets its own subdirectory (
MODELS_DIR/<model_id>)
When you change MODELS_DIR at runtime, Modly triggers a bulk path rewrite and optionally unloads all models so the new directory becomes the source of truth (api/services/generator_registry.py#L24-L33).
Cache Validation: How Modly Detects Existing Models
Before any download begins, Modly checks whether a model already exists locally. This prevents redundant downloads and enables offline operation.
The is_downloaded() Method
The BaseGenerator class in api/services/generators/base.py defines the cache validation logic (L61-L70):
# From api/services/generators/base.py
def is_downloaded(self) -> bool:
"""
Checks if the model files are present in the cache.
Uses manifest-defined download_check or falls back to directory contents.
"""
check_path = self.MODELS_DIR / self.model_id / self.manifest.get("download_check", "")
if check_path.exists():
return True
# Fallback: check if any files exist in model directory
model_dir = self.MODELS_DIR / self.model_id
return model_dir.exists() and any(model_dir.iterdir())
Key behaviors:
- Manifest-defined check: If
download_checkis specified in the model manifest, Modly verifies that exact file exists - Fallback detection: Without a specific check file, Modly validates that the model directory contains any files
Automatic Downloads on First Use
Modly implements a lazy download pattern: models fetch automatically when first requested, not at startup.
The Download Trigger Flow
- Registry checks cache:
GeneratorRegistry.get_active()callsgen.is_downloaded()(api/services/generator_registry.py#L252-L254) - Download if missing: When not cached,
_auto_download()executes - Generator instantiation: Metadata from the model manifest (e.g.,
hf_repo,hf_skip_prefixes,hf_include_prefixes) injects into the generator
The _auto_download() Implementation
# From api/services/generators/base.py#L45-L68
def _auto_download(self) -> None:
"""
Default download implementation using huggingface_hub.
Respects hf_skip_prefixes and hf_include_prefixes from manifest.
"""
from huggingface_hub import snapshot_download
repo_id = self.manifest.get("hf_repo")
if not repo_id:
raise ValueError(f"No hf_repo specified for model {self.model_id}")
local_dir = self.MODELS_DIR / self.model_id
snapshot_download(
repo_id=repo_id,
local_dir=str(local_dir),
allow_patterns=self.manifest.get("hf_include_prefixes"),
ignore_patterns=self.manifest.get("hf_skip_prefixes"),
local_dir_use_symlinks=False,
)
Configuration options via manifest:
hf_repo: Required HuggingFace repository identifierhf_include_prefixes: Only download files matching these patternshf_skip_prefixes: Exclude files matching these patternsdownload_check: Specific file to verify for cache validity
Streaming Downloads with Progress Control
For large models or UI-driven installations, Modly exposes a streaming download endpoint with Server-Sent Events (SSE).
The Streaming API Endpoint
The /api/model/hf-download router in api/routers/model.py (L25-L78) provides:
# Client example: Stream download with progress
import sseclient # pip install sseclient-py
import json
import requests
url = "http://localhost:8000/api/model/hf-download"
params = {
"repo_id": "stabilityai/stable-fast-3d",
"model_id": "sf3d",
}
resp = requests.get(url, params=params, stream=True)
client = sseclient.SSEClient(resp)
for event in client.events():
data = json.loads(event.data)
print(f"{data.get('percent', 0):.1f}% – {data.get('status')}")
if data.get("error"):
print(f"Download failed: {data['error']}")
break
elif data.get("percent") == 100:
print("Download complete")
break
Pause and Cancel Controls
Modly provides thread-safe controls for active downloads:
# Pause a running download
import requests
requests.post(
"http://localhost:8000/api/model/hf-download/pause",
json={"model_id": "sf3d"}
).json()
# → {"paused": True}
# Resume from pause
requests.post(
"http://localhost:8000/api/model/hf-download/resume",
json={"model_id": "sf3d"}
).json()
# Cancel and clean up
requests.post(
"http://localhost:8000/api/model/hf-download/cancel",
json={"model_id": "sf3d"}
).json()
# → {"cancelled": True}
Model Lifecycle: Memory vs. Cache
Modly distinguishes between memory-resident loaded models and disk-cached weights. This separation enables flexible resource management.
Unload from Memory (Keep Cache)
# Remove model from RAM but preserve downloaded files
import requests
requests.post(
"http://localhost:8000/api/model/unload",
json={"model_id": "sf3d"}
).json()
This calls gen.unload() on the generator, freeing GPU/CPU memory without touching MODELS_DIR/<model_id>.
Clear Cache via API
The unload endpoint with path parameter removes both memory and disk copies (api/routers/model.py#L100-L108):
# DELETE /api/model/unload/{model_id} - removes cache entirely
requests.delete("http://localhost:8000/api/model/unload/sf3d").json()
Relocating the Cache Directory
You can move the entire model cache to a different location—useful for external drives or network storage.
# Change MODELS_DIR at runtime
import requests
new_dir = "/mnt/nas/modly_models"
result = requests.post(
"http://localhost:8000/api/settings",
json={"models_dir": new_dir}
).json()
# → {"models_dir": "/mnt/nas/modly_models", ...}
Side effects:
- All active models unload from memory
- New models download to the updated location
- Existing cached models remain at the old path (not auto-migrated)
Practical Code Examples
Complete Model Switch with Auto-Download
import requests
def switch_model(model_id: str) -> dict:
"""
Switches to a model, automatically downloading if not cached.
Returns the active model configuration.
"""
response = requests.post(
"http://localhost:8000/api/model/switch",
json={"model_id": model_id}
)
result = response.json()
if result.get("active") == model_id:
print(f"Model {model_id} is ready")
if result.get("downloaded"):
print(" (fetched from cache)")
else:
print(" (downloaded fresh)")
return result
# Usage
switch_model("sf3d")
Checking Model Status Before Switching
def get_model_status(model_id: str) -> dict:
"""Check if a model is downloaded, loaded, or needs fetch."""
response = requests.get("http://localhost:8000/api/model/status")
models = response.json().get("models", {})
info = models.get(model_id, {})
return {
"downloaded": info.get("is_downloaded", False),
"loaded": info.get("is_loaded", False),
"active": info.get("is_active", False),
"manifest": info.get("manifest", {}),
}
status = get_model_status("sf3d")
print(f"Downloaded: {status['downloaded']}, Loaded: {status['loaded']}")
Key Source Files
| File | Purpose |
|---|---|
api/services/generator_registry.py |
Central registry; discovers extensions, stores MODELS_DIR, orchestrates loading/unloading |
api/services/generators/base.py |
Abstract generator with is_downloaded(), _auto_download(), and lifecycle hooks |
api/routers/model.py |
FastAPI endpoints for status, switching, streaming HF download, and cache management |
api/services/extension_process.py |
Subprocess-based extensions for models with custom download logic |
tools/modly-cli/agent.py |
CLI wrapper forwarding MODELS_DIR environment variable to backend |
Summary
- Cache location: Configurable via
MODELS_DIR(default~/.modly/models), relocatable at runtime - Cache validation:
BaseGenerator.is_downloaded()checks manifest-defineddownload_checkfile or directory contents - Auto-download: Lazy fetch on first use via
_auto_download()usinghuggingface_hub.snapshot_download - Filter options: Manifest supports
hf_skip_prefixesandhf_include_prefixesfor selective downloads - Streaming API: SSE-based
/api/model/hf-downloadwith pause, resume, and cancel controls - Memory separation:
unload()frees RAM without deleting cache; DELETE endpoint clears both
Frequently Asked Questions
How does Modly handle interrupted downloads?
Modly relies on huggingface_hub.snapshot_download which implements resume capability. For streaming downloads via the SSE endpoint, pause/resume controls let you suspend and continue without restart. Cancelled streaming downloads clean up partial files automatically.
Can I use a model without HuggingFace integration?
Yes. The BaseGenerator in api/services/generators/base.py supports subclasses overriding _auto_download() with custom logic. Models can also ship as subprocess-based extensions handled by api/services/extension_process.py for completely independent download mechanisms.
What happens if I change MODELS_DIR while models are loaded?
The registry unloads all models from memory (api/services/generator_registry.py#L24-L33). The new directory becomes active immediately, though existing cached weights at the old location remain untouched. You would need to manually migrate or re-download to the new path.
How do I verify which files were actually cached?
Check the MODELS_DIR/<model_id> directory directly. Modly does not maintain a separate manifest of cached files—it validates presence using is_downloaded() each time. For debugging, query /api/model/status to see is_downloaded state per model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →