# How Modly Manages Model Downloads and Caching: A Complete Technical Guide

> Discover how Modly manages model downloads and caching efficiently. Learn about lazy download on first use, streaming API with pause/cancel, and user-controlled storage for seamless model management.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: how-to-guide
- Published: 2026-08-20

---

**Modly stores all model weights in a user-controlled `MODELS_DIR` directory (default: `~/.modly/models`), uses lazy download-on-first-use via `huggingface_hub.snapshot_download`, and provides a streaming API with pause/cancel support for fine-grained progress control.**

This guide examines how the `lightningpixel/modly` repository implements its model download and caching system. Understanding Modly's approach to model downloads and caching helps developers integrate custom models, optimize storage, and build responsive UIs that track download progress.

## How Modly Organizes Model Storage

Modly centralizes all model artifacts under a single configurable directory. The **GeneratorRegistry** in [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) maintains this path and orchestrates model lifecycle operations.

### The MODELS_DIR Configuration

- **Default location**: `~/.modly/models`
- **Runtime relocation**: Supported via settings API
- **Structure**: Each model gets its own subdirectory (`MODELS_DIR/<model_id>`)

When you change `MODELS_DIR` at runtime, Modly triggers a bulk path rewrite and optionally unloads all models so the new directory becomes the source of truth (`api/services/generator_registry.py#L24-L33`).

## Cache Validation: How Modly Detects Existing Models

Before any download begins, Modly checks whether a model already exists locally. This prevents redundant downloads and enables offline operation.

### The is_downloaded() Method

The **BaseGenerator** class in [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) defines the cache validation logic (`L61-L70`):

```python

# From api/services/generators/base.py

def is_downloaded(self) -> bool:
    """
    Checks if the model files are present in the cache.
    Uses manifest-defined download_check or falls back to directory contents.
    """
    check_path = self.MODELS_DIR / self.model_id / self.manifest.get("download_check", "")
    if check_path.exists():
        return True
    # Fallback: check if any files exist in model directory

    model_dir = self.MODELS_DIR / self.model_id
    return model_dir.exists() and any(model_dir.iterdir())

```

**Key behaviors:**
- **Manifest-defined check**: If `download_check` is specified in the model manifest, Modly verifies that exact file exists
- **Fallback detection**: Without a specific check file, Modly validates that the model directory contains any files

## Automatic Downloads on First Use

Modly implements a **lazy download** pattern: models fetch automatically when first requested, not at startup.

### The Download Trigger Flow

1. **Registry checks cache**: `GeneratorRegistry.get_active()` calls `gen.is_downloaded()` (`api/services/generator_registry.py#L252-L254`)
2. **Download if missing**: When not cached, `_auto_download()` executes
3. **Generator instantiation**: Metadata from the model manifest (e.g., `hf_repo`, `hf_skip_prefixes`, `hf_include_prefixes`) injects into the generator

### The _auto_download() Implementation

```python

# From api/services/generators/base.py#L45-L68

def _auto_download(self) -> None:
    """
    Default download implementation using huggingface_hub.
    Respects hf_skip_prefixes and hf_include_prefixes from manifest.
    """
    from huggingface_hub import snapshot_download
    
    repo_id = self.manifest.get("hf_repo")
    if not repo_id:
        raise ValueError(f"No hf_repo specified for model {self.model_id}")
    
    local_dir = self.MODELS_DIR / self.model_id
    
    snapshot_download(
        repo_id=repo_id,
        local_dir=str(local_dir),
        allow_patterns=self.manifest.get("hf_include_prefixes"),
        ignore_patterns=self.manifest.get("hf_skip_prefixes"),
        local_dir_use_symlinks=False,
    )

```

**Configuration options via manifest:**
- `hf_repo`: Required HuggingFace repository identifier
- `hf_include_prefixes`: Only download files matching these patterns
- `hf_skip_prefixes`: Exclude files matching these patterns
- `download_check`: Specific file to verify for cache validity

## Streaming Downloads with Progress Control

For large models or UI-driven installations, Modly exposes a **streaming download endpoint** with Server-Sent Events (SSE).

### The Streaming API Endpoint

The `/api/model/hf-download` router in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) (`L25-L78`) provides:

```python

# Client example: Stream download with progress

import sseclient  # pip install sseclient-py

import json
import requests

url = "http://localhost:8000/api/model/hf-download"
params = {
    "repo_id": "stabilityai/stable-fast-3d",
    "model_id": "sf3d",
}

resp = requests.get(url, params=params, stream=True)
client = sseclient.SSEClient(resp)

for event in client.events():
    data = json.loads(event.data)
    print(f"{data.get('percent', 0):.1f}% – {data.get('status')}")
    
    if data.get("error"):
        print(f"Download failed: {data['error']}")
        break
    elif data.get("percent") == 100:
        print("Download complete")
        break

```

### Pause and Cancel Controls

Modly provides thread-safe controls for active downloads:

```python

# Pause a running download

import requests

requests.post(
    "http://localhost:8000/api/model/hf-download/pause",
    json={"model_id": "sf3d"}
).json()

# → {"paused": True}

# Resume from pause

requests.post(
    "http://localhost:8000/api/model/hf-download/resume",
    json={"model_id": "sf3d"}
).json()

# Cancel and clean up

requests.post(
    "http://localhost:8000/api/model/hf-download/cancel",
    json={"model_id": "sf3d"}
).json()

# → {"cancelled": True}

```

## Model Lifecycle: Memory vs. Cache

Modly distinguishes between **memory-resident loaded models** and **disk-cached weights**. This separation enables flexible resource management.

### Unload from Memory (Keep Cache)

```python

# Remove model from RAM but preserve downloaded files

import requests

requests.post(
    "http://localhost:8000/api/model/unload",
    json={"model_id": "sf3d"}
).json()

```

This calls `gen.unload()` on the generator, freeing GPU/CPU memory without touching `MODELS_DIR/<model_id>`.

### Clear Cache via API

The unload endpoint with path parameter removes both memory and disk copies (`api/routers/model.py#L100-L108`):

```python

# DELETE /api/model/unload/{model_id} - removes cache entirely

requests.delete("http://localhost:8000/api/model/unload/sf3d").json()

```

## Relocating the Cache Directory

You can move the entire model cache to a different location—useful for external drives or network storage.

```python

# Change MODELS_DIR at runtime

import requests

new_dir = "/mnt/nas/modly_models"

result = requests.post(
    "http://localhost:8000/api/settings",
    json={"models_dir": new_dir}
).json()

# → {"models_dir": "/mnt/nas/modly_models", ...}

```

**Side effects:**
- All active models unload from memory
- New models download to the updated location
- Existing cached models remain at the old path (not auto-migrated)

## Practical Code Examples

### Complete Model Switch with Auto-Download

```python
import requests

def switch_model(model_id: str) -> dict:
    """
    Switches to a model, automatically downloading if not cached.
    Returns the active model configuration.
    """
    response = requests.post(
        "http://localhost:8000/api/model/switch",
        json={"model_id": model_id}
    )
    result = response.json()
    
    if result.get("active") == model_id:
        print(f"Model {model_id} is ready")
        if result.get("downloaded"):
            print("  (fetched from cache)")
        else:
            print("  (downloaded fresh)")
    
    return result

# Usage

switch_model("sf3d")

```

### Checking Model Status Before Switching

```python
def get_model_status(model_id: str) -> dict:
    """Check if a model is downloaded, loaded, or needs fetch."""
    response = requests.get("http://localhost:8000/api/model/status")
    models = response.json().get("models", {})
    
    info = models.get(model_id, {})
    return {
        "downloaded": info.get("is_downloaded", False),
        "loaded": info.get("is_loaded", False),
        "active": info.get("is_active", False),
        "manifest": info.get("manifest", {}),
    }

status = get_model_status("sf3d")
print(f"Downloaded: {status['downloaded']}, Loaded: {status['loaded']}")

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) | Central registry; discovers extensions, stores `MODELS_DIR`, orchestrates loading/unloading |
| [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) | Abstract generator with `is_downloaded()`, `_auto_download()`, and lifecycle hooks |
| [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) | FastAPI endpoints for status, switching, streaming HF download, and cache management |
| [`api/services/extension_process.py`](https://github.com/lightningpixel/modly/blob/main/api/services/extension_process.py) | Subprocess-based extensions for models with custom download logic |
| [`tools/modly-cli/agent.py`](https://github.com/lightningpixel/modly/blob/main/tools/modly-cli/agent.py) | CLI wrapper forwarding `MODELS_DIR` environment variable to backend |

## Summary

- **Cache location**: Configurable via `MODELS_DIR` (default `~/.modly/models`), relocatable at runtime
- **Cache validation**: `BaseGenerator.is_downloaded()` checks manifest-defined `download_check` file or directory contents
- **Auto-download**: Lazy fetch on first use via `_auto_download()` using `huggingface_hub.snapshot_download`
- **Filter options**: Manifest supports `hf_skip_prefixes` and `hf_include_prefixes` for selective downloads
- **Streaming API**: SSE-based `/api/model/hf-download` with pause, resume, and cancel controls
- **Memory separation**: `unload()` frees RAM without deleting cache; DELETE endpoint clears both

## Frequently Asked Questions

### How does Modly handle interrupted downloads?

Modly relies on `huggingface_hub.snapshot_download` which implements resume capability. For streaming downloads via the SSE endpoint, pause/resume controls let you suspend and continue without restart. Cancelled streaming downloads clean up partial files automatically.

### Can I use a model without HuggingFace integration?

Yes. The `BaseGenerator` in [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) supports subclasses overriding `_auto_download()` with custom logic. Models can also ship as subprocess-based extensions handled by [`api/services/extension_process.py`](https://github.com/lightningpixel/modly/blob/main/api/services/extension_process.py) for completely independent download mechanisms.

### What happens if I change MODELS_DIR while models are loaded?

The registry unloads all models from memory (`api/services/generator_registry.py#L24-L33`). The new directory becomes active immediately, though existing cached weights at the old location remain untouched. You would need to manually migrate or re-download to the new path.

### How do I verify which files were actually cached?

Check the `MODELS_DIR/<model_id>` directory directly. Modly does not maintain a separate manifest of cached files—it validates presence using `is_downloaded()` each time. For debugging, query `/api/model/status` to see `is_downloaded` state per model.