# How Modly Integrates with HuggingFace for Model Downloads

> Learn how Modly integrates with HuggingFace for seamless model downloads. Modly uses huggingface_hub to fetch weights and stream files on demand.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: how-to-guide
- Published: 2026-08-15

---

**Modly leverages the `huggingface_hub` Python library to fetch model weights automatically via `snapshot_download` and stream individual files on-demand through authenticated REST endpoints.**

The `lightningpixel/modly` repository implements a dual-path HuggingFace integration that handles both bulk model synchronization and granular file streaming. This architecture allows generator extensions to declare their source repositories while giving users control over authentication, file filtering, and download strategies.

## Core HuggingFace Integration Architecture

Modly splits its HuggingFace integration into two complementary workflows. The **automatic download** system clones entire repositories when generators initialize, while the **streaming API** provides on-demand access to specific files without full local mirroring.

Both paths rely on the official `huggingface_hub` SDK and support token-based authentication through environment variables or query parameters. The integration centers on three core files: [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) for automatic downloads, [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) for streaming endpoints, and [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) for manifest injection.

## Automatic Weight Download with snapshot_download

When a generator instance first runs, Modly checks whether the model exists locally. If `is_downloaded()` returns false, the `BaseGenerator` class triggers `_auto_download()` defined in [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) (lines 44-68).

This method constructs an ignore pattern from the generator's manifest and calls `snapshot_download()`:

```python

# From api/services/generators/base.py

def _auto_download(self):
    target_dir = self._get_model_path()
    target_dir.mkdir(parents=True, exist_ok=True)
    
    ignore_patterns = self.hf_skip_prefixes if hasattr(self, 'hf_skip_prefixes') else None
    
    snapshot_download(
        repo_id=self.hf_repo,
        local_dir=str(target_dir),
        ignore_patterns=ignore_patterns,
        token=self._get_hf_token()
    )

```

The `hf_repo` identifier and skip lists are injected at runtime by [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py), which reads the extension's [`manifest.json`](https://github.com/lightningpixel/modly/blob/main/manifest.json) during initialization. This ensures the generator knows exactly which HuggingFace repository to mirror and which files to exclude.

## Streaming File Downloads via REST API

For scenarios requiring direct file access without full repository cloning, Modly exposes the `/model/{model_id}/download` endpoint in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) (lines 85-92). This endpoint implements `_download_file_streamed()` (lines 81-120) to handle chunked transfers with resume support.

The streaming workflow follows these steps:

1. **Enumerate files**: Uses `list_repo_files(repo_id, token=hf_token)` to retrieve the repository index.
2. **Apply filters**: Respects `hf_skip_prefixes` and `hf_include_prefixes` defined in the manifest (lines 56-69).
3. **Generate URLs**: Builds download links via `hf_hub_url(repo_id, filename)`.
4. **Stream chunks**: Downloads bytes incrementally while reporting progress through Server-Sent Events (SSE).

```python

# Conceptual flow from api/routers/model.py

files = list_repo_files(repo_id, token=token)
filtered = [f for f in files if not any(f.startswith(p) for p in skip_prefixes)]
for filename in filtered:
    url = hf_hub_url(repo_id, filename)
    yield from _download_file_streamed(url, token)

```

## Authentication and Token Management

Modly supports multiple authentication vectors for private HuggingFace repositories. The system checks for tokens in the following priority order:

- Query parameter: `?token=...` passed to the download endpoint
- Environment variables: `HUGGING_FACE_HUB_TOKEN` or `HF_TOKEN`

Token extraction occurs in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) at lines 71-73 for the endpoint handler and lines 100-106 within the streaming function. The token is passed directly to all `huggingface_hub` SDK calls, ensuring gated model access works seamlessly.

## Selective File Filtering

Extensions control their download footprint through [`manifest.json`](https://github.com/lightningpixel/modly/blob/main/manifest.json) entries:

- **`hf_skip_prefixes`**: Array of path prefixes to exclude (e.g., `["samples/", "README.md"]`)
- **`hf_include_prefixes`**: Array of path prefixes to exclusively include (e.g., `["weights/"]`)

These lists are read in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) (lines 56-69) and converted into `ignore_patterns` for `snapshot_download` or used to filter the file list in streaming operations. This selective approach reduces disk usage and download times by fetching only necessary model artifacts.

```json
{
  "hf_repo": "username/example-model",
  "hf_skip_prefixes": ["samples/", "README.md"],
  "hf_include_prefixes": ["weights/"]
}

```

## Usage Examples

### Triggering Automatic Download via Python API

When implementing a generator class, no explicit download code is required. The base class handles synchronization automatically:

```python
from modly.api.services.generators.base import BaseGenerator

class MyModel(BaseGenerator):
    MODEL_ID = "example-model"
    # hf_repo injected from manifest at runtime

generator = MyModel()

# This triggers _auto_download() if model not present

generator.generate(prompt="A sunny beach", progress_cb=lambda p, s: print(p, s))

```

### Requesting Streaming Downloads via HTTP

Clients can download specific filtered files without local storage requirements:

```bash
curl -N "http://localhost:8000/model/example-model/download?token=$HF_TOKEN"

```

The response streams Server-Sent Events reporting per-file progress and handles resumable transfers internally.

## Summary

- **Modly integrates HuggingFace through the `huggingface_hub` SDK**, using `snapshot_download` for bulk operations and `list_repo_files` with `hf_hub_url` for streaming.
- **Automatic downloads occur in [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py)** via the `_auto_download()` method, which clones repositories based on manifest-declared `hf_repo` values.
- **Streaming endpoints live in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py)**, implementing chunked file transfers with support for resume capabilities and progress callbacks.
- **Authentication supports both query parameters and environment variables**, with token handling at lines 71-73 and 100-106 of the router file.
- **File filtering uses `hf_skip_prefixes` and `hf_include_prefixes`** defined in extension manifests, parsed at lines 56-69 of [`model.py`](https://github.com/lightningpixel/modly/blob/main/model.py) to minimize unnecessary downloads.

## Frequently Asked Questions

### How does Modly authenticate with HuggingFace for private models?

Modly accepts HuggingFace tokens via the `token` query parameter in download requests or through the `HUGGING_FACE_HUB_TOKEN` and `HF_TOKEN` environment variables. The system extracts these credentials in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) and passes them to every `huggingface_hub` SDK call, ensuring gated repository access works for both automatic and streaming downloads.

### What is the difference between automatic and streaming downloads in Modly?

Automatic downloads use `snapshot_download` in [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) to mirror entire HuggingFace repositories locally when a generator first initializes. Streaming downloads, handled by `/model/{model_id}/download` in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py), enumerate specific files via `list_repo_files` and transfer them chunk-by-chunk without persisting a full local copy, ideal for ephemeral model access.

### Can I exclude specific files from HuggingFace downloads?

Yes. Define `hf_skip_prefixes` in your extension's [`manifest.json`](https://github.com/lightningpixel/modly/blob/main/manifest.json) to ignore specific paths, or use `hf_include_prefixes` to whitelist only necessary directories. Modly applies these filters in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) (lines 56-69) during both bulk `snapshot_download` operations and individual file streaming, reducing bandwidth and storage requirements.

### Where is the HuggingFace integration logic located in the Modly codebase?

The integration spans three primary files: [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) contains the `BaseGenerator` class with `_auto_download()` for automatic weight syncing; [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) implements the streaming download endpoint and token management; and [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) injects repository metadata from extension manifests into generator instances at runtime.