How Modly Integrates with HuggingFace for Model Downloads
Modly leverages the huggingface_hub Python library to fetch model weights automatically via snapshot_download and stream individual files on-demand through authenticated REST endpoints.
The lightningpixel/modly repository implements a dual-path HuggingFace integration that handles both bulk model synchronization and granular file streaming. This architecture allows generator extensions to declare their source repositories while giving users control over authentication, file filtering, and download strategies.
Core HuggingFace Integration Architecture
Modly splits its HuggingFace integration into two complementary workflows. The automatic download system clones entire repositories when generators initialize, while the streaming API provides on-demand access to specific files without full local mirroring.
Both paths rely on the official huggingface_hub SDK and support token-based authentication through environment variables or query parameters. The integration centers on three core files: api/services/generators/base.py for automatic downloads, api/routers/model.py for streaming endpoints, and api/services/generator_registry.py for manifest injection.
Automatic Weight Download with snapshot_download
When a generator instance first runs, Modly checks whether the model exists locally. If is_downloaded() returns false, the BaseGenerator class triggers _auto_download() defined in api/services/generators/base.py (lines 44-68).
This method constructs an ignore pattern from the generator's manifest and calls snapshot_download():
# From api/services/generators/base.py
def _auto_download(self):
target_dir = self._get_model_path()
target_dir.mkdir(parents=True, exist_ok=True)
ignore_patterns = self.hf_skip_prefixes if hasattr(self, 'hf_skip_prefixes') else None
snapshot_download(
repo_id=self.hf_repo,
local_dir=str(target_dir),
ignore_patterns=ignore_patterns,
token=self._get_hf_token()
)
The hf_repo identifier and skip lists are injected at runtime by api/services/generator_registry.py, which reads the extension's manifest.json during initialization. This ensures the generator knows exactly which HuggingFace repository to mirror and which files to exclude.
Streaming File Downloads via REST API
For scenarios requiring direct file access without full repository cloning, Modly exposes the /model/{model_id}/download endpoint in api/routers/model.py (lines 85-92). This endpoint implements _download_file_streamed() (lines 81-120) to handle chunked transfers with resume support.
The streaming workflow follows these steps:
- Enumerate files: Uses
list_repo_files(repo_id, token=hf_token)to retrieve the repository index. - Apply filters: Respects
hf_skip_prefixesandhf_include_prefixesdefined in the manifest (lines 56-69). - Generate URLs: Builds download links via
hf_hub_url(repo_id, filename). - Stream chunks: Downloads bytes incrementally while reporting progress through Server-Sent Events (SSE).
# Conceptual flow from api/routers/model.py
files = list_repo_files(repo_id, token=token)
filtered = [f for f in files if not any(f.startswith(p) for p in skip_prefixes)]
for filename in filtered:
url = hf_hub_url(repo_id, filename)
yield from _download_file_streamed(url, token)
Authentication and Token Management
Modly supports multiple authentication vectors for private HuggingFace repositories. The system checks for tokens in the following priority order:
- Query parameter:
?token=...passed to the download endpoint - Environment variables:
HUGGING_FACE_HUB_TOKENorHF_TOKEN
Token extraction occurs in api/routers/model.py at lines 71-73 for the endpoint handler and lines 100-106 within the streaming function. The token is passed directly to all huggingface_hub SDK calls, ensuring gated model access works seamlessly.
Selective File Filtering
Extensions control their download footprint through manifest.json entries:
hf_skip_prefixes: Array of path prefixes to exclude (e.g.,["samples/", "README.md"])hf_include_prefixes: Array of path prefixes to exclusively include (e.g.,["weights/"])
These lists are read in api/routers/model.py (lines 56-69) and converted into ignore_patterns for snapshot_download or used to filter the file list in streaming operations. This selective approach reduces disk usage and download times by fetching only necessary model artifacts.
{
"hf_repo": "username/example-model",
"hf_skip_prefixes": ["samples/", "README.md"],
"hf_include_prefixes": ["weights/"]
}
Usage Examples
Triggering Automatic Download via Python API
When implementing a generator class, no explicit download code is required. The base class handles synchronization automatically:
from modly.api.services.generators.base import BaseGenerator
class MyModel(BaseGenerator):
MODEL_ID = "example-model"
# hf_repo injected from manifest at runtime
generator = MyModel()
# This triggers _auto_download() if model not present
generator.generate(prompt="A sunny beach", progress_cb=lambda p, s: print(p, s))
Requesting Streaming Downloads via HTTP
Clients can download specific filtered files without local storage requirements:
curl -N "http://localhost:8000/model/example-model/download?token=$HF_TOKEN"
The response streams Server-Sent Events reporting per-file progress and handles resumable transfers internally.
Summary
- Modly integrates HuggingFace through the
huggingface_hubSDK, usingsnapshot_downloadfor bulk operations andlist_repo_fileswithhf_hub_urlfor streaming. - Automatic downloads occur in
api/services/generators/base.pyvia the_auto_download()method, which clones repositories based on manifest-declaredhf_repovalues. - Streaming endpoints live in
api/routers/model.py, implementing chunked file transfers with support for resume capabilities and progress callbacks. - Authentication supports both query parameters and environment variables, with token handling at lines 71-73 and 100-106 of the router file.
- File filtering uses
hf_skip_prefixesandhf_include_prefixesdefined in extension manifests, parsed at lines 56-69 ofmodel.pyto minimize unnecessary downloads.
Frequently Asked Questions
How does Modly authenticate with HuggingFace for private models?
Modly accepts HuggingFace tokens via the token query parameter in download requests or through the HUGGING_FACE_HUB_TOKEN and HF_TOKEN environment variables. The system extracts these credentials in api/routers/model.py and passes them to every huggingface_hub SDK call, ensuring gated repository access works for both automatic and streaming downloads.
What is the difference between automatic and streaming downloads in Modly?
Automatic downloads use snapshot_download in api/services/generators/base.py to mirror entire HuggingFace repositories locally when a generator first initializes. Streaming downloads, handled by /model/{model_id}/download in api/routers/model.py, enumerate specific files via list_repo_files and transfer them chunk-by-chunk without persisting a full local copy, ideal for ephemeral model access.
Can I exclude specific files from HuggingFace downloads?
Yes. Define hf_skip_prefixes in your extension's manifest.json to ignore specific paths, or use hf_include_prefixes to whitelist only necessary directories. Modly applies these filters in api/routers/model.py (lines 56-69) during both bulk snapshot_download operations and individual file streaming, reducing bandwidth and storage requirements.
Where is the HuggingFace integration logic located in the Modly codebase?
The integration spans three primary files: api/services/generators/base.py contains the BaseGenerator class with _auto_download() for automatic weight syncing; api/routers/model.py implements the streaming download endpoint and token management; and api/services/generator_registry.py injects repository metadata from extension manifests into generator instances at runtime.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →