How model_identity and sha256_file Compose a Deterministic Request Hash in YuE

The YuE inference pipeline generates a deterministic request hash by combining a model-identity hash of all static artifact files with a separate runtime SHA of the Python source code, enabling precise cache invalidation and full reproducibility.

In the multimodal-art-projection/YuE repository, every model generation request is assigned a unique, deterministic identifier derived from both the model weights and the inference code itself. This system relies on two core functions in src/yue2/storage.py—sha256_file() and model_identity()—to fingerprint static model artifacts, while src/yue2/pipeline.py maintains a separate runtime SHA to track the executing code version.

Fingerprinting Model Artifacts with model_identity

The model_identity() function creates a cryptographic fingerprint of the entire model directory by recursively hashing every relevant file and aggregating the results into a single digest.

Hashing Individual Files with sha256_file

Inside src/yue2/storage.py, the sha256_file() utility computes SHA-256 digests using a memory-efficient chunked approach. It reads files in 8192-byte blocks to handle multi-gigabyte model weights without exhausting RAM.


# src/yue2/storage.py

def sha256_file(path: Path) -> str:
    """Return the SHA-256 digest of a file's bytes."""
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for chunk in iter(lambda: f.read(8192), b""):
            h.update(chunk)
    return h.hexdigest()

Aggregating Directory Contents

The model_identity() function walks the model directory, collecting hashes for all files including *.safetensors and config.json. It builds a mapping of relative paths to file metadata (SHA-256 and byte size), sorts the dictionary to ensure determinism, and passes the sorted JSON to the identity() helper which produces the final digest.


# src/yue2/storage.py

def model_identity(model_dir: Path) -> str:
    """
    Build a deterministic hash that represents the *entire* model directory.
    """
    # Collect per-file hashes

    entries = {
        str(p.relative_to(model_dir)): {
            "sha256": sha256_file(p),
            "bytes": p.stat().st_size,
        }
        for p in model_dir.rglob("*") if p.is_file()
    }
    # Serialize deterministically (sorted keys) and hash

    return identity(entries)  # identity → json.dumps(..., sort_keys=True) → sha256

Isolating Code Versions with the Runtime SHA

While model artifacts remain static after training, the inference pipeline in src/yue2/pipeline.py evolves through bug fixes and feature updates. The Pipeline class constructor computes self.runtime_sha256 by hashing all *.py files in its package directory, creating a separate fingerprint for the executing code.


# src/yue2/pipeline.py

class Pipeline:
    def __init__(self, model_dir: Path):
        # Model identity (static for a given model)

        self.model_hash = model_identity(model_dir)

        # Runtime SHA (hash of all *.py files in the pipeline package)

        self.runtime_sha256 = identity(
            {
                p.name: sha256_file(p)
                for p in sorted(Path(__file__).parent.glob("*.py"))
            }
        )

Composing the Deterministic Request Hash

The final request hash combines both components in the request_hash() method. The model identity and runtime SHA are concatenated and hashed together, producing a unique key that changes if either the weights or the code changes.


# src/yue2/pipeline.py

def request_hash(self) -> str:
    """
    Deterministic request hash = hash(model_hash + runtime_sha256)
    """
    return hashlib.sha256(
        (self.model_hash + self.runtime_sha256).encode()
    ).hexdigest()

Why the Runtime SHA Is Stored Separately

Separating the runtime SHA from the model identity provides three critical advantages for production deployments:

  • Efficient Cache Invalidation: When only the pipeline code changes, the system updates the cache key by swapping the runtime component without recomputing expensive model hashes across gigabytes of weights.
  • Auditability and Provenance: The runtime SHA precisely identifies which version of the inference code produced a given output, independent of the model weights, enabling exact reproduction of results.
  • Separation of Concerns: Model artifacts follow an immutable lifecycle after training, while runtime code undergoes frequent updates; independent tracking respects these distinct lifecycle patterns.

Summary

  • sha256_file() in src/yue2/storage.py computes file digests using chunked 8192-byte reads for memory efficiency when processing large model weights.
  • model_identity() aggregates these digests into a sorted JSON structure hashed by the identity() helper to fingerprint the entire model directory including *.safetensors and config.json.
  • src/yue2/pipeline.py computes a separate runtime_sha256 by hashing all Python source files in the pipeline package to track code versions.
  • The request_hash() method deterministically combines both values, enabling reproducible generation and granular cache management that distinguishes between model updates and code changes.

Frequently Asked Questions

What is the difference between model_identity and runtime_sha256?

model_identity() produces a hash representing the static model artifacts (weights and configurations) in the model directory, while runtime_sha256 represents the dynamic Python source code that executes the inference. The former changes when model files are modified; the latter changes when the pipeline implementation is updated.

How does sha256_file handle large model files without memory issues?

The function implements chunked reading using iter(lambda: f.read(8192), b""), which processes files in 8192-byte blocks rather than loading entire multi-gigabyte tensors into memory. This constant-memory approach ensures stable operation regardless of file size.

Why does YuE separate the runtime hash from the model identity instead of hashing everything together?

The separation allows the system to invalidate caches or track provenance at the appropriate granularity. When the pipeline code changes but the model weights remain identical, only the runtime SHA changes, avoiding the computational cost of re-hashing gigabytes of model data while still providing a unique cache key.

Where is the deterministic request hash actually used in the codebase?

The request hash appears in src/yue2/fast.py and related caching layers where it serves as a lookup key for storing and retrieving pre-computed model artifacts, ensuring that cached results are only returned when both the specific model version and the specific pipeline code version match exactly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →