# Handling Large Binary Data Efficiently with sdata's Blob Class

> Efficiently handle large binary data with sdata's Blob class. This memory-efficient wrapper keeps payloads off-heap until accessed and supports unified storage via fsspec.

- Repository: [lepy/sdata](https://github.com/lepy/sdata)
- Tags: how-to-guide
- Published: 2026-03-05

---

**The `Blob` class in the `sdata` package provides a memory-efficient, lazy-loading wrapper for binary assets that keeps large payloads off-heap until explicitly accessed while supporting unified storage access via `fsspec`.**

The `sdata` library solves the challenge of managing massive binary files—PDFs, images, video, or scientific datasets—without exhausting system memory. Implemented in [`sdata/sclass/blob.py`](https://github.com/lepy/sdata/blob/main/sdata/sclass/blob.py), the `Blob` class inherits from `Base` and uses a descriptor-based architecture to separate lightweight metadata from heavy binary payloads. This design enables handling large binary data efficiently across local filesystems, cloud storage, and archive formats through a single Pythonic interface.

## Core Architecture and Lazy Loading

The `Blob` class minimizes memory footprint by storing only a **content descriptor** rather than the actual bytes. This descriptor tracks three critical fields in `self.data['content']`: `type` (either `'bytes'` or `'uri'`), `filetype` (e.g., `'pdf'`, `'png'`), and `value` (Base64-encoded string or URI).

### Content Descriptor Structure

When you call `set_content()`, the class populates this descriptor without necessarily loading data into RAM:

```python
from sdata import Blob

# Reference a 10GB file without loading it

blob = Blob(name="large_dataset", filetype="bin")
blob.set_content('uri', 's3://my-bucket/dataset.bin')

```

The actual binary payload remains on external storage until you access the `content_bytes` property. At that point, `Blob` dispatches to the appropriate loader: `base64.b64decode` for in-memory bytes or `fsspec.open` for URI-based resources.

### The content_bytes Property

The lazy-loading mechanism lives in the `content_bytes` property implementation. For URI-type content, the class executes:

```python
with fsspec.open(self.data['content']['value'], 'rb') as f:
    loaded_bytes = f.read()

```

This streams data on-demand, ensuring that merely creating or serializing a `Blob` object never triggers heavy I/O operations.

## Multi-Protocol Storage Integration

The `Blob` class leverages **`fsspec`** to abstract disparate storage systems behind a unified `open()` API. This integration enables transparent access to local files, S3 buckets, GCS, Zip archives, and HTTP endpoints without changing application code.

### URI-Based Access Patterns

When `type='uri'`, the `Blob` class delegates existence checks and read operations to `fsspec`. The `exists` method calls `fsspec.core.get_fs_token_paths` to verify resource accessibility without downloading content:

```python
blob = Blob(name="cloud_backup", filetype="tar.gz")
blob.set_content('uri', 's3://my-bucket/backup.tar.gz')

# Verify accessibility without network transfer

print(blob.exists())  # Delegates to underlying filesystem

```

Supported schemes include `file://`, `s3://`, `gcs://`, `zip://`, and `http://`, provided the corresponding `fsspec` extension packages (e.g., `s3fs`, `gcsfs`) are installed.

## Integrity Checking and Metadata

Beyond storage abstraction, the `Blob` class provides cryptographic verification through **streaming hash computation**. The `sha1` and `md5` properties calculate digests without permanently caching the entire byte sequence in memory.

### Streaming Hash Computation

The internal `_update_hash` method processes data in **64 kB chunks**, making it safe to hash multi-gigabyte files:

```python
blob = Blob(name="critical_data", filetype="iso")
blob.set_content('uri', '/mnt/data/large_image.iso')

# Compute SHA-1 without loading entire file

print(blob.sha1)  # Streams through 64kB buffers

```

### Metadata Inheritance

As a subclass of `Base` (defined in [`sdata/base.py`](https://github.com/lepy/sdata/blob/main/sdata/base.py)), `Blob` automatically integrates with `sdata`'s metadata system. The `DEFAULT_METADATA` constant inside [`blob.py`](https://github.com/lepy/sdata/blob/main/blob.py) pre-defines fields like `checksum`, `mime_type`, and `source_uri`, ensuring provenance tracking accompanies every binary asset.

## Serialization Without Payload Bloat

The `to_dict()` and `from_dict()` methods enable persistence while maintaining memory efficiency. When serializing, the class exports only the content descriptor—the Base64 string for small in-memory payloads or the URI string for external files—never the raw decoded bytes.

```python

# Create and serialize

blob = Blob(name="document", filetype="pdf")
blob.set_content('bytes', b'%PDF-1.4...')
descriptor = blob.to_dict()  # JSON-safe, no raw bytes

# Rehydrate later

restored = Blob.from_dict(descriptor)
assert restored.content_bytes == b'%PDF-1.4...'

```

This approach guarantees that databases, REST APIs, or message queues can store `Blob` references without suffering payload size explosions from embedded binary data.

## Practical Implementation Examples

### Referencing Large External Files

Use URI mode to handle files too large for system memory:

```python
from sdata import Blob
import pathlib

path = pathlib.Path("/data/500mb_dataset.bin").as_uri()
blob = Blob(name="big_data", filetype="bin")
blob.set_content('uri', path)

# Check existence without I/O

print(f"File accessible: {blob.exists()}")

# Load only when necessary

data = blob.content_bytes  # Triggered read

```

### Working with S3 Objects

After installing `fsspec` and `s3fs`, access cloud storage identically to local files:

```python
from sdata import Blob

blob = Blob(name="s3_asset", filetype="mp4")
blob.set_content('uri', 's3://production-bucket/video.mp4')

# Stream content on demand

video_bytes = blob.content_bytes

```

### Embedding Small In-Memory Data

For small payloads under a few megabytes, embed directly using bytes mode:

```python
pdf_bytes = b"%PDF-1.4\n1 0 obj\n<< /Type /Catalog >>\nendobj"
blob = Blob(name="sample", filetype="pdf")
blob.set_content('bytes', pdf_bytes)

print(f"SHA-1: {blob.sha1}")

```

## Summary

- **Lazy Loading**: The `Blob` class in [`sdata/sclass/blob.py`](https://github.com/lepy/sdata/blob/main/sdata/sclass/blob.py) keeps large binaries off-heap until `content_bytes` is accessed, preventing memory exhaustion.
- **Protocol Agnostic**: Via `fsspec`, the class supports `file://`, `s3://`, `gcs://`, and other schemes through a unified interface.
- **Streaming Integrity**: Hash computation via `sha1` and `md5` properties uses 64 kB chunks to handle multi-gigabyte files safely.
- **Lightweight Serialization**: The `to_dict()` method persists only descriptors (Base64 or URIs), making objects safe for database storage and API transmission.
- **Metadata Rich**: Automatic provenance tracking through the `Base` class integration stores checksums, MIME types, and source URIs.

## Frequently Asked Questions

### How does sdata's Blob class prevent memory issues with large files?

The `Blob` class stores only a lightweight content descriptor containing the URI or Base64 reference, not the actual bytes. Memory allocation occurs only when you access the `content_bytes` property, which streams data through `fsspec` or decodes Base64 on demand. This architecture allows applications to reference terabyte-scale files while consuming only kilobytes of RAM for the `Blob` object itself.

### What storage backends are compatible with the Blob class?

Any storage system supported by `fsspec` works transparently, including local filesystems (`file://`), Amazon S3 (`s3://`), Google Cloud Storage (`gcs://`), Zip archives (`zip://`), and HTTP/HTTPS endpoints. You must install the corresponding optional dependencies (e.g., `s3fs`, `gcsfs`) for cloud protocols, but the `Blob` API remains identical across all backends.

### How do I verify file integrity without loading the entire blob?

Access the `sha1` or `md5` properties to trigger streaming hash computation. The internal `_update_hash` method reads the file in 64 kB chunks, computes the digest, and discards each buffer immediately. This verifies integrity against the stored `checksum` metadata field without ever holding the complete file in memory.

### Can Blob objects be safely stored in databases or JSON APIs?

Yes. The `to_dict()` method serializes only the content descriptor—either a Base64 string (for small embedded data) or a URI string (for external storage)—ensuring the JSON representation remains compact. The `from_dict()` class method reconstructs the `Blob` with identical lazy-loading behavior, making the class ideal for ORM mapping or REST payload transmission.