Handling Large Binary Data Efficiently with sdata's Blob Class

The Blob class in the sdata package provides a memory-efficient, lazy-loading wrapper for binary assets that keeps large payloads off-heap until explicitly accessed while supporting unified storage access via fsspec.

The sdata library solves the challenge of managing massive binary files—PDFs, images, video, or scientific datasets—without exhausting system memory. Implemented in sdata/sclass/blob.py, the Blob class inherits from Base and uses a descriptor-based architecture to separate lightweight metadata from heavy binary payloads. This design enables handling large binary data efficiently across local filesystems, cloud storage, and archive formats through a single Pythonic interface.

Core Architecture and Lazy Loading

The Blob class minimizes memory footprint by storing only a content descriptor rather than the actual bytes. This descriptor tracks three critical fields in self.data['content']: type (either 'bytes' or 'uri'), filetype (e.g., 'pdf', 'png'), and value (Base64-encoded string or URI).

Content Descriptor Structure

When you call set_content(), the class populates this descriptor without necessarily loading data into RAM:

from sdata import Blob

# Reference a 10GB file without loading it

blob = Blob(name="large_dataset", filetype="bin")
blob.set_content('uri', 's3://my-bucket/dataset.bin')

The actual binary payload remains on external storage until you access the content_bytes property. At that point, Blob dispatches to the appropriate loader: base64.b64decode for in-memory bytes or fsspec.open for URI-based resources.

The content_bytes Property

The lazy-loading mechanism lives in the content_bytes property implementation. For URI-type content, the class executes:

with fsspec.open(self.data['content']['value'], 'rb') as f:
    loaded_bytes = f.read()

This streams data on-demand, ensuring that merely creating or serializing a Blob object never triggers heavy I/O operations.

Multi-Protocol Storage Integration

The Blob class leverages fsspec to abstract disparate storage systems behind a unified open() API. This integration enables transparent access to local files, S3 buckets, GCS, Zip archives, and HTTP endpoints without changing application code.

URI-Based Access Patterns

When type='uri', the Blob class delegates existence checks and read operations to fsspec. The exists method calls fsspec.core.get_fs_token_paths to verify resource accessibility without downloading content:

blob = Blob(name="cloud_backup", filetype="tar.gz")
blob.set_content('uri', 's3://my-bucket/backup.tar.gz')

# Verify accessibility without network transfer

print(blob.exists())  # Delegates to underlying filesystem

Supported schemes include file://, s3://, gcs://, zip://, and http://, provided the corresponding fsspec extension packages (e.g., s3fs, gcsfs) are installed.

Integrity Checking and Metadata

Beyond storage abstraction, the Blob class provides cryptographic verification through streaming hash computation. The sha1 and md5 properties calculate digests without permanently caching the entire byte sequence in memory.

Streaming Hash Computation

The internal _update_hash method processes data in 64 kB chunks, making it safe to hash multi-gigabyte files:

blob = Blob(name="critical_data", filetype="iso")
blob.set_content('uri', '/mnt/data/large_image.iso')

# Compute SHA-1 without loading entire file

print(blob.sha1)  # Streams through 64kB buffers

Metadata Inheritance

As a subclass of Base (defined in sdata/base.py), Blob automatically integrates with sdata's metadata system. The DEFAULT_METADATA constant inside blob.py pre-defines fields like checksum, mime_type, and source_uri, ensuring provenance tracking accompanies every binary asset.

Serialization Without Payload Bloat

The to_dict() and from_dict() methods enable persistence while maintaining memory efficiency. When serializing, the class exports only the content descriptor—the Base64 string for small in-memory payloads or the URI string for external files—never the raw decoded bytes.


# Create and serialize

blob = Blob(name="document", filetype="pdf")
blob.set_content('bytes', b'%PDF-1.4...')
descriptor = blob.to_dict()  # JSON-safe, no raw bytes

# Rehydrate later

restored = Blob.from_dict(descriptor)
assert restored.content_bytes == b'%PDF-1.4...'

This approach guarantees that databases, REST APIs, or message queues can store Blob references without suffering payload size explosions from embedded binary data.

Practical Implementation Examples

Referencing Large External Files

Use URI mode to handle files too large for system memory:

from sdata import Blob
import pathlib

path = pathlib.Path("/data/500mb_dataset.bin").as_uri()
blob = Blob(name="big_data", filetype="bin")
blob.set_content('uri', path)

# Check existence without I/O

print(f"File accessible: {blob.exists()}")

# Load only when necessary

data = blob.content_bytes  # Triggered read

Working with S3 Objects

After installing fsspec and s3fs, access cloud storage identically to local files:

from sdata import Blob

blob = Blob(name="s3_asset", filetype="mp4")
blob.set_content('uri', 's3://production-bucket/video.mp4')

# Stream content on demand

video_bytes = blob.content_bytes

Embedding Small In-Memory Data

For small payloads under a few megabytes, embed directly using bytes mode:

pdf_bytes = b"%PDF-1.4\n1 0 obj\n<< /Type /Catalog >>\nendobj"
blob = Blob(name="sample", filetype="pdf")
blob.set_content('bytes', pdf_bytes)

print(f"SHA-1: {blob.sha1}")

Summary

  • Lazy Loading: The Blob class in sdata/sclass/blob.py keeps large binaries off-heap until content_bytes is accessed, preventing memory exhaustion.
  • Protocol Agnostic: Via fsspec, the class supports file://, s3://, gcs://, and other schemes through a unified interface.
  • Streaming Integrity: Hash computation via sha1 and md5 properties uses 64 kB chunks to handle multi-gigabyte files safely.
  • Lightweight Serialization: The to_dict() method persists only descriptors (Base64 or URIs), making objects safe for database storage and API transmission.
  • Metadata Rich: Automatic provenance tracking through the Base class integration stores checksums, MIME types, and source URIs.

Frequently Asked Questions

How does sdata's Blob class prevent memory issues with large files?

The Blob class stores only a lightweight content descriptor containing the URI or Base64 reference, not the actual bytes. Memory allocation occurs only when you access the content_bytes property, which streams data through fsspec or decodes Base64 on demand. This architecture allows applications to reference terabyte-scale files while consuming only kilobytes of RAM for the Blob object itself.

What storage backends are compatible with the Blob class?

Any storage system supported by fsspec works transparently, including local filesystems (file://), Amazon S3 (s3://), Google Cloud Storage (gcs://), Zip archives (zip://), and HTTP/HTTPS endpoints. You must install the corresponding optional dependencies (e.g., s3fs, gcsfs) for cloud protocols, but the Blob API remains identical across all backends.

How do I verify file integrity without loading the entire blob?

Access the sha1 or md5 properties to trigger streaming hash computation. The internal _update_hash method reads the file in 64 kB chunks, computes the digest, and discards each buffer immediately. This verifies integrity against the stored checksum metadata field without ever holding the complete file in memory.

Can Blob objects be safely stored in databases or JSON APIs?

Yes. The to_dict() method serializes only the content descriptor—either a Base64 string (for small embedded data) or a URI string (for external storage)—ensuring the JSON representation remains compact. The from_dict() class method reconstructs the Blob with identical lazy-loading behavior, making the class ideal for ORM mapping or REST payload transmission.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →