Handling Large Binary Data Efficiently with sdata's Blob Class
The Blob class in the sdata package provides a memory-efficient, lazy-loading wrapper for binary assets that keeps large payloads off-heap until explicitly accessed while supporting unified storage access via fsspec.
The sdata library solves the challenge of managing massive binary files—PDFs, images, video, or scientific datasets—without exhausting system memory. Implemented in sdata/sclass/blob.py, the Blob class inherits from Base and uses a descriptor-based architecture to separate lightweight metadata from heavy binary payloads. This design enables handling large binary data efficiently across local filesystems, cloud storage, and archive formats through a single Pythonic interface.
Core Architecture and Lazy Loading
The Blob class minimizes memory footprint by storing only a content descriptor rather than the actual bytes. This descriptor tracks three critical fields in self.data['content']: type (either 'bytes' or 'uri'), filetype (e.g., 'pdf', 'png'), and value (Base64-encoded string or URI).
Content Descriptor Structure
When you call set_content(), the class populates this descriptor without necessarily loading data into RAM:
from sdata import Blob
# Reference a 10GB file without loading it
blob = Blob(name="large_dataset", filetype="bin")
blob.set_content('uri', 's3://my-bucket/dataset.bin')
The actual binary payload remains on external storage until you access the content_bytes property. At that point, Blob dispatches to the appropriate loader: base64.b64decode for in-memory bytes or fsspec.open for URI-based resources.
The content_bytes Property
The lazy-loading mechanism lives in the content_bytes property implementation. For URI-type content, the class executes:
with fsspec.open(self.data['content']['value'], 'rb') as f:
loaded_bytes = f.read()
This streams data on-demand, ensuring that merely creating or serializing a Blob object never triggers heavy I/O operations.
Multi-Protocol Storage Integration
The Blob class leverages fsspec to abstract disparate storage systems behind a unified open() API. This integration enables transparent access to local files, S3 buckets, GCS, Zip archives, and HTTP endpoints without changing application code.
URI-Based Access Patterns
When type='uri', the Blob class delegates existence checks and read operations to fsspec. The exists method calls fsspec.core.get_fs_token_paths to verify resource accessibility without downloading content:
blob = Blob(name="cloud_backup", filetype="tar.gz")
blob.set_content('uri', 's3://my-bucket/backup.tar.gz')
# Verify accessibility without network transfer
print(blob.exists()) # Delegates to underlying filesystem
Supported schemes include file://, s3://, gcs://, zip://, and http://, provided the corresponding fsspec extension packages (e.g., s3fs, gcsfs) are installed.
Integrity Checking and Metadata
Beyond storage abstraction, the Blob class provides cryptographic verification through streaming hash computation. The sha1 and md5 properties calculate digests without permanently caching the entire byte sequence in memory.
Streaming Hash Computation
The internal _update_hash method processes data in 64 kB chunks, making it safe to hash multi-gigabyte files:
blob = Blob(name="critical_data", filetype="iso")
blob.set_content('uri', '/mnt/data/large_image.iso')
# Compute SHA-1 without loading entire file
print(blob.sha1) # Streams through 64kB buffers
Metadata Inheritance
As a subclass of Base (defined in sdata/base.py), Blob automatically integrates with sdata's metadata system. The DEFAULT_METADATA constant inside blob.py pre-defines fields like checksum, mime_type, and source_uri, ensuring provenance tracking accompanies every binary asset.
Serialization Without Payload Bloat
The to_dict() and from_dict() methods enable persistence while maintaining memory efficiency. When serializing, the class exports only the content descriptor—the Base64 string for small in-memory payloads or the URI string for external files—never the raw decoded bytes.
# Create and serialize
blob = Blob(name="document", filetype="pdf")
blob.set_content('bytes', b'%PDF-1.4...')
descriptor = blob.to_dict() # JSON-safe, no raw bytes
# Rehydrate later
restored = Blob.from_dict(descriptor)
assert restored.content_bytes == b'%PDF-1.4...'
This approach guarantees that databases, REST APIs, or message queues can store Blob references without suffering payload size explosions from embedded binary data.
Practical Implementation Examples
Referencing Large External Files
Use URI mode to handle files too large for system memory:
from sdata import Blob
import pathlib
path = pathlib.Path("/data/500mb_dataset.bin").as_uri()
blob = Blob(name="big_data", filetype="bin")
blob.set_content('uri', path)
# Check existence without I/O
print(f"File accessible: {blob.exists()}")
# Load only when necessary
data = blob.content_bytes # Triggered read
Working with S3 Objects
After installing fsspec and s3fs, access cloud storage identically to local files:
from sdata import Blob
blob = Blob(name="s3_asset", filetype="mp4")
blob.set_content('uri', 's3://production-bucket/video.mp4')
# Stream content on demand
video_bytes = blob.content_bytes
Embedding Small In-Memory Data
For small payloads under a few megabytes, embed directly using bytes mode:
pdf_bytes = b"%PDF-1.4\n1 0 obj\n<< /Type /Catalog >>\nendobj"
blob = Blob(name="sample", filetype="pdf")
blob.set_content('bytes', pdf_bytes)
print(f"SHA-1: {blob.sha1}")
Summary
- Lazy Loading: The
Blobclass insdata/sclass/blob.pykeeps large binaries off-heap untilcontent_bytesis accessed, preventing memory exhaustion. - Protocol Agnostic: Via
fsspec, the class supportsfile://,s3://,gcs://, and other schemes through a unified interface. - Streaming Integrity: Hash computation via
sha1andmd5properties uses 64 kB chunks to handle multi-gigabyte files safely. - Lightweight Serialization: The
to_dict()method persists only descriptors (Base64 or URIs), making objects safe for database storage and API transmission. - Metadata Rich: Automatic provenance tracking through the
Baseclass integration stores checksums, MIME types, and source URIs.
Frequently Asked Questions
How does sdata's Blob class prevent memory issues with large files?
The Blob class stores only a lightweight content descriptor containing the URI or Base64 reference, not the actual bytes. Memory allocation occurs only when you access the content_bytes property, which streams data through fsspec or decodes Base64 on demand. This architecture allows applications to reference terabyte-scale files while consuming only kilobytes of RAM for the Blob object itself.
What storage backends are compatible with the Blob class?
Any storage system supported by fsspec works transparently, including local filesystems (file://), Amazon S3 (s3://), Google Cloud Storage (gcs://), Zip archives (zip://), and HTTP/HTTPS endpoints. You must install the corresponding optional dependencies (e.g., s3fs, gcsfs) for cloud protocols, but the Blob API remains identical across all backends.
How do I verify file integrity without loading the entire blob?
Access the sha1 or md5 properties to trigger streaming hash computation. The internal _update_hash method reads the file in 64 kB chunks, computes the digest, and discards each buffer immediately. This verifies integrity against the stored checksum metadata field without ever holding the complete file in memory.
Can Blob objects be safely stored in databases or JSON APIs?
Yes. The to_dict() method serializes only the content descriptor—either a Base64 string (for small embedded data) or a URI string (for external storage)—ensuring the JSON representation remains compact. The from_dict() class method reconstructs the Blob with identical lazy-loading behavior, making the class ideal for ORM mapping or REST payload transmission.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →