Configuring sdata for Cloud Storage Backends: MinIO, S3, and OwnCloud Integration Guide
sdata provides ready-to-use wrapper classes—MinioBucket for S3-compatible storage and OwnCloudVault for WebDAV—that enable seamless cloud storage integration through environment variables or explicit configuration.
The lepy/sdata library simplifies data serialization by offering native support for popular cloud storage protocols. Configuring sdata for cloud storage backends allows you to persist large binary assets and datasets directly to MinIO, Amazon S3-compatible services, or OwnCloud shares without modifying your core application logic. Both implementations follow a thin-wrapper pattern that delegates heavy lifting to underlying libraries while exposing Pythonic APIs for Path, bytes, and BinaryIO types.
MinIO and S3-Compatible Storage Configuration
The MinioBucket class in sdata/iolib/minio.py provides a complete abstraction over MinIO and other S3-compatible object stores, including Amazon S3, Wasabi, and DigitalOcean Spaces.
Authentication and Client Initialization
The MinioBucket.__init__ method establishes the client connection and guarantees bucket existence upon instantiation. You can supply credentials explicitly or rely on environment variables for secure deployment.
Configuration options:
- Explicit: Pass
access_keyandsecret_keydirectly to the constructor - Environment variables: Set
MINIO_ACCESS_KEY_IDandMINIO_SECRET_ACCESS_KEYin your environment; the class automatically falls back to these when credentials are omitted - Endpoint: Specify hostname and port (e.g.,
"minio.mycompany.com:9000") - Secure flag: Set
secure=Falsefor HTTP connections to local or internal MinIO instances
from sdata.iolib.minio import MinioBucket
# Initialize with explicit credentials or env var fallback
minio = MinioBucket(
endpoint="minio.mycompany.com:9000",
bucket="my-data",
access_key="my-key", # Optional: falls back to MINIO_ACCESS_KEY_ID
secret_key="my-secret", # Optional: falls back to MINIO_SECRET_ACCESS_KEY
secure=False
)
As implemented in the source code, the constructor automatically checks bucket_exists and creates the bucket if missing, ensuring the storage location is ready for immediate use.
Core Storage Operations
The class provides high-level methods that translate between Python types and the storage client API. According to sdata/iolib/minio.py, key operations include:
objectnames(): Returns a list of object names in the bucketlist_df(): Returns a pandas DataFrame with filename, size, modification time, and suffix (source: lines 81-85)upload(): Uploads local files with optionalforceparameter to control overwrite behaviorupload_bytes(): Streams binary data directly without temporary files (source: lines 97-108)download()anddelete(): Retrieve and remove objects by key
from pathlib import Path
# List objects as DataFrame for inspection
print(minio.list_df().head())
# Upload local file (force=False skips existing files)
minio.upload(Path("samples/data.csv"), force=False)
# Stream large in-memory payloads directly
payload = b"binary data" * 100_000
minio.upload_bytes(payload, "large_blob.bin", force=True)
# Download to specific path
minio.download("data.csv", Path("downloads/local_copy.csv"))
# Clean up
minio.delete("large_blob.bin")
OwnCloud WebDAV Integration
For OwnCloud integration, sdata provides OwnCloudVault in sdata/iolib/owncloudfs.py, which connects to public shares via the WebDAV protocol.
Connecting to Public Shares
The OwnCloudVault.__init__ method (source: lines 19-33) establishes the connection using three parameters: hostname, share_id, and optional password. Unlike MinIO, this backend requires no bucket creation step, as the public share URL is immutable.
from sdata.iolib.owncloudfs import OwnCloudVault
# Connect to a public OwnCloud share
vault = OwnCloudVault(
hostname="https://cloud.example.com",
share_id="abcdef1234567890", # Token from the public link
password="" # Leave empty if the share has no password
)
File Synchronization
The implementation supports bidirectional folder synchronization and single-file operations:
upload_file(): Uploads a single local file to a remote pathdownload_file(): Retrieves a remote file to a local pathsync_local_to_remote(): Recursively uploads an entire local directory tree (source: lines 64-74)sync_remote_to_local(): Downloads the complete remote folder structure to a local directory
from pathlib import Path
# Inspect remote structure
print(vault.list_remote("/"))
# Single file operations
vault.upload_file(Path("report.pdf"), "/reports/annual.pdf")
vault.download_file("/reports/annual.pdf", Path("local_report.pdf"))
# Batch synchronization
vault.sync_local_to_remote(Path("datasets/"), remote_folder="/backup/datasets/")
vault.sync_remote_to_local(remote_folder="/backup/datasets/", local_folder=Path("restored_data/"))
Integrating Cloud Backends with sdata Objects
Higher-level sdata objects accept a storage_backend parameter, enabling seamless persistence without protocol-specific code. The storage_backend hook is defined in the core classes (see sdata/base.py), allowing objects like Workbook to offload binary asset storage to the configured cloud provider.
from sdata import Workbook
from sdata.iolib.minio import MinioBucket
import pandas as pd
# Create workbook backed by MinIO
wb = Workbook(
storage_backend=MinioBucket("minio.mycompany.com:9000", "workbook-bucket")
)
# Add data (DataFrame stored locally; large binaries routed to bucket)
df = pd.DataFrame({"x": range(1000), "y": range(1000)})
wb.add_table(df, name="large_dataset")
# Persist—automatically uploads to configured backend
wb.save("analysis.sdata")
This pattern allows you to switch between MinIO and OwnCloud (or any custom backend implementing the storage interface) by changing a single line of configuration code.
Summary
MinioBucket(sdata/iolib/minio.py) wraps S3-compatible storage with automatic bucket creation, environment variable credential fallback (MINIO_ACCESS_KEY_ID/MINIO_SECRET_ACCESS_KEY), and pandas-friendly listing methods.OwnCloudVault(sdata/iolib/owncloudfs.py) provides WebDAV access to public OwnCloud shares with bidirectional recursive synchronization capabilities.- Both classes handle translation between Python native types (
Path,bytes,BinaryIO) and underlying library APIs, re-raisingS3Erroror WebDAV exceptions as standard PythonExceptionobjects. - Integration requires only passing the configured backend instance to sdata objects via the
storage_backendparameter, enabling cloud-native data serialization without application changes.
Frequently Asked Questions
How do I configure MinIO credentials securely in sdata?
You can pass access_key and secret_key directly to MinioBucket(), or set the MINIO_ACCESS_KEY_ID and MINIO_SECRET_ACCESS_KEY environment variables. The constructor automatically checks for environment variables when explicit credentials are omitted, allowing secure deployment in containerized environments without hardcoded secrets.
Does sdata support Amazon S3 directly, or only MinIO?
The MinioBucket class uses the standard MinIO client library, which is fully compatible with Amazon S3 and other S3-compatible services. Configure your AWS endpoint (e.g., "s3.amazonaws.com"), region, and IAM credentials exactly as you would for MinIO—no code changes are required to switch between providers.
Can I synchronize entire directories with OwnCloud using sdata?
Yes. The OwnCloudVault.sync_local_to_remote() method recursively uploads a local directory tree to the remote share, while sync_remote_to_local() performs the reverse operation. These methods handle directory traversal and file comparison automatically, as implemented in sdata/iolib/owncloudfs.py lines 64-74.
What error handling does sdata implement for cloud storage operations?
Both wrapper classes catch underlying library exceptions—S3Error for MinIO and WebDAV-specific errors for OwnCloud—and re-raise them as generic Exception instances with descriptive messages. This allows consistent try/except blocks in your application regardless of which storage backend is configured.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →