Common Issues Running TileDB-VCF on Cloud Storage (S3, Azure, GCS): Troubleshooting Guide
TileDB-VCF requires explicit TileDB configuration for cloud credentials, proper URI formatting with trailing slashes, and tuned concurrency settings to avoid rate limits when accessing genomic variant data on S3, Azure Blob, or Google Cloud Storage.
When running TileDB-VCF against remote object stores, the library delegates all storage operations to the TileDB-Python configuration layer, which abstracts S3, Azure, and GCS behind a unified virtual file system (VFS). While this architecture enables seamless cloud deployment, several common configuration pitfalls can trigger authentication errors, permission denials, or silent query failures. This guide synthesizes solutions from the K-Dense-AI/scientific-agent-skills repository to resolve the most frequent issues encountered when executing variant queries against cloud-backed datasets.
Authentication and Credential Configuration
TileDB-VCF does not automatically read standard cloud credential files like ~/.aws/credentials on every platform. Instead, it relies on parameters passed directly through the TileDB Config object.
Explicit TileDB Config vs. Standard AWS Files
You must populate the tiledb.Config dictionary with specific keys before initializing your dataset context. According to the Cloud Storage Integration section in scientific-skills/tiledbvcf/SKILL.md (line 214), the required parameters are:
import tiledb, tiledbvcf
cfg = tiledb.Config()
cfg["vfs.s3.aws_access_key_id"] = "<YOUR_ACCESS_KEY>"
cfg["vfs.s3.aws_secret_access_key"] = "<YOUR_SECRET>"
ctx = tiledb.Ctx(cfg)
vcf = tiledbvcf.VcfDataset(uri="s3://my-bucket/vcf/", ctx=ctx)
Temporary Credentials and Session Tokens
When using temporary credentials from AWS STS or similar services, include the session token in the configuration:
cfg["vfs.s3.aws_session_token"] = "<TOKEN>"
Without this parameter, TileDB-VCF will authenticate successfully but receive 403 errors on subsequent operations when the token expires or is required by the IAM policy.
URI Scheme and Path Formatting
Incorrect URI construction is the most common source of "Array not found" errors when migrating from local filesystems to cloud storage.
The Trailing Slash Requirement
TileDB-VCF interprets the URI as a directory path. A missing trailing slash causes the library to treat the bucket root or parent path as the array itself, resulting in schema detection failures.
# Correct - includes trailing slash
uri = "s3://my-bucket/vcf/"
# Incorrect - leads to "Array not found"
uri = "s3://my-bucket/vcf"
Endpoint Overrides for S3-Compatible Services
For S3-compatible services like Wasabi, MinIO, or Azure Blob with S3 API enabled, you must specify the endpoint override to prevent TileDB from routing requests to the default AWS endpoint:
cfg["vfs.s3.endpoint_override"] = "https://my-wasabi-endpoint.com"
Without this configuration, TileDB-VCF will attempt to resolve the bucket against s3.amazonaws.com, resulting in 404 or 403 errors even with valid credentials.
Permission and Network Issues
Even with correct credentials, network policies and TLS configurations can block access.
IAM Policies and Bucket Permissions
The IAM role or user must possess explicit permissions for s3:GetObject, s3:PutObject, and s3:ListBucket actions. TileDB-VCF surfaces generic "Permission denied" errors when any of these actions are restricted by bucket policies. Grant the least-privilege subset that includes these specific actions, or use managed policies like AmazonS3FullAccess for development environments.
TLS/SSL Verification in Private Clouds
Private cloud deployments using self-signed certificates will trigger TLS validation failures. Disable verification only for trusted internal networks:
cfg["vfs.s3.ssl_verify"] = "false"
Warning: Never disable SSL verification in production public cloud environments.
Performance and Query Optimization
Cloud providers throttle HTTP request rates, which conflicts with TileDB's parallel query engine that generates thousands of concurrent small reads against cloud objects.
Rate Limiting and Concurrent Requests
Reduce the parallelism to avoid 503 Slow Down errors:
cfg["vfs.s3.concurrent_requests"] = 4 # Default is often 8-16
cfg["vfs.s3.retry_count"] = 5 # Enable automatic retries
Caching Configuration
Enable the S3 request cache to minimize redundant GET requests during repetitive queries:
cfg["vfs.s3.enable_cache"] = "true"
This is particularly effective when scanning genomic regions with overlapping tile extents across multiple query iterations.
Data Integrity and Coordinate Systems
TileDB-VCF stores variant coordinates in 1-based format, consistent with the VCF specification. However, many downstream bioinformatics tools (particularly BED file processors) assume 0-based indexing.
1-Based vs 0-Based Coordinate Confusion
If your queries return empty result sets despite valid region strings, verify your coordinate system:
# Correct: TileDB-VCF expects 1-based VCF-style coordinates
result = vcf.query(region="chr1:100000-200000")
# Incorrect: 0-based coordinates will miss variants
# Convert by adding 1 to start positions before querying
Always convert 0-based indices by adding 1 before passing them to vcf.query() to ensure variant records are retrieved correctly from the sparse array.
Version Compatibility and TileDB Cloud
Version mismatches between the core TileDB library and TileDB-VCF can produce "Unsupported array schema" errors or silent corruption when writing to cloud storage.
Package Version Mismatches
Newer TileDB-VCF releases require TileDB-Python 0.12 or higher. Upgrade both packages simultaneously to ensure compatibility:
pip install -U tiledb tiledbvcf
TileDB Cloud URI Registration
Attempting to access a locally-created dataset using a tiledb:// URI without registration results in "Dataset not found on TileDB-Cloud" errors. The dataset must be registered via the TileDB Cloud UI or CLI before cloud URIs function. For pure-cloud deployments, initialize the dataset directly with the tiledb:// scheme from creation rather than migrating from local storage.
Practical Implementation Examples
Creating a Dataset on Amazon S3
import tiledb, tiledbvcf
cfg = tiledb.Config()
cfg["vfs.s3.aws_access_key_id"] = "<ACCESS_KEY>"
cfg["vfs.s3.aws_secret_access_key"] = "<SECRET_KEY>"
cfg["vfs.s3.region"] = "us-west-2"
ctx = tiledb.Ctx(cfg)
uri = "s3://my-genomics-bucket/vcf-dataset/"
tiledbvcf.create_dataset(uri=uri, ctx=ctx)
Ingesting VCF Files from Google Cloud Storage
vcf = tiledbvcf.VcfDataset(uri=uri, ctx=ctx)
vcf.store("gs://my-bucket/samples/sample1.vcf.gz")
Querying with Optimized Concurrency
cfg["vfs.s3.concurrent_requests"] = 2
cfg["vfs.s3.retry_count"] = 5
vcf = tiledbvcf.VcfDataset(uri=uri, ctx=tiledb.Ctx(cfg))
result = vcf.query(
region="chr1:100000-200000",
fields=["sample_name", "genotype"],
return_all=True
)
Summary
- Configure credentials explicitly using
tiledb.Configkeys (vfs.s3.aws_access_key_id, etc.) rather than relying on environment variables or AWS credential files. - Always terminate cloud URIs with a trailing slash to indicate directory paths and prevent "Array not found" errors.
- Specify endpoint overrides for S3-compatible services (Wasabi, MinIO, Azure S3 API) to route requests correctly.
- Tune concurrency with
vfs.s3.concurrent_requeststo avoid rate limiting on large-scale genomic scans. - Maintain version parity between
tiledbandtiledbvcfpackages to prevent schema compatibility issues. - Use 1-based coordinates exclusively when querying to match VCF specifications and the sparse array storage layout.
Frequently Asked Questions
Why does TileDB-VCF fail to find my array on S3 when local files work fine?
TileDB-VCF requires a trailing slash at the end of cloud URIs to distinguish between the array directory and its parent path. Without the slash (e.g., s3://bucket/vcf instead of s3://bucket/vcf/), the library attempts to open the parent container as an array, resulting in a "Array not found" error. Additionally, ensure you have passed a configured tiledb.Ctx with valid credentials, as cloud storage cannot use standard filesystem authentication.
How do I fix "Permission denied" errors when my AWS credentials are correct?
Even with valid access keys, the IAM policy attached to those credentials must explicitly allow s3:GetObject, s3:PutObject, and s3:ListBucket actions on the target bucket. TileDB-VCF also requires ListBucket permissions for the parent path to enumerate fragments. If using temporary credentials, verify that vfs.s3.aws_session_token is configured in your TileDB config object, as omission of the session token commonly manifests as permission errors after initial authentication.
What causes slow query performance on cloud storage, and how can I improve it?
TileDB's parallel query engine generates many concurrent HTTP requests to cloud object stores, which can trigger provider rate limits (HTTP 503 errors) and throttling. Reduce vfs.s3.concurrent_requests to 2-4 and enable vfs.s3.enable_cache to minimize redundant requests. The sparse array engine in TileDB-VCF (as documented in scientific-skills/tiledbvcf/SKILL.md at line 166) stores variants in a multi-dimensional layout that skips empty genomic regions, but this benefit is negated if the VFS layer is throttled by excessive concurrent requests.
Can I use TileDB-VCF with Azure Blob Storage or Google Cloud Storage instead of S3?
Yes. TileDB-VCF uses the TileDB VFS abstraction layer, which supports Azure Blob (azure:// URIs) and Google Cloud Storage (gs:// URIs) using analogous configuration parameters (vfs.azure.* and vfs.gcs.* respectively). The same troubleshooting principles apply: use trailing slashes, configure explicit credentials, and set appropriate concurrency limits. For Azure, you may need to specify vfs.azure.blob_endpoint if using private endpoints or non-standard regions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →