# Common Issues Running TileDB-VCF on Cloud Storage (S3, Azure, GCS): Troubleshooting Guide

> Troubleshoot common TileDB-VCF issues on cloud storage S3 Azure GCS. Learn essential configurations for credentials URI formatting and concurrency to optimize performance and avoid errors.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: how-to-guide
- Published: 2026-05-14

---

**TileDB-VCF requires explicit TileDB configuration for cloud credentials, proper URI formatting with trailing slashes, and tuned concurrency settings to avoid rate limits when accessing genomic variant data on S3, Azure Blob, or Google Cloud Storage.**

When running TileDB-VCF against remote object stores, the library delegates all storage operations to the TileDB-Python configuration layer, which abstracts S3, Azure, and GCS behind a unified virtual file system (VFS). While this architecture enables seamless cloud deployment, several common configuration pitfalls can trigger authentication errors, permission denials, or silent query failures. This guide synthesizes solutions from the `K-Dense-AI/scientific-agent-skills` repository to resolve the most frequent issues encountered when executing variant queries against cloud-backed datasets.

## Authentication and Credential Configuration

TileDB-VCF does not automatically read standard cloud credential files like `~/.aws/credentials` on every platform. Instead, it relies on parameters passed directly through the **TileDB Config object**.

### Explicit TileDB Config vs. Standard AWS Files

You must populate the `tiledb.Config` dictionary with specific keys before initializing your dataset context. According to the Cloud Storage Integration section in [`scientific-skills/tiledbvcf/SKILL.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/tiledbvcf/SKILL.md) (line 214), the required parameters are:

```python
import tiledb, tiledbvcf

cfg = tiledb.Config()
cfg["vfs.s3.aws_access_key_id"] = "<YOUR_ACCESS_KEY>"
cfg["vfs.s3.aws_secret_access_key"] = "<YOUR_SECRET>"

ctx = tiledb.Ctx(cfg)
vcf = tiledbvcf.VcfDataset(uri="s3://my-bucket/vcf/", ctx=ctx)

```

### Temporary Credentials and Session Tokens

When using temporary credentials from AWS STS or similar services, include the session token in the configuration:

```python
cfg["vfs.s3.aws_session_token"] = "<TOKEN>"

```

Without this parameter, TileDB-VCF will authenticate successfully but receive 403 errors on subsequent operations when the token expires or is required by the IAM policy.

## URI Scheme and Path Formatting

Incorrect URI construction is the most common source of "Array not found" errors when migrating from local filesystems to cloud storage.

### The Trailing Slash Requirement

TileDB-VCF interprets the URI as a **directory** path. A missing trailing slash causes the library to treat the bucket root or parent path as the array itself, resulting in schema detection failures.

```python

# Correct - includes trailing slash

uri = "s3://my-bucket/vcf/"

# Incorrect - leads to "Array not found"

uri = "s3://my-bucket/vcf"

```

### Endpoint Overrides for S3-Compatible Services

For S3-compatible services like Wasabi, MinIO, or Azure Blob with S3 API enabled, you must specify the endpoint override to prevent TileDB from routing requests to the default AWS endpoint:

```python
cfg["vfs.s3.endpoint_override"] = "https://my-wasabi-endpoint.com"

```

Without this configuration, TileDB-VCF will attempt to resolve the bucket against `s3.amazonaws.com`, resulting in 404 or 403 errors even with valid credentials.

## Permission and Network Issues

Even with correct credentials, network policies and TLS configurations can block access.

### IAM Policies and Bucket Permissions

The IAM role or user must possess explicit permissions for `s3:GetObject`, `s3:PutObject`, and `s3:ListBucket` actions. TileDB-VCF surfaces generic "Permission denied" errors when any of these actions are restricted by bucket policies. Grant the least-privilege subset that includes these specific actions, or use managed policies like `AmazonS3FullAccess` for development environments.

### TLS/SSL Verification in Private Clouds

Private cloud deployments using self-signed certificates will trigger TLS validation failures. Disable verification only for trusted internal networks:

```python
cfg["vfs.s3.ssl_verify"] = "false"

```

**Warning:** Never disable SSL verification in production public cloud environments.

## Performance and Query Optimization

Cloud providers throttle HTTP request rates, which conflicts with TileDB's parallel query engine that generates thousands of concurrent small reads against cloud objects.

### Rate Limiting and Concurrent Requests

Reduce the parallelism to avoid 503 Slow Down errors:

```python
cfg["vfs.s3.concurrent_requests"] = 4  # Default is often 8-16

cfg["vfs.s3.retry_count"] = 5          # Enable automatic retries

```

### Caching Configuration

Enable the S3 request cache to minimize redundant GET requests during repetitive queries:

```python
cfg["vfs.s3.enable_cache"] = "true"

```

This is particularly effective when scanning genomic regions with overlapping tile extents across multiple query iterations.

## Data Integrity and Coordinate Systems

TileDB-VCF stores variant coordinates in **1-based** format, consistent with the VCF specification. However, many downstream bioinformatics tools (particularly BED file processors) assume 0-based indexing.

### 1-Based vs 0-Based Coordinate Confusion

If your queries return empty result sets despite valid region strings, verify your coordinate system:

```python

# Correct: TileDB-VCF expects 1-based VCF-style coordinates

result = vcf.query(region="chr1:100000-200000")

# Incorrect: 0-based coordinates will miss variants

# Convert by adding 1 to start positions before querying

```

Always convert 0-based indices by adding 1 before passing them to `vcf.query()` to ensure variant records are retrieved correctly from the sparse array.

## Version Compatibility and TileDB Cloud

Version mismatches between the core TileDB library and TileDB-VCF can produce "Unsupported array schema" errors or silent corruption when writing to cloud storage.

### Package Version Mismatches

Newer TileDB-VCF releases require TileDB-Python 0.12 or higher. Upgrade both packages simultaneously to ensure compatibility:

```bash
pip install -U tiledb tiledbvcf

```

### TileDB Cloud URI Registration

Attempting to access a locally-created dataset using a `tiledb://` URI without registration results in "Dataset not found on TileDB-Cloud" errors. The dataset must be registered via the TileDB Cloud UI or CLI before cloud URIs function. For pure-cloud deployments, initialize the dataset directly with the `tiledb://` scheme from creation rather than migrating from local storage.

## Practical Implementation Examples

### Creating a Dataset on Amazon S3

```python
import tiledb, tiledbvcf

cfg = tiledb.Config()
cfg["vfs.s3.aws_access_key_id"] = "<ACCESS_KEY>"
cfg["vfs.s3.aws_secret_access_key"] = "<SECRET_KEY>"
cfg["vfs.s3.region"] = "us-west-2"

ctx = tiledb.Ctx(cfg)
uri = "s3://my-genomics-bucket/vcf-dataset/"
tiledbvcf.create_dataset(uri=uri, ctx=ctx)

```

### Ingesting VCF Files from Google Cloud Storage

```python
vcf = tiledbvcf.VcfDataset(uri=uri, ctx=ctx)
vcf.store("gs://my-bucket/samples/sample1.vcf.gz")

```

### Querying with Optimized Concurrency

```python
cfg["vfs.s3.concurrent_requests"] = 2
cfg["vfs.s3.retry_count"] = 5

vcf = tiledbvcf.VcfDataset(uri=uri, ctx=tiledb.Ctx(cfg))
result = vcf.query(
    region="chr1:100000-200000",
    fields=["sample_name", "genotype"],
    return_all=True
)

```

## Summary

- **Configure credentials explicitly** using `tiledb.Config` keys (`vfs.s3.aws_access_key_id`, etc.) rather than relying on environment variables or AWS credential files.
- **Always terminate cloud URIs with a trailing slash** to indicate directory paths and prevent "Array not found" errors.
- **Specify endpoint overrides** for S3-compatible services (Wasabi, MinIO, Azure S3 API) to route requests correctly.
- **Tune concurrency** with `vfs.s3.concurrent_requests` to avoid rate limiting on large-scale genomic scans.
- **Maintain version parity** between `tiledb` and `tiledbvcf` packages to prevent schema compatibility issues.
- **Use 1-based coordinates** exclusively when querying to match VCF specifications and the sparse array storage layout.

## Frequently Asked Questions

### Why does TileDB-VCF fail to find my array on S3 when local files work fine?

TileDB-VCF requires a trailing slash at the end of cloud URIs to distinguish between the array directory and its parent path. Without the slash (e.g., `s3://bucket/vcf` instead of `s3://bucket/vcf/`), the library attempts to open the parent container as an array, resulting in a "Array not found" error. Additionally, ensure you have passed a configured `tiledb.Ctx` with valid credentials, as cloud storage cannot use standard filesystem authentication.

### How do I fix "Permission denied" errors when my AWS credentials are correct?

Even with valid access keys, the IAM policy attached to those credentials must explicitly allow `s3:GetObject`, `s3:PutObject`, and `s3:ListBucket` actions on the target bucket. TileDB-VCF also requires `ListBucket` permissions for the parent path to enumerate fragments. If using temporary credentials, verify that `vfs.s3.aws_session_token` is configured in your TileDB config object, as omission of the session token commonly manifests as permission errors after initial authentication.

### What causes slow query performance on cloud storage, and how can I improve it?

TileDB's parallel query engine generates many concurrent HTTP requests to cloud object stores, which can trigger provider rate limits (HTTP 503 errors) and throttling. Reduce `vfs.s3.concurrent_requests` to 2-4 and enable `vfs.s3.enable_cache` to minimize redundant requests. The sparse array engine in TileDB-VCF (as documented in [`scientific-skills/tiledbvcf/SKILL.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/tiledbvcf/SKILL.md) at line 166) stores variants in a multi-dimensional layout that skips empty genomic regions, but this benefit is negated if the VFS layer is throttled by excessive concurrent requests.

### Can I use TileDB-VCF with Azure Blob Storage or Google Cloud Storage instead of S3?

Yes. TileDB-VCF uses the TileDB VFS abstraction layer, which supports Azure Blob (`azure://` URIs) and Google Cloud Storage (`gs://` URIs) using analogous configuration parameters (`vfs.azure.*` and `vfs.gcs.*` respectively). The same troubleshooting principles apply: use trailing slashes, configure explicit credentials, and set appropriate concurrency limits. For Azure, you may need to specify `vfs.azure.blob_endpoint` if using private endpoints or non-standard regions.