# How to Contribute Experimentally-Validated Simulation Data to the SQuADDS HuggingFace Dataset

> Learn to contribute experimentally-validated simulation data to SQuADDS HuggingFace dataset using the ExistingConfigData API for structured JSON submissions. Enhance the community dataset.

- Repository: [Levenson-Falk Lab/squadds](https://github.com/lfl-lab/squadds)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Use the `ExistingConfigData` API from [`squadds/database/contributor.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/database/contributor.py) to prepare, validate, and submit structured JSON entries to the `SQuADDS/SQuADDS_DB` repository on Hugging Face.**

The SQuADDS (Superconducting Qubit And Device Design Suite) project maintains a community-curated dataset of quantum hardware designs and simulation results on the Hugging Face Hub. Contributing experimentally-validated simulation data to the SQuADDS HuggingFace dataset involves programmatically constructing validated entries using the Python API and submitting them via a Pull Request workflow.

## Prerequisites and Environment Setup

Before instantiating the contributor class, you must define your institutional metadata and API credentials. The system expects environment variables stored in a `.env` file located at the repository root, as specified in [`squadds/core/globals.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/core/globals.py) via `ENV_FILE_PATH`.

Create a `.env` file based on the template at [`.env.example`](https://github.com/lfl-lab/squadds/blob/master/.env.example) and populate the following variables:

```bash
GROUP_NAME="your_group"
PI_NAME="principal_investigator"
INSTITUTION="university_name"
USER_NAME="your_username"
CONTRIB_MISC="optional_metadata"
HUGGINGFACE_API_KEY="hf_..."
GITHUB_TOKEN="ghp_..."

```

These values are automatically ingested when the `ExistingConfigData` class initializes, ensuring every contribution carries proper attribution.

## The ExistingConfigData Contribution Workflow

The contribution pipeline is implemented in [`squadds/database/contributor.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/database/contributor.py). The `ExistingConfigData` class orchestrates the entire process from data entry to repository synchronization.

### Instantiate the Configuration Object

First, create an instance targeting a specific configuration, such as `qubit-TransmonCross-cap_matrix`. The constructor validates the config string against the available dataset configurations using `get_dataset_config_names`.

```python
from squadds.database.contributor import ExistingConfigData

cfg = "qubit-TransmonCross-cap_matrix"
data = ExistingConfigData(cfg)

```

### Populate Design and Simulation Data

The API provides discrete methods for each component of the entry. All inputs are checked against the schema returned by `get_config_schema` (defined in [`squadds/core/utils.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/core/utils.py)).

**Design information** via `add_design()`:

```python
design = {
    "design_tool": "qiskit_metal",
    "design_options": {
        "pos_x": "-1500um",
        "pos_y": "1200um",
        "orientation": "-90",
        "cross_width": "30um"
    }
}
data.add_design(design)

```

**Simulation results** via `add_sim_result(name, value, unit)`:

```python
data.add_sim_result("cross_to_ground", 157.6063, "fF")
data.add_sim_result("claw_to_ground", 101.24431, "fF")

```

**Simulation setup** via `add_sim_setup()`:

```python
sim_setup = {
    "setup": {
        "name": "sweep_setup",
        "freq_ghz": 5.0,
        "max_passes": 12
    },
    "simulator": "ANSYS HFSS",
    "renderer_options": {"Lj": "10nH", "Cj": 0}
}
data.add_sim_setup(sim_setup)

```

**Optional metadata** via `add_notes()`:

```python
data.add_notes({"message": "Experimental validation of the transmon capacitance matrix."})

```

### Validate Entry Structure and Content

The `validate()` method performs three distinct checks defined in [`contributor.py`](https://github.com/lfl-lab/squadds/blob/main/contributor.py):

1. **Structure validation** (`_validate_structure`): Ensures required top-level keys (`design`, `sim_options`, `sim_results`, `contributor`, `notes`) are present.
2. **Type validation** (`_validate_types`): Uses `validate_types` from [`squadds/core/utils.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/core/utils.py) to verify Python types match the schema.
3. **Content validation** (`_validate_content`): Compares the nested structure of `design_options` and `sim_options.setup` against a reference entry in the existing dataset.

```python
data.validate()  # Prints: "Structure validated...", "Types validated...", etc.

```

### Synchronize and Submit to Hugging Face

The final phase clones the `SQuADDS/SQuADDS_DB` repository, creates a new branch, and writes the JSON entry. The `contribute()` method combines `update_repo()` and `update_db()`:

```python
repo_path = "/tmp/SQuADDS_DB_clone"
data.contribute(repo_path)  # Prints: "Contribution ready for PR"

```

Under the hood, `update_repo(path_to_repo)` handles the `git clone` or `git pull`, while `update_db(path_to_repo)` appends the serialized dictionary (produced by `to_dict()`) to the appropriate `<config>.json` file with proper indentation. The actual `git push` is currently disabled in the source (the `upload_to_HF` call is commented out), so you must manually push the branch using your `GITHUB_TOKEN` and open a Pull Request on the Hugging Face Hub.

## Handling Batch Contributions

For parametric sweeps or multiple experimental runs, use the sweep mode. The `from_json()` method loads all `*.json` files from a directory, and `validate_sweep()` checks each entry individually.

```python
data = ExistingConfigData("qubit-TransmonCross-cap_matrix")
data.from_json("tutorials/examples/sweep_data/", is_sweep=True)
data.validate_sweep()
data.contribute("/tmp/SQuADDS_DB_clone", is_sweep=True)

```

This workflow is demonstrated comprehensively in [**Tutorial 3**](https://github.com/lfl-lab/squadds/blob/master/tutorials/Tutorial-3_Contributing_Validated_Simulation_Data_to_SQuADDS.ipynb), which walks through every method call and shows the final JSON structure.

## Key Source Files and Architecture

Understanding the following files helps debug validation errors or extend the contribution pipeline:

- **[`squadds/database/contributor.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/database/contributor.py)**: Core API containing `ExistingConfigData`, validation logic, and repository handling.
- **[`squadds/core/utils.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/core/utils.py)**: Helper functions for schema extraction (`get_config_schema`), type checking (`validate_types`), and caching.
- **[`squadds/core/globals.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/core/globals.py)**: Defines `ENV_FILE_PATH` used to locate the `.env` configuration.
- **`.env.example`**: Template for required contributor metadata and API tokens.

## Summary

- **Configure environment variables** in a root `.env` file to set contributor attribution and API access.
- **Use `ExistingConfigData`** from [`squadds/database/contributor.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/database/contributor.py) to instantiate a target configuration.
- **Populate entries** using `add_design()`, `add_sim_result()`, `add_sim_setup()`, and `add_notes()`, all of which enforce schema compliance.
- **Validate entries** with `validate()` to check structure, types, and content against existing dataset references.
- **Submit via `contribute()`**, which clones `SQuADDS/SQuADDS_DB`, writes the JSON, and prepares a branch for a Pull Request on Hugging Face.
- **Batch process** multiple entries using `from_json()` with `is_sweep=True` and `validate_sweep()`.

## Frequently Asked Questions

### What validation does the SQuADDS contribution API perform before allowing a submission?

The `validate()` method in [`squadds/database/contributor.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/database/contributor.py) runs three sequential checks: `_validate_structure` ensures required keys (`design`, `sim_options`, `sim_results`, `contributor`, `notes`) exist; `_validate_types` uses `validate_types` from [`squadds/core/utils.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/core/utils.py) to confirm Python types match the schema; and `_validate_content` compares the nested dictionary structure of your `design_options` and `sim_options.setup` against a reference entry to ensure field compatibility.

### How do I set up my credentials to contribute to the SQuADDS dataset?

Create a `.env` file in the repository root (as specified in [`squadds/core/globals.py`](https://github.com/lfl-lab/squadds/blob/main/squadds/core/globals.py) by `ENV_FILE_PATH`) containing `GROUP_NAME`, `PI_NAME`, `INSTITUTION`, `USER_NAME`, `HUGGINGFACE_API_KEY`, and `GITHUB_TOKEN`. The `ExistingConfigData` class automatically loads these variables to attribute the contribution and authenticate repository operations.

### Can I contribute multiple simulation results at once instead of one by one?

Yes. For batch contributions, use the `from_json()` method with `is_sweep=True` to load an entire directory of JSON files, then call `validate_sweep()` and `contribute(repo_path, is_sweep=True)`. This processes each entry through the same validation pipeline before updating the local repository clone.

### Where is the contributed data physically stored after I run the `contribute()` method?

The `contribute()` method serializes your entry to a dictionary via `to_dict()` and appends it to a JSON file named after the configuration (e.g., [`qubit-TransmonCross-cap_matrix.json`](https://github.com/lfl-lab/squadds/blob/main/qubit-TransmonCross-cap_matrix.json)) within your local clone of the `SQuADDS/SQuADDS_DB` repository. You must then manually push this branch to your Hugging Face fork and create a Pull Request, as the automated `git push` is currently commented out in the source code.