How to Contribute Experimentally-Validated Simulation Data to the SQuADDS HuggingFace Dataset
Use the ExistingConfigData API from squadds/database/contributor.py to prepare, validate, and submit structured JSON entries to the SQuADDS/SQuADDS_DB repository on Hugging Face.
The SQuADDS (Superconducting Qubit And Device Design Suite) project maintains a community-curated dataset of quantum hardware designs and simulation results on the Hugging Face Hub. Contributing experimentally-validated simulation data to the SQuADDS HuggingFace dataset involves programmatically constructing validated entries using the Python API and submitting them via a Pull Request workflow.
Prerequisites and Environment Setup
Before instantiating the contributor class, you must define your institutional metadata and API credentials. The system expects environment variables stored in a .env file located at the repository root, as specified in squadds/core/globals.py via ENV_FILE_PATH.
Create a .env file based on the template at .env.example and populate the following variables:
GROUP_NAME="your_group"
PI_NAME="principal_investigator"
INSTITUTION="university_name"
USER_NAME="your_username"
CONTRIB_MISC="optional_metadata"
HUGGINGFACE_API_KEY="hf_..."
GITHUB_TOKEN="ghp_..."
These values are automatically ingested when the ExistingConfigData class initializes, ensuring every contribution carries proper attribution.
The ExistingConfigData Contribution Workflow
The contribution pipeline is implemented in squadds/database/contributor.py. The ExistingConfigData class orchestrates the entire process from data entry to repository synchronization.
Instantiate the Configuration Object
First, create an instance targeting a specific configuration, such as qubit-TransmonCross-cap_matrix. The constructor validates the config string against the available dataset configurations using get_dataset_config_names.
from squadds.database.contributor import ExistingConfigData
cfg = "qubit-TransmonCross-cap_matrix"
data = ExistingConfigData(cfg)
Populate Design and Simulation Data
The API provides discrete methods for each component of the entry. All inputs are checked against the schema returned by get_config_schema (defined in squadds/core/utils.py).
Design information via add_design():
design = {
"design_tool": "qiskit_metal",
"design_options": {
"pos_x": "-1500um",
"pos_y": "1200um",
"orientation": "-90",
"cross_width": "30um"
}
}
data.add_design(design)
Simulation results via add_sim_result(name, value, unit):
data.add_sim_result("cross_to_ground", 157.6063, "fF")
data.add_sim_result("claw_to_ground", 101.24431, "fF")
Simulation setup via add_sim_setup():
sim_setup = {
"setup": {
"name": "sweep_setup",
"freq_ghz": 5.0,
"max_passes": 12
},
"simulator": "ANSYS HFSS",
"renderer_options": {"Lj": "10nH", "Cj": 0}
}
data.add_sim_setup(sim_setup)
Optional metadata via add_notes():
data.add_notes({"message": "Experimental validation of the transmon capacitance matrix."})
Validate Entry Structure and Content
The validate() method performs three distinct checks defined in contributor.py:
- Structure validation (
_validate_structure): Ensures required top-level keys (design,sim_options,sim_results,contributor,notes) are present. - Type validation (
_validate_types): Usesvalidate_typesfromsquadds/core/utils.pyto verify Python types match the schema. - Content validation (
_validate_content): Compares the nested structure ofdesign_optionsandsim_options.setupagainst a reference entry in the existing dataset.
data.validate() # Prints: "Structure validated...", "Types validated...", etc.
Synchronize and Submit to Hugging Face
The final phase clones the SQuADDS/SQuADDS_DB repository, creates a new branch, and writes the JSON entry. The contribute() method combines update_repo() and update_db():
repo_path = "/tmp/SQuADDS_DB_clone"
data.contribute(repo_path) # Prints: "Contribution ready for PR"
Under the hood, update_repo(path_to_repo) handles the git clone or git pull, while update_db(path_to_repo) appends the serialized dictionary (produced by to_dict()) to the appropriate <config>.json file with proper indentation. The actual git push is currently disabled in the source (the upload_to_HF call is commented out), so you must manually push the branch using your GITHUB_TOKEN and open a Pull Request on the Hugging Face Hub.
Handling Batch Contributions
For parametric sweeps or multiple experimental runs, use the sweep mode. The from_json() method loads all *.json files from a directory, and validate_sweep() checks each entry individually.
data = ExistingConfigData("qubit-TransmonCross-cap_matrix")
data.from_json("tutorials/examples/sweep_data/", is_sweep=True)
data.validate_sweep()
data.contribute("/tmp/SQuADDS_DB_clone", is_sweep=True)
This workflow is demonstrated comprehensively in Tutorial 3, which walks through every method call and shows the final JSON structure.
Key Source Files and Architecture
Understanding the following files helps debug validation errors or extend the contribution pipeline:
squadds/database/contributor.py: Core API containingExistingConfigData, validation logic, and repository handling.squadds/core/utils.py: Helper functions for schema extraction (get_config_schema), type checking (validate_types), and caching.squadds/core/globals.py: DefinesENV_FILE_PATHused to locate the.envconfiguration..env.example: Template for required contributor metadata and API tokens.
Summary
- Configure environment variables in a root
.envfile to set contributor attribution and API access. - Use
ExistingConfigDatafromsquadds/database/contributor.pyto instantiate a target configuration. - Populate entries using
add_design(),add_sim_result(),add_sim_setup(), andadd_notes(), all of which enforce schema compliance. - Validate entries with
validate()to check structure, types, and content against existing dataset references. - Submit via
contribute(), which clonesSQuADDS/SQuADDS_DB, writes the JSON, and prepares a branch for a Pull Request on Hugging Face. - Batch process multiple entries using
from_json()withis_sweep=Trueandvalidate_sweep().
Frequently Asked Questions
What validation does the SQuADDS contribution API perform before allowing a submission?
The validate() method in squadds/database/contributor.py runs three sequential checks: _validate_structure ensures required keys (design, sim_options, sim_results, contributor, notes) exist; _validate_types uses validate_types from squadds/core/utils.py to confirm Python types match the schema; and _validate_content compares the nested dictionary structure of your design_options and sim_options.setup against a reference entry to ensure field compatibility.
How do I set up my credentials to contribute to the SQuADDS dataset?
Create a .env file in the repository root (as specified in squadds/core/globals.py by ENV_FILE_PATH) containing GROUP_NAME, PI_NAME, INSTITUTION, USER_NAME, HUGGINGFACE_API_KEY, and GITHUB_TOKEN. The ExistingConfigData class automatically loads these variables to attribute the contribution and authenticate repository operations.
Can I contribute multiple simulation results at once instead of one by one?
Yes. For batch contributions, use the from_json() method with is_sweep=True to load an entire directory of JSON files, then call validate_sweep() and contribute(repo_path, is_sweep=True). This processes each entry through the same validation pipeline before updating the local repository clone.
Where is the contributed data physically stored after I run the contribute() method?
The contribute() method serializes your entry to a dictionary via to_dict() and appends it to a JSON file named after the configuration (e.g., qubit-TransmonCross-cap_matrix.json) within your local clone of the SQuADDS/SQuADDS_DB repository. You must then manually push this branch to your Hugging Face fork and create a Pull Request, as the automated git push is currently commented out in the source code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →