How to Initialize kcmd for Multiple BigQuery Datasets: A Complete Guide
Yes, you can initialize kcmd for multiple BigQuery datasets by passing the --bigquery-dataset flag multiple times in a single kcmd init command.
The kcmd command-line tool in the GoogleCloudPlatform/knowledge-catalog repository supports bulk dataset initialization through its init sub-command. When you initialize kcmd for multiple BigQuery datasets, the tool creates a unified local snapshot containing metadata exports from all specified sources.
How Multiple Dataset Initialization Works
The multi-dataset capability is implemented across two key components in the knowledge-catalog codebase.
Command Line Parsing
In toolbox/mdcode/src/tool/commands.ts, the init sub-command definition handles the --bigquery-dataset flag as a repeatable argument. The CLI collects all occurrences of this flag into an array before forwarding them to the manifest initialization logic.
Manifest Generation
The collected dataset identifiers are passed to CatalogManifest.initWithBigQuery(datasets, ctx) in toolbox/mdcode/src/libts/catalog/manifest.ts. This function accepts an array of dataset strings and builds a manifest that references each BigQuery dataset. According to the source code, the function processes the entire array to create a single coherent catalog structure.
Practical Examples
You can specify multiple datasets by repeating the --bigquery-dataset flag for each target.
Initialize a single dataset:
kcmd init --bigquery-dataset my-project.sales_data
Initialize two datasets in one command:
kcmd init \
--bigquery-dataset my-project.sales_data \
--bigquery-dataset my-project.marketing_insights
For workflows involving many datasets, use line continuation for readability:
kcmd init \
--bigquery-dataset proj1.dataset_a \
--bigquery-dataset proj2.dataset_b \
--bigquery-dataset proj3.dataset_c
Each dataset identifier must follow the <project-id>.<dataset-id> format as documented in toolbox/mdcode/docs/spec.md.
Output Structure
After running kcmd init with multiple datasets, the tool creates two artifacts in your current working directory:
- catalog.json: The manifest file containing references to all specified BigQuery datasets
- snapshot/: A directory containing the exported metadata for every dataset
You can then execute standard kcmd operations (pull, push, status) against this unified snapshot rather than managing separate configurations.
Summary
- Repeatable flag: Pass
--bigquery-datasetmultiple times tokcmd initto process several datasets at once - Array processing: The
CatalogManifest.initWithBigQueryfunction intoolbox/mdcode/src/libts/catalog/manifest.tshandles the dataset array internally - Required argument: You must provide at least one
--bigquery-datasetflag; otherwise, the command exits with a usage error - Unified output: A single
catalog.jsonandsnapshot/directory contain combined metadata from all specified datasets - Format requirement: Use
<project-id>.<dataset-id>syntax for each dataset identifier
Frequently Asked Questions
Can I initialize kcmd for multiple BigQuery datasets in one command?
Yes. The kcmd init command accepts the --bigquery-dataset flag multiple times. Each occurrence specifies a different dataset, and the tool processes all of them in a single execution, creating one unified manifest and snapshot directory.
What is the correct format for BigQuery dataset identifiers?
Dataset identifiers must use the format <project-id>.<dataset-id>. For example, my-project.sales_data or company-analytics.marketing_data. This format is required by the parser in toolbox/mdcode/src/tool/commands.ts and documented in the project specification.
Where does kcmd store metadata when initializing multiple datasets?
The tool creates a catalog.json manifest file and a snapshot/ directory in your current working directory. These contain the exported metadata for all datasets specified in the init command, allowing you to manage multiple BigQuery datasets as a single unit.
What happens if I run kcmd init without specifying any datasets?
The command fails with a usage message. The init sub-command requires at least one --bigquery-dataset argument to function, as confirmed by the implementation in toolbox/mdcode/src/tool/commands.ts and the user-facing specification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →