How to Execute a Pipeline Defined with Semantica’s DSL: CLI and YAML Guide
To execute a pipeline defined with Semantica’s DSL, run semantica pipeline run --config path/to/pipe.yaml after defining your workflow in a declarative YAML file containing ordered stages with type and config blocks.
Semantica provides a lightweight domain-specific language (DSL) for describing data-processing workflows in declarative YAML files. This guide explains how to execute a pipeline defined with Semantica’s DSL using the built-in CLI, referencing the actual implementation in semantica/pipeline/cli.py and validation logic defined in semantica/pipeline/schema.yaml.
Understanding the Semantica DSL Structure
The Semantica DSL uses a YAML-based schema to chain semantic-aware components without writing Python code. Each pipeline is an ordered list of stages defined under a top-level stages key.
A valid pipeline declaration requires:
version: Schema version identifier (e.g.,1)stages: Ordered list of processing steps, each containing:name: Human-readable identifier for the stagetype: The component to invoke (e.g.,vector_store.loader,embedding_generator)config: Parameter block specific to that component type
# pipe.yaml
version: 1
stages:
- name: load_documents
type: vector_store.loader
config:
source: s3://my-bucket/docs/
format: parquet
- name: embed
type: embedding_generator
config:
model: openai/text-embedding-ada-002
weight: 0.7
- name: enrich_kg
type: kg_enricher
config:
graph_store: neo4j
enable_community_detection: true
When parsed, the CLI constructs a Pipeline object (exposed in semantica/pipeline/__init__.py) that wires components together according to the dependency graph implied by the stage order.
Running Pipelines via the Semantica CLI
The pipeline run command in semantica/pipeline/cli.py provides the primary interface for executing DSL definitions.
Basic Execution Syntax
Run a pipeline with the minimum required argument:
semantica pipeline run --config path/to/pipe.yaml
The CLI performs four internal steps:
- Load and validate the YAML file against
semantica/pipeline/schema.yaml - Instantiate each stage’s implementation found in
semantica/pipeline/stages/*.py - Execute stages sequentially, streaming output from one stage as input to the next
- Persist results according to the final stage’s configuration (e.g., writing embeddings back to a vector store)
Key CLI Flags
--config: Path to the DSL YAML file (required)--dry-run: Preview the execution plan without performing any work--log-level: Adjust verbosity (e.g.,INFO,DEBUG) for troubleshooting
Use --dry-run to validate your pipeline logic before consuming API credits or processing large datasets:
semantica pipeline run --config examples/decision_pipeline.yaml --dry-run
This outputs the execution plan including stage names and component types without invoking the actual implementations.
Complete Pipeline Examples
Simple Embedding Pipeline
This example loads documents from local storage and generates embeddings using OpenAI’s model:
# examples/simple_embedding.yaml
version: 1
stages:
- name: load
type: vector_store.loader
config:
source: ./data/articles.parquet
- name: embed
type: embedding_generator
config:
model: openai/text-embedding-ada-002
Execute with:
semantica pipeline run --config examples/simple_embedding.yaml
Knowledge Graph-Enriched Pipeline
For workflows requiring semantic and structural embeddings, combine vector stores with knowledge graph enrichment:
# examples/decision_pipeline.yaml
version: 1
stages:
- name: load_decisions
type: vector_store.loader
config:
source: ./data/decisions.json
- name: semantic_embed
type: embedding_generator
config:
model: all-MiniLM-L6-v2
weight: 0.7
- name: structural_embed
type: kg_embedder
config:
graph_store: neo4j
weight: 0.3
- name: combine
type: combine_embeddings
config:
method: weighted_average
- name: store
type: vector_store.saver
config:
destination: qdrant
This pipeline is tested in the repository via semantica/tests/verify_rich_cli.py, ensuring the CLI correctly handles multi-stage workflows with KG integration.
Validating Execution Plans
Preview complex workflows before running:
semantica pipeline run --config examples/decision_pipeline.yaml --dry-run
Expected output structure:
Pipeline version: 1
Stages:
1️⃣ load_decisions → vector_store.loader
2️⃣ semantic_embed → embedding_generator
3️⃣ structural_embed → kg_embedder
4️⃣ combine → combine_embeddings
5️⃣ store → vector_store.saver
Summary
- Semantica’s DSL uses declarative YAML to define data-processing pipelines through ordered
stageswithtypeandconfigblocks. - Execute pipelines using
semantica pipeline run --config <file>as implemented insemantica/pipeline/cli.py. - Validate schemas against
semantica/pipeline/schema.yamlbefore execution to catch configuration errors early. - Use
--dry-runto preview execution plans without consuming computational resources. - Stage implementations reside in
semantica/pipeline/stages/*.pyand handle specific component logic likeembedding_generatororkg_enricher.
Frequently Asked Questions
What file format does the Semantica DSL use?
The Semantica DSL uses YAML files with a required version key and a stages list. The CLI validates these files against the JSON Schema defined in semantica/pipeline/schema.yaml before execution begins.
How do I validate a pipeline without running it?
Pass the --dry-run flag to semantica pipeline run. This parses the YAML, validates the schema, and prints the execution plan showing each stage name and component type without instantiating the actual stage implementations or processing data.
Where are pipeline stage implementations defined?
Concrete stage implementations (such as vector_store.loader, embedding_generator, and kg_enricher) are defined in Python files within semantica/pipeline/stages/. The CLI dynamically instantiates these classes based on the type field in your YAML configuration.
Can I adjust logging verbosity during pipeline execution?
Yes. Use the --log-level flag followed by a standard Python logging level (e.g., DEBUG, INFO, WARNING). This is handled by the argument parser in semantica/pipeline/cli.py and controls the detail level of execution logs and error tracebacks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →