How to Execute a Pipeline Defined with Semantica’s DSL: CLI and YAML Guide

To execute a pipeline defined with Semantica’s DSL, run semantica pipeline run --config path/to/pipe.yaml after defining your workflow in a declarative YAML file containing ordered stages with type and config blocks.

Semantica provides a lightweight domain-specific language (DSL) for describing data-processing workflows in declarative YAML files. This guide explains how to execute a pipeline defined with Semantica’s DSL using the built-in CLI, referencing the actual implementation in semantica/pipeline/cli.py and validation logic defined in semantica/pipeline/schema.yaml.

Understanding the Semantica DSL Structure

The Semantica DSL uses a YAML-based schema to chain semantic-aware components without writing Python code. Each pipeline is an ordered list of stages defined under a top-level stages key.

A valid pipeline declaration requires:

  • version: Schema version identifier (e.g., 1)
  • stages: Ordered list of processing steps, each containing:
    • name: Human-readable identifier for the stage
    • type: The component to invoke (e.g., vector_store.loader, embedding_generator)
    • config: Parameter block specific to that component type

# pipe.yaml

version: 1
stages:
  - name: load_documents
    type: vector_store.loader
    config:
      source: s3://my-bucket/docs/
      format: parquet
  - name: embed
    type: embedding_generator
    config:
      model: openai/text-embedding-ada-002
      weight: 0.7
  - name: enrich_kg
    type: kg_enricher
    config:
      graph_store: neo4j
      enable_community_detection: true

When parsed, the CLI constructs a Pipeline object (exposed in semantica/pipeline/__init__.py) that wires components together according to the dependency graph implied by the stage order.

Running Pipelines via the Semantica CLI

The pipeline run command in semantica/pipeline/cli.py provides the primary interface for executing DSL definitions.

Basic Execution Syntax

Run a pipeline with the minimum required argument:

semantica pipeline run --config path/to/pipe.yaml

The CLI performs four internal steps:

  1. Load and validate the YAML file against semantica/pipeline/schema.yaml
  2. Instantiate each stage’s implementation found in semantica/pipeline/stages/*.py
  3. Execute stages sequentially, streaming output from one stage as input to the next
  4. Persist results according to the final stage’s configuration (e.g., writing embeddings back to a vector store)

Key CLI Flags

  • --config: Path to the DSL YAML file (required)
  • --dry-run: Preview the execution plan without performing any work
  • --log-level: Adjust verbosity (e.g., INFO, DEBUG) for troubleshooting

Use --dry-run to validate your pipeline logic before consuming API credits or processing large datasets:

semantica pipeline run --config examples/decision_pipeline.yaml --dry-run

This outputs the execution plan including stage names and component types without invoking the actual implementations.

Complete Pipeline Examples

Simple Embedding Pipeline

This example loads documents from local storage and generates embeddings using OpenAI’s model:


# examples/simple_embedding.yaml

version: 1
stages:
  - name: load
    type: vector_store.loader
    config:
      source: ./data/articles.parquet
  - name: embed
    type: embedding_generator
    config:
      model: openai/text-embedding-ada-002

Execute with:

semantica pipeline run --config examples/simple_embedding.yaml

Knowledge Graph-Enriched Pipeline

For workflows requiring semantic and structural embeddings, combine vector stores with knowledge graph enrichment:


# examples/decision_pipeline.yaml

version: 1
stages:
  - name: load_decisions
    type: vector_store.loader
    config:
      source: ./data/decisions.json
  - name: semantic_embed
    type: embedding_generator
    config:
      model: all-MiniLM-L6-v2
      weight: 0.7
  - name: structural_embed
    type: kg_embedder
    config:
      graph_store: neo4j
      weight: 0.3
  - name: combine
    type: combine_embeddings
    config:
      method: weighted_average
  - name: store
    type: vector_store.saver
    config:
      destination: qdrant

This pipeline is tested in the repository via semantica/tests/verify_rich_cli.py, ensuring the CLI correctly handles multi-stage workflows with KG integration.

Validating Execution Plans

Preview complex workflows before running:

semantica pipeline run --config examples/decision_pipeline.yaml --dry-run

Expected output structure:


Pipeline version: 1
Stages:
  1️⃣ load_decisions → vector_store.loader
  2️⃣ semantic_embed → embedding_generator
  3️⃣ structural_embed → kg_embedder
  4️⃣ combine → combine_embeddings
  5️⃣ store → vector_store.saver

Summary

  • Semantica’s DSL uses declarative YAML to define data-processing pipelines through ordered stages with type and config blocks.
  • Execute pipelines using semantica pipeline run --config <file> as implemented in semantica/pipeline/cli.py.
  • Validate schemas against semantica/pipeline/schema.yaml before execution to catch configuration errors early.
  • Use --dry-run to preview execution plans without consuming computational resources.
  • Stage implementations reside in semantica/pipeline/stages/*.py and handle specific component logic like embedding_generator or kg_enricher.

Frequently Asked Questions

What file format does the Semantica DSL use?

The Semantica DSL uses YAML files with a required version key and a stages list. The CLI validates these files against the JSON Schema defined in semantica/pipeline/schema.yaml before execution begins.

How do I validate a pipeline without running it?

Pass the --dry-run flag to semantica pipeline run. This parses the YAML, validates the schema, and prints the execution plan showing each stage name and component type without instantiating the actual stage implementations or processing data.

Where are pipeline stage implementations defined?

Concrete stage implementations (such as vector_store.loader, embedding_generator, and kg_enricher) are defined in Python files within semantica/pipeline/stages/. The CLI dynamically instantiates these classes based on the type field in your YAML configuration.

Can I adjust logging verbosity during pipeline execution?

Yes. Use the --log-level flag followed by a standard Python logging level (e.g., DEBUG, INFO, WARNING). This is handled by the argument parser in semantica/pipeline/cli.py and controls the detail level of execution logs and error tracebacks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →