How to Use the Pipeline DSL in Semantica to Compose AI Workflows

Semantica provides a fluent Domain-Specific Language (DSL) through the PipelineBuilder class that lets you declare steps, configure parameters, manage dependencies, and control parallelism using a chainable, readable syntax.

The Semantica framework offers a powerful pipeline DSL that abstracts away low-level workflow orchestration, allowing you to focus on AI logic rather than infrastructure plumbing. This declarative interface enables rapid composition of complex processing graphs using method chaining and built-in validation.

Core DSL Components and Workflow

Instantiating the PipelineBuilder

Create a new PipelineBuilder instance to initialize the workflow construction context. According to the source in semantica/pipeline/pipeline_builder.py (lines 95-106), the constructor sets up internal state including a logger, progress tracker, and a PipelineValidator for structural integrity checks.

from semantica.pipeline import PipelineBuilder

builder = PipelineBuilder()

Declaring Steps with add_step

Use the add_step(name, type, **config) method to append processing units to your workflow. As implemented in lines 20-33 of semantica/pipeline/pipeline_builder.py, each invocation creates a PipelineStep dataclass capturing the step name, type identifier, optional configuration dictionary, dependency list, and runtime handler reference.

builder.add_step("ingest", "file_ingest", source="data/")
builder.add_step("parse", "document_parse", formats=["pdf", "txt"])
builder.add_step("embed", "text_embed", model="text-embedding-3-large")

Establishing Dependencies with connect_steps

Define execution order using connect_steps(from_step, to_step), which internally appends the source step name to the target step's dependencies list. The implementation in lines 63-82 of pipeline_builder.py ensures proper graph construction without manual dependency manipulation.

builder.connect_steps("ingest", "parse")
builder.connect_steps("parse", "embed")

Configuring Parallelism

Control concurrent execution with set_parallelism(level). The builder stores this value in its internal pipeline_config dictionary (lines 87-100), and the validator later warns if the level is non-positive.

builder.set_parallelism(2)  # Run up to 2 steps concurrently when possible

Validation and Finalization

The build(name, validate=True) method (lines 108-126) produces a final Pipeline dataclass containing the ordered step list, configuration, and metadata. By default, this invokes PipelineValidator to check for duplicate step names, missing dependencies, and circular references (as defined in semantica/pipeline/pipeline_validator.py, lines 90-114).

pipeline = builder.build(name="document_processing_pipeline")

Practical Implementation Patterns

Building a Linear Document Processing Pipeline

Chain methods to create a complete workflow:

from semantica.pipeline import PipelineBuilder

builder = PipelineBuilder()
pipeline = (
    builder
    .add_step("ingest", "file_ingest", source="data/")
    .add_step("parse", "document_parse", formats=["pdf", "txt"])
    .add_step("embed", "text_embed", model="text-embedding-3-large")
    .connect_steps("ingest", "parse")
    .connect_steps("parse", "embed")
    .set_parallelism(2)
    .build(name="simple_doc_pipeline")
)

Leveraging Pipeline Templates

Use PipelineTemplateManager to instantiate pre-configured workflows and override specific parameters:

from semantica.pipeline import PipelineTemplateManager

mgr = PipelineTemplateManager()
builder = mgr.create_pipeline_from_template(
    "document_processing",
    ingest={"source": "my_docs/"},
    embed={"model": "text-embedding-3-small"},
)
pipeline = builder.build(name="custom_doc_pipeline")

Registering Custom Step Handlers

Extend the DSL with domain-specific logic by registering custom handlers before building:

def my_custom_handler(step_input):
    # Custom AI processing logic here

    return step_input

builder = PipelineBuilder()
builder.register_step_handler("my_step", my_custom_handler)
pipeline = (
    builder
    .add_step("start", "my_step", param=42)
    .build()
)

Key Source Files

Summary

  • PipelineBuilder provides the core DSL interface for fluent workflow construction in Semantica.
  • Use add_step to declare processing units with type identifiers and configuration parameters.
  • connect_steps establishes execution dependencies by modifying step dependency lists internally.
  • set_parallelism controls concurrent execution levels stored in the pipeline configuration.
  • build finalizes the workflow into a Pipeline dataclass and triggers validation unless explicitly disabled.
  • The PipelineTemplateManager enables rapid instantiation of common workflow patterns with selective overrides.
  • Custom step handlers can be registered via register_step_handler to extend built-in functionality.

Frequently Asked Questions

How do I validate a Semantica pipeline before execution?

The DSL automatically validates pipelines during the build() method call via the PipelineValidator class (lines 90-114 in semantica/pipeline/pipeline_validator.py). This checks for duplicate step names, missing dependencies, and circular references. You can disable validation by passing validate=False to build(), though this is not recommended for production workflows.

Can I modify step dependencies after adding them to the builder?

While the connect_steps(from_step, to_step) method is the primary interface for establishing dependencies, the underlying PipelineStep dataclass stores dependencies as a list that is appended to during connection (lines 63-82 in pipeline_builder.py). However, direct mutation of internal step structures is discouraged; use the builder's methods to maintain structural integrity.

What parallelism levels does Semantica support?

The set_parallelism(level) method accepts any positive integer, storing it in the internal pipeline_config dictionary (lines 87-100). The validator warns if you specify a non-positive value. The actual concurrency implementation respects this configuration during pipeline execution, running up to the specified number of independent steps simultaneously.

How do I reuse common pipeline patterns across projects?

Use the PipelineTemplateManager from semantica/pipeline/pipeline_templates.py to load pre-defined workflow templates. Call create_pipeline_from_template(template_name, **overrides) to instantiate a builder with default configurations, then selectively override specific step parameters before calling build().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →