How to Use the Pipeline DSL in Semantica to Compose AI Workflows
Semantica provides a fluent Domain-Specific Language (DSL) through the PipelineBuilder class that lets you declare steps, configure parameters, manage dependencies, and control parallelism using a chainable, readable syntax.
The Semantica framework offers a powerful pipeline DSL that abstracts away low-level workflow orchestration, allowing you to focus on AI logic rather than infrastructure plumbing. This declarative interface enables rapid composition of complex processing graphs using method chaining and built-in validation.
Core DSL Components and Workflow
Instantiating the PipelineBuilder
Create a new PipelineBuilder instance to initialize the workflow construction context. According to the source in semantica/pipeline/pipeline_builder.py (lines 95-106), the constructor sets up internal state including a logger, progress tracker, and a PipelineValidator for structural integrity checks.
from semantica.pipeline import PipelineBuilder
builder = PipelineBuilder()
Declaring Steps with add_step
Use the add_step(name, type, **config) method to append processing units to your workflow. As implemented in lines 20-33 of semantica/pipeline/pipeline_builder.py, each invocation creates a PipelineStep dataclass capturing the step name, type identifier, optional configuration dictionary, dependency list, and runtime handler reference.
builder.add_step("ingest", "file_ingest", source="data/")
builder.add_step("parse", "document_parse", formats=["pdf", "txt"])
builder.add_step("embed", "text_embed", model="text-embedding-3-large")
Establishing Dependencies with connect_steps
Define execution order using connect_steps(from_step, to_step), which internally appends the source step name to the target step's dependencies list. The implementation in lines 63-82 of pipeline_builder.py ensures proper graph construction without manual dependency manipulation.
builder.connect_steps("ingest", "parse")
builder.connect_steps("parse", "embed")
Configuring Parallelism
Control concurrent execution with set_parallelism(level). The builder stores this value in its internal pipeline_config dictionary (lines 87-100), and the validator later warns if the level is non-positive.
builder.set_parallelism(2) # Run up to 2 steps concurrently when possible
Validation and Finalization
The build(name, validate=True) method (lines 108-126) produces a final Pipeline dataclass containing the ordered step list, configuration, and metadata. By default, this invokes PipelineValidator to check for duplicate step names, missing dependencies, and circular references (as defined in semantica/pipeline/pipeline_validator.py, lines 90-114).
pipeline = builder.build(name="document_processing_pipeline")
Practical Implementation Patterns
Building a Linear Document Processing Pipeline
Chain methods to create a complete workflow:
from semantica.pipeline import PipelineBuilder
builder = PipelineBuilder()
pipeline = (
builder
.add_step("ingest", "file_ingest", source="data/")
.add_step("parse", "document_parse", formats=["pdf", "txt"])
.add_step("embed", "text_embed", model="text-embedding-3-large")
.connect_steps("ingest", "parse")
.connect_steps("parse", "embed")
.set_parallelism(2)
.build(name="simple_doc_pipeline")
)
Leveraging Pipeline Templates
Use PipelineTemplateManager to instantiate pre-configured workflows and override specific parameters:
from semantica.pipeline import PipelineTemplateManager
mgr = PipelineTemplateManager()
builder = mgr.create_pipeline_from_template(
"document_processing",
ingest={"source": "my_docs/"},
embed={"model": "text-embedding-3-small"},
)
pipeline = builder.build(name="custom_doc_pipeline")
Registering Custom Step Handlers
Extend the DSL with domain-specific logic by registering custom handlers before building:
def my_custom_handler(step_input):
# Custom AI processing logic here
return step_input
builder = PipelineBuilder()
builder.register_step_handler("my_step", my_custom_handler)
pipeline = (
builder
.add_step("start", "my_step", param=42)
.build()
)
Key Source Files
semantica/pipeline/pipeline_builder.py: Implements thePipelineBuilderclass,PipelineStepdataclass, and DSL methods includingadd_step,connect_steps, andbuild.semantica/pipeline/pipeline_validator.py: ContainsPipelineValidatorwhich performs structural validation checks for duplicates, missing dependencies, and circular references.semantica/pipeline/pipeline_templates.py: ProvidesPipelineTemplateManagerand ready-made pipeline configurations.
Summary
PipelineBuilderprovides the core DSL interface for fluent workflow construction in Semantica.- Use
add_stepto declare processing units with type identifiers and configuration parameters. connect_stepsestablishes execution dependencies by modifying step dependency lists internally.set_parallelismcontrols concurrent execution levels stored in the pipeline configuration.buildfinalizes the workflow into aPipelinedataclass and triggers validation unless explicitly disabled.- The
PipelineTemplateManagerenables rapid instantiation of common workflow patterns with selective overrides. - Custom step handlers can be registered via
register_step_handlerto extend built-in functionality.
Frequently Asked Questions
How do I validate a Semantica pipeline before execution?
The DSL automatically validates pipelines during the build() method call via the PipelineValidator class (lines 90-114 in semantica/pipeline/pipeline_validator.py). This checks for duplicate step names, missing dependencies, and circular references. You can disable validation by passing validate=False to build(), though this is not recommended for production workflows.
Can I modify step dependencies after adding them to the builder?
While the connect_steps(from_step, to_step) method is the primary interface for establishing dependencies, the underlying PipelineStep dataclass stores dependencies as a list that is appended to during connection (lines 63-82 in pipeline_builder.py). However, direct mutation of internal step structures is discouraged; use the builder's methods to maintain structural integrity.
What parallelism levels does Semantica support?
The set_parallelism(level) method accepts any positive integer, storing it in the internal pipeline_config dictionary (lines 87-100). The validator warns if you specify a non-positive value. The actual concurrency implementation respects this configuration during pipeline execution, running up to the specified number of independent steps simultaneously.
How do I reuse common pipeline patterns across projects?
Use the PipelineTemplateManager from semantica/pipeline/pipeline_templates.py to load pre-defined workflow templates. Call create_pipeline_from_template(template_name, **overrides) to instantiate a builder with default configurations, then selectively override specific step parameters before calling build().
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →