Internal Architecture of Semantica's PipelineBuilder DSL: Core Components and Data Flow
Semantica's PipelineBuilder DSL provides a fluent interface for constructing, validating, and executing directed acyclic graph (DAG) workflows through three core data structures—Pipeline, PipelineStep, and PipelineBuilder—enabling declarative pipeline definition with static validation and extensible step handlers.
The semantica-agi/semantica repository implements a powerful Domain-Specific Language (DSL) for orchestrating knowledge-graph and vector-store workflows. This internal architecture analysis examines how the PipelineBuilder DSL transforms declarative pipeline definitions into executable DAGs through its layered validation, serialization, and execution systems.
Core Data Model of the PipelineBuilder DSL
The Pipeline Dataclass
In semantica/pipeline/pipeline_builder.py, the Pipeline dataclass serves as the immutable container for workflow definitions. It stores a list of PipelineStep objects, global configuration parameters, and runtime metadata required for execution tracking and status monitoring across distributed environments.
PipelineStep and Execution Status
Each PipelineStep instance represents a discrete operation within the DAG, encapsulating the step name, type identifier, configuration dictionary, downstream connections, and current execution state via the StepStatus enumeration. This structure enables the ExecutionEngine to traverse dependency graphs and manage step lifecycle transitions from pending through completion or failure states.
PipelineBuilder as DSL Entry Point
The PipelineBuilder class provides the fluent API surface for constructing pipelines programmatically. According to the source code in semantica/pipeline/pipeline_builder.py, it exposes methods including add_step(), connect_steps(), set_parallelism(), and register_step_handler() that ultimately produce a validated Pipeline object through the build() method. This builder pattern abstracts the complexity of DAG construction while maintaining type safety and structural constraints.
Validation and Serialization Infrastructure
Structural Validation Layer
Before a pipeline becomes executable, semantica/pipeline/pipeline_validator.py performs static analysis to enforce structural integrity. The validation logic checks for orphaned steps without connections, validates parallelism constraints against hardware limits, and ensures all referenced step types have registered handlers, catching configuration errors prior to runtime deployment.
Cross-Environment Persistence
The PipelineSerializer class in semantica/pipeline/pipeline_serializer.py handles bidirectional conversion between Pipeline objects and JSON/YAML representations. This enables version-controlled pipeline definitions that can be stored in object storage, transmitted across environments, and reconstructed identically via deserialization methods for reproducible execution.
Execution Engine and Handler Registration
DAG Traversal in ExecutionEngine
Located in semantica/core/orchestrator.py, the ExecutionEngine consumes built Pipeline objects and orchestrates their execution. It walks the DAG of PipelineStep instances in topological order, invoking appropriate handlers while respecting parallelism constraints defined during the building phase through set_parallelism().
Extensible Step Handlers
The DSL supports custom logic injection through the PipelineBuilder.register_step_handler(step_type, handler) method. Developers override default behaviors to implement domain-specific operations such as vector-store ingestion, knowledge-graph analysis, or external API integration without modifying core framework code in the semantica package.
Template System and Common Patterns
Beyond manual construction, semantica/pipeline/pipeline_templates.py provides pre-configured templates for recurring workflow patterns including "construct-template", "vector-store ingestion", and "knowledge-graph enrichment". These templates accelerate development by bundling validated step sequences with optimized default configurations that comply with the DSL's structural requirements.
Practical Implementation Examples
Basic Two-Step Pipeline Construction
from semantica.pipeline import PipelineBuilder
builder = PipelineBuilder()
builder.add_step("load_data", "vector_store_ingest", source="s3://my-bucket/data")
builder.add_step("embed", "embedding", model="openai/text-embedding-ada-002")
builder.connect_steps("load_data", "embed") # load_data → embed
pipeline = builder.build(name="my_simple_pipeline")
Parallel Execution with Custom Handlers
from semantica.pipeline import PipelineBuilder
def my_custom_handler(step_config, context):
# … user‑defined logic …
return {"status": "ok"}
builder = PipelineBuilder()
builder.register_step_handler("my_custom", my_custom_handler)
builder.add_step("step_a", "my_custom", param=1)
builder.add_step("step_b", "my_custom", param=2)
builder.set_parallelism(level=2) # run step_a and step_b concurrently
pipeline = builder.build("parallel_demo")
Serialization and Deserialization Workflows
import json
from semantica.pipeline import PipelineBuilder, PipelineSerializer
builder = PipelineBuilder()
builder.add_step("extract", "semantic_extractor", model="gpt-4")
pipeline = builder.build("exportable")
# Serialize
json_repr = PipelineSerializer.serialize(pipeline, format="json")
print(json_repr)
# Deserialize back to a Pipeline object
restored = PipelineSerializer.deserialize(json_repr, format="json")
Summary
- The PipelineBuilder DSL centers on three core structures:
Pipeline,PipelineStep, andPipelineBuilderdefined insemantica/pipeline/pipeline_builder.py. - Static validation occurs through
pipeline_validator.pybefore execution to prevent runtime structural failures and orphaned steps. - The
ExecutionEngineinsemantica/core/orchestrator.pytraverses the DAG and invokes registered step handlers in topological order. - Custom handlers enable extensibility without core code modification via
register_step_handler(). - JSON/YAML serialization in
pipeline_serializer.pysupports reproducible, version-controlled pipelines across environments.
Frequently Asked Questions
What are the three core data structures in Semantica's PipelineBuilder DSL?
The architecture relies on the Pipeline dataclass (container for workflow definitions), PipelineStep (individual operations with connection metadata and StepStatus), and PipelineBuilder (fluent API entry point), all implemented in semantica/pipeline/pipeline_builder.py. These structures work together to represent the DAG topology and execution state.
How does the ExecutionEngine process a built pipeline?
The ExecutionEngine located in semantica/core/orchestrator.py receives a validated Pipeline object, performs topological sorting of the PipelineStep nodes, and sequentially invokes registered handlers while respecting parallelism settings configured via PipelineBuilder.set_parallelism(). It manages step lifecycle transitions and tracks execution status throughout the workflow run.
Can custom step handlers be registered in the PipelineBuilder DSL?
Yes. Developers call PipelineBuilder.register_step_handler(step_type, handler) to map custom Python functions to specific step type identifiers, enabling integration of vector-store operations, knowledge-graph algorithms, or external API calls without modifying the framework's core execution logic in the semantica package.
Where is pipeline validation logic implemented in the codebase?
Structural validation logic resides in semantica/pipeline/pipeline_validator.py, which checks for orphaned steps, validates that parallelism levels are correctly configured, and verifies that all step types have associated handlers before the PipelineBuilder.build() method returns a concrete Pipeline instance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →