Available Step Types in Semantica's Pipeline DSL: Complete Reference
Semantica's pipeline DSL supports built-in step types including construct_template, branch, and validation, alongside specialized conflict-detection types (value, entity_wide, type, temporal, logical), while allowing unlimited custom types via the register_step_handler API in semantica/pipeline/pipeline_builder.py.
The semantica-agi/semantica repository implements a declarative pipeline DSL for orchestrating knowledge-graph construction and validation workflows. Every processing node in a pipeline must specify a step_type string that determines which handler function executes the step's logic, making the catalog of available Semantica pipeline DSL step types critical for developers extending the framework.
How the Pipeline DSL Resolves Step Types
The DSL operates on a registration pattern. When you call add_step(step_name, step_type, ...) in semantica/pipeline/pipeline_builder.py, the builder stores the step_type identifier in the step's configuration. At build time, the validator defined in semantica/pipeline/pipeline_validator.py ensures that a callable handler exists for that type—either provided explicitly via the handler= argument or registered ahead of time via register_step_handler(step_type, handler).
This architecture means that any string can serve as a step type provided you register a matching handler. The framework ships with several default handlers that cover the most common knowledge-graph operations.
Core Built-in Step Types
These step types are guaranteed to work out-of-the-box because their handlers are registered by default in the standard PipelineBuilder.
construct_template
The construct_template step type generates template nodes in the triplet store and serves as the primary mechanism for graph construction pipelines.
In semantica/triplet_store/construct_templates.py, the handler for this type creates structured entities based on input data sources. This is the workhorse step for ingesting raw data into the semantic graph.
from semantica.pipeline import PipelineBuilder
builder = PipelineBuilder()
pipeline = (
builder.add_step("ingest_entities", "construct_template", source="s3://data/raw")
.add_step("link_relations", "construct_template", config={"merge_strategy": "upsert"})
.build()
)
branch
The branch step type enables conditional or parallel execution paths within a pipeline. Each branch instance receives its own sub-module, allowing the same step_type to appear multiple times with different configurations.
This pattern is demonstrated in tests/pipeline/test_pipeline_parallel.py, where branching logic coordinates concurrent processing streams.
builder = PipelineBuilder()
pipeline = (
builder.add_step("start", "construct_template")
.add_step("parallel_path_a", "branch", parallel_safe=True)
.add_step("parallel_path_b", "branch", parallel_safe=True)
.build()
)
validation
The validation step type performs integrity checks on the graph or intermediate results before persistence. Typically used as a final gate in processing chains, this step ensures data quality constraints are met.
You can see practical usage in tests/pipeline/test_pipeline_comprehensive.py, where validation steps verify graph consistency after construction phases.
pipeline = (
builder.add_step("load_data", "construct_template")
.add_step("check_integrity", "validation", rules=["unique_ids", "temporal_bounds"])
.build()
)
Specialized Conflict Detection Step Types
In addition to general pipeline steps, the framework defines five specialized step types used exclusively by the conflict detection subsystem in semantica/conflicts/conflict_detector.py. These are not general-purpose pipeline steps but rather identifiers for specific contradiction checks:
value– Detects value-level conflicts between entity attributesentity_wide– Identifies inconsistencies spanning entire entitiestype– Validates type-level contradictions in the graph schematemporal– Detects timeline contradictions and invalid date sequenceslogical– Catches logical impossibilities in semantic relationships
These types are invoked internally by the conflict resolution engine and are not registered in the standard PipelineBuilder unless you explicitly configure a conflict-detection pipeline.
Creating Custom Step Types
To extend beyond the built-in set, use the register_step_handler method. This binds any string identifier to a Python callable, making it available throughout your pipeline definition.
def custom_enrichment_handler(data, context):
"""Custom processing logic for 'enrich' step type."""
return data.enrich(context.config["fields"])
builder = PipelineBuilder()
builder.register_step_handler("enrich", custom_enrichment_handler)
pipeline = (
builder.add_step("raw_input", "construct_template")
.add_step("enrich_data", "enrich", config={"fields": ["timestamp", "source"]})
.build()
)
The validator in semantica/pipeline/pipeline_validator.py will reject any step whose step_type lacks a registered handler or explicit handler argument, preventing runtime errors from undefined step types.
Summary
- Semantica's pipeline DSL uses string-based
step_typeidentifiers to route execution to registered handler functions. - Built-in types include
construct_template(graph construction),branch(parallel/conditional logic), andvalidation(integrity checking). - Conflict detection types (
value,entity_wide,type,temporal,logical) are specialized identifiers used by the internal conflict detector. - Custom step types can be added via
PipelineBuilder.register_step_handler()insemantica/pipeline/pipeline_builder.py. - The pipeline validator ensures every step type has a corresponding handler before the pipeline executes.
Frequently Asked Questions
What is the primary step type for building knowledge graphs in Semantica?
The construct_template step type is the default mechanism for ingesting data and generating graph structures. Defined in semantica/triplet_store/construct_templates.py, this handler converts raw inputs into templated graph nodes and edges, serving as the foundation for most construction pipelines.
Can I define my own step types without modifying the core library?
Yes. The DSL is designed for extensibility. Call builder.register_step_handler("my_type", handler_function) before invoking add_step(). As long as you provide a callable that accepts the step's data and context, the validator in semantica/pipeline/pipeline_validator.py will recognize your custom type as valid.
How does Semantica handle unknown or misspelled step types?
The pipeline validator performs a strict check during build(). If a step specifies a step_type that has no registered handler and no explicit handler= argument, the validator raises a configuration error immediately, preventing deployment of pipelines with undefined step types.
Are conflict detection step types interchangeable with standard pipeline steps?
No. While they share the same step_type metadata structure, types like value, entity_wide, and temporal are consumed by the conflict detection engine in semantica/conflicts/conflict_detector.py rather than the general pipeline executor. Standard workflows should use construct_template, branch, or validation unless specifically implementing custom conflict resolution logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →