Is Semantica Deterministic by Default? LLM Integration and Architecture Explained
Semantica is deterministic by default for all core operations, using hash-based algorithms for entity IDs, embeddings, and exports, while LLM integration remains optional and isolated behind an abstraction layer that supports deterministic outputs via temperature controls.
The semantica-agi/semantica repository provides a knowledge graph construction platform engineered for reproducibility. While the framework guarantees deterministic behavior for all algorithmic operations—from node identification to export generation—it integrates large language models through an opt-in abstraction layer that preserves full auditability. Understanding how Semantica maintains determinism by default while accommodating non-deterministic LLM calls is essential for building reliable, auditable AI systems.
How Semantica Achieves Determinism by Default
Every core operation that does not involve a language model is generated from stable, hash-based algorithms. Identifiers, normalizations, splits, exports, and vector-store look-ups all produce identical output artifacts when given identical input data.
Hash-Based Entity Identification
In semantica/graph_store/graph_store.py, entity and graph identifiers are created using MD5 hashes of the entity's canonical contents. The implementation generates repeatable IRIs through the pattern hash_id = hashlib.md5(...).hexdigest()[:8]. This approach guarantees that the same node receives the same identifier across different execution runs, ensuring graph merging and updates remain idempotent.
Deterministic Export and Embedding Generation
Export modules located in paths such as semantica/export/json_exporter.py reuse existing IDs or derive deterministic ones directly from node contents, producing repeatable RDF, CSV, and JSON-LD files. For vector operations, semantica/embeddings/text_embedder.py implements deterministic embeddings based on SHA-256 hashing of input text. This enables reproducible similarity calculations without invoking a neural embedding model.
Algorithmic Processing Without Randomness
Chunking algorithms in semantica/split/kg_chunkers.py rely on pure-Python graph analysis—for example, graph-centrality calculations—without initializing random seeds. Similarly, semantica/kg/temporal_query_rewriter.py performs deterministic temporal filtering on text before any potential LLM invocation, as explicitly documented by the inline comment "deterministic, zero LLM calls". The provenance system in semantica/provenance/integrity.py constructs checksums from lexicographically sorted keys, ensuring that identical provenance records always yield identical hash values.
LLM Integration Architecture
When language model capabilities are required, Semantica delegates calls to the abstraction layer under semantica.llms. Each provider implements a thin wrapper around vendor SDKs, exposing a unified interface while maintaining separation from deterministic core operations.
The Abstraction Layer
Provider implementations—including OpenAI, Anthropic, Groq, HuggingFace, and LiteLLM—reside in semantica/semantic_extract/providers.py and provider-specific files like semantica/llms/openai.py. These modules expose standardized generate and generate_structured methods that normalize interaction patterns across different model vendors. The LLM layer does not alter deterministic semantics; it serves as an optional component invoked only for tasks that cannot be solved purely algorithmically, such as free-form summarization, ontology generation, or complex reasoning.
Optional Non-Deterministic Steps
Determinism of LLM outputs is controlled explicitly via generation parameters. If the user sets temperature=0—or utilizes a provider-specific deterministic mode—the model returns identical text for identical prompts. This parameter is documented in every provider's generate method signature, allowing users to enforce reproducibility even when using generative models.
Provenance and Auditability
All LLM calls are wrapped by provenance mixins defined in semantica/llms/llms_provenance. These mixins record the prompt text, model identifier, temperature setting, and raw response payload. This architecture ensures that any non-deterministic step remains fully traceable and auditable, creating an immutable record of generative AI involvement in the pipeline.
Practical Implementation Examples
Building a Deterministic Knowledge Graph
The following pipeline ingests a document and constructs a graph using only deterministic operations:
from semantica.ingest.file_ingestor import FileIngestor
from semantica.graph_store.graph_store import GraphStore
# Ingest a PDF → raw documents (deterministic)
ingestor = FileIngestor()
raw_docs = ingestor.ingest("example.pdf")
# Parse → normalize → split → extract (all deterministic)
parser = ... # use semantica.parse.DocumentParser
graph = parser.build_graph(raw_docs) # node IDs are MD5 hashes
# Store the graph – IDs are reproducible across runs
store = GraphStore(backend="neo4j")
store.save(graph) # same graph → same IDs
Re-running this script on the same PDF yields an identical graph structure with identical node IRIs.
Configuring Deterministic LLM Outputs
To obtain deterministic responses from OpenAI models, explicitly set the temperature parameter to zero:
from semantica.llms.openai import OpenAI
# Initialise the provider (API key taken from env)
llm = OpenAI(model="gpt-4")
# Request a deterministic response (temperature=0)
prompt = "Summarize the following paragraph in one sentence."
text = "Semantica is a platform for building deterministic knowledge graphs..."
summary = llm.generate(prompt + "\n\n" + text, temperature=0)
print(summary)
Because temperature=0, the model returns the same summary every time it receives the identical prompt.
Auditing LLM Calls with Provenance
Wrap providers with provenance mixins to automatically capture audit trails:
from semantica.llms.llms_provenance import OpenAILLMWithProvenance
# Wrap the provider to automatically log provenance
llm = OpenAILLMWithProvenance(model="gpt-4")
response = llm.generate("Explain RDF in two lines.", temperature=0)
# Access the provenance record
print(llm.last_provenance) # shows prompt, model, temperature, raw output
The provenance mixin captures the non-deterministic step's parameters and output, ensuring compliance and reproducibility tracking.
Summary
- Hash-based identity:
semantica/graph_store/graph_store.pyuses MD5 hashes to generate reproducible entity IRIs. - Deterministic embeddings:
semantica/embeddings/text_embedder.pyimplements SHA-256-based vectors for consistent similarity calculations. - Pure algorithmic processing: Splitting and temporal reasoning operate without random seeds or LLM calls.
- Optional LLM layer: The
semantica.llmsabstraction treats generative AI as an opt-in enrichment step. - Deterministic LLM configuration: Setting
temperature=0in providergeneratemethods ensures repeatable model outputs. - Full auditability:
semantica/llms/llms_provenance.pyrecords every LLM invocation for compliance and debugging.
Frequently Asked Questions
Is Semantica deterministic without any LLM configuration?
Yes, Semantica guarantees determinism by default for all core knowledge graph operations. The system utilizes hash-based algorithms in graph_store.py for entity identification and text_embedder.py for vector generation, ensuring that identical inputs always produce identical outputs without requiring any LLM parameters or API keys.
How does Semantica handle temperature settings for OpenAI models?
The OpenAI wrapper in semantica/llms/openai.py accepts a temperature parameter in its generate method signature. Setting temperature=0 configures the model to return the same text for identical prompts, effectively neutralizing the randomness inherent in neural sampling and making the LLM step deterministic.
Can I use Semantica without any LLM integration?
Yes, all core pipeline functions—including document ingestion, entity extraction, graph construction, and export—operate without invoking language models. The framework explicitly reserves LLM calls for optional tasks such as summarization or ontology generation, while components like temporal_query_rewriter.py perform complex operations with "zero LLM calls" as documented in the source.
How does Semantica ensure auditability of LLM calls?
The semantica.llms.llms_provenance module provides mixins that automatically capture the prompt text, model identifier, temperature setting, and raw response for every LLM invocation. This creates an immutable audit trail that tracks exactly when and how non-deterministic generative steps influenced the knowledge graph construction process.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →