# Is Semantica Deterministic by Default? LLM Integration and Architecture Explained

> Discover if Semantica is deterministic by default. Learn how its LLM integration and architecture ensure controlled, predictable AI operations for your projects.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-13

---

**Semantica is deterministic by default for all core operations, using hash-based algorithms for entity IDs, embeddings, and exports, while LLM integration remains optional and isolated behind an abstraction layer that supports deterministic outputs via temperature controls.**

The `semantica-agi/semantica` repository provides a knowledge graph construction platform engineered for reproducibility. While the framework guarantees deterministic behavior for all algorithmic operations—from node identification to export generation—it integrates large language models through an opt-in abstraction layer that preserves full auditability. Understanding how Semantica maintains determinism by default while accommodating non-deterministic LLM calls is essential for building reliable, auditable AI systems.

## How Semantica Achieves Determinism by Default

Every core operation that does not involve a language model is generated from stable, hash-based algorithms. Identifiers, normalizations, splits, exports, and vector-store look-ups all produce identical output artifacts when given identical input data.

### Hash-Based Entity Identification

In [`semantica/graph_store/graph_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/graph_store.py), entity and graph identifiers are created using MD5 hashes of the entity's canonical contents. The implementation generates repeatable IRIs through the pattern `hash_id = hashlib.md5(...).hexdigest()[:8]`. This approach guarantees that the same node receives the same identifier across different execution runs, ensuring graph merging and updates remain idempotent.

### Deterministic Export and Embedding Generation

Export modules located in paths such as [`semantica/export/json_exporter.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/export/json_exporter.py) reuse existing IDs or derive deterministic ones directly from node contents, producing repeatable RDF, CSV, and JSON-LD files. For vector operations, [`semantica/embeddings/text_embedder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/embeddings/text_embedder.py) implements deterministic embeddings based on SHA-256 hashing of input text. This enables reproducible similarity calculations without invoking a neural embedding model.

### Algorithmic Processing Without Randomness

Chunking algorithms in [`semantica/split/kg_chunkers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/kg_chunkers.py) rely on pure-Python graph analysis—for example, graph-centrality calculations—without initializing random seeds. Similarly, [`semantica/kg/temporal_query_rewriter.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/temporal_query_rewriter.py) performs deterministic temporal filtering on text before any potential LLM invocation, as explicitly documented by the inline comment "deterministic, zero LLM calls". The provenance system in [`semantica/provenance/integrity.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/integrity.py) constructs checksums from lexicographically sorted keys, ensuring that identical provenance records always yield identical hash values.

## LLM Integration Architecture

When language model capabilities are required, Semantica delegates calls to the abstraction layer under `semantica.llms`. Each provider implements a thin wrapper around vendor SDKs, exposing a unified interface while maintaining separation from deterministic core operations.

### The Abstraction Layer

Provider implementations—including OpenAI, Anthropic, Groq, HuggingFace, and LiteLLM—reside in [`semantica/semantic_extract/providers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/providers.py) and provider-specific files like [`semantica/llms/openai.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/llms/openai.py). These modules expose standardized `generate` and `generate_structured` methods that normalize interaction patterns across different model vendors. The LLM layer does not alter deterministic semantics; it serves as an optional component invoked only for tasks that cannot be solved purely algorithmically, such as free-form summarization, ontology generation, or complex reasoning.

### Optional Non-Deterministic Steps

Determinism of LLM outputs is controlled explicitly via generation parameters. If the user sets `temperature=0`—or utilizes a provider-specific deterministic mode—the model returns identical text for identical prompts. This parameter is documented in every provider's `generate` method signature, allowing users to enforce reproducibility even when using generative models.

### Provenance and Auditability

All LLM calls are wrapped by provenance mixins defined in `semantica/llms/llms_provenance`. These mixins record the prompt text, model identifier, temperature setting, and raw response payload. This architecture ensures that any non-deterministic step remains fully traceable and auditable, creating an immutable record of generative AI involvement in the pipeline.

## Practical Implementation Examples

### Building a Deterministic Knowledge Graph

The following pipeline ingests a document and constructs a graph using only deterministic operations:

```python
from semantica.ingest.file_ingestor import FileIngestor
from semantica.graph_store.graph_store import GraphStore

# Ingest a PDF → raw documents (deterministic)

ingestor = FileIngestor()
raw_docs = ingestor.ingest("example.pdf")

# Parse → normalize → split → extract (all deterministic)

parser = ...  # use semantica.parse.DocumentParser

graph = parser.build_graph(raw_docs)   # node IDs are MD5 hashes

# Store the graph – IDs are reproducible across runs

store = GraphStore(backend="neo4j")
store.save(graph)                      # same graph → same IDs

```

Re-running this script on the same PDF yields an identical graph structure with identical node IRIs.

### Configuring Deterministic LLM Outputs

To obtain deterministic responses from OpenAI models, explicitly set the temperature parameter to zero:

```python
from semantica.llms.openai import OpenAI

# Initialise the provider (API key taken from env)

llm = OpenAI(model="gpt-4")

# Request a deterministic response (temperature=0)

prompt = "Summarize the following paragraph in one sentence."
text   = "Semantica is a platform for building deterministic knowledge graphs..."
summary = llm.generate(prompt + "\n\n" + text, temperature=0)

print(summary)

```

Because `temperature=0`, the model returns the same summary every time it receives the identical prompt.

### Auditing LLM Calls with Provenance

Wrap providers with provenance mixins to automatically capture audit trails:

```python
from semantica.llms.llms_provenance import OpenAILLMWithProvenance

# Wrap the provider to automatically log provenance

llm = OpenAILLMWithProvenance(model="gpt-4")
response = llm.generate("Explain RDF in two lines.", temperature=0)

# Access the provenance record

print(llm.last_provenance)   # shows prompt, model, temperature, raw output

```

The provenance mixin captures the non-deterministic step's parameters and output, ensuring compliance and reproducibility tracking.

## Summary

- **Hash-based identity**: [`semantica/graph_store/graph_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/graph_store.py) uses MD5 hashes to generate reproducible entity IRIs.
- **Deterministic embeddings**: [`semantica/embeddings/text_embedder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/embeddings/text_embedder.py) implements SHA-256-based vectors for consistent similarity calculations.
- **Pure algorithmic processing**: Splitting and temporal reasoning operate without random seeds or LLM calls.
- **Optional LLM layer**: The `semantica.llms` abstraction treats generative AI as an opt-in enrichment step.
- **Deterministic LLM configuration**: Setting `temperature=0` in provider `generate` methods ensures repeatable model outputs.
- **Full auditability**: [`semantica/llms/llms_provenance.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/llms/llms_provenance.py) records every LLM invocation for compliance and debugging.

## Frequently Asked Questions

### Is Semantica deterministic without any LLM configuration?

Yes, Semantica guarantees determinism by default for all core knowledge graph operations. The system utilizes hash-based algorithms in [`graph_store.py`](https://github.com/semantica-agi/semantica/blob/main/graph_store.py) for entity identification and [`text_embedder.py`](https://github.com/semantica-agi/semantica/blob/main/text_embedder.py) for vector generation, ensuring that identical inputs always produce identical outputs without requiring any LLM parameters or API keys.

### How does Semantica handle temperature settings for OpenAI models?

The OpenAI wrapper in [`semantica/llms/openai.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/llms/openai.py) accepts a `temperature` parameter in its `generate` method signature. Setting `temperature=0` configures the model to return the same text for identical prompts, effectively neutralizing the randomness inherent in neural sampling and making the LLM step deterministic.

### Can I use Semantica without any LLM integration?

Yes, all core pipeline functions—including document ingestion, entity extraction, graph construction, and export—operate without invoking language models. The framework explicitly reserves LLM calls for optional tasks such as summarization or ontology generation, while components like [`temporal_query_rewriter.py`](https://github.com/semantica-agi/semantica/blob/main/temporal_query_rewriter.py) perform complex operations with "zero LLM calls" as documented in the source.

### How does Semantica ensure auditability of LLM calls?

The `semantica.llms.llms_provenance` module provides mixins that automatically capture the prompt text, model identifier, temperature setting, and raw response for every LLM invocation. This creates an immutable audit trail that tracks exactly when and how non-deterministic generative steps influenced the knowledge graph construction process.