# How to Implement an End-to-End GraphRAG Pipeline with Semantica: A Complete Guide

> Learn how to implement an end-to-end GraphRAG pipeline with Semantica. Discover a modular architecture combining vector stores, knowledge graphs, and memory for advanced retrieval.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Semantica provides a modular architecture that combines vector-store retrieval, knowledge-graph traversal, and memory stores into a single unified GraphRAG pipeline through its `ContextRetriever` class and pipeline templates.**

The `semantica-agi/semantica` repository offers a production-ready framework for building Retrieval-Augmented Generation systems that leverage both dense vector similarity and structured graph relationships. By combining the `rag_pipeline` template with the hybrid `ContextRetriever`, you can implement sophisticated GraphRAG workflows that retrieve context from unstructured documents and structured knowledge graphs simultaneously.

## Core Components of the Semantica GraphRAG Architecture

Semantica's GraphRAG implementation relies on five key components that work together to provide hybrid retrieval capabilities:

- **`PipelineTemplate`** ([`semantica/pipeline/pipeline_templates.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_templates.py)) – Defines reusable workflows for ingestion, chunking, embedding, and storage
- **`ContextRetriever`** ([`semantica/context/context_retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context/context_retriever.py)) – The core hybrid retrieval engine that merges vector search results with graph traversals
- **`SemanticaRetriever`** ([`semantica/integrations/langchain/retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/integrations/langchain/retriever.py)) – A LangChain-compatible wrapper that exposes the hybrid retriever to existing LLM chains
- **Knowledge Graph Adapter** ([`semantica/integrations/agno/knowledge_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/integrations/agno/knowledge_graph.py)) – Implements the `GraphStore` interface with methods like `query` and `get_neighbors`
- **Vector Store** – Any compatible store (Weaviate, pgvector) providing dense embedding search via a standard `search`/`embed` API

## Step 1: Initialize the RAG Pipeline Template

Start by loading the pre-defined RAG template from [`pipeline_templates.py`](https://github.com/semantica-agi/semantica/blob/main/pipeline_templates.py), which handles document ingestion through vector storage automatically.

```python
from semantica.pipeline import PipelineTemplateManager

# Initialize the template manager

manager = PipelineTemplateManager()

# Create a pipeline builder from the RAG template

builder = manager.create_pipeline_from_template(
    "rag_pipeline",
    chunk={"chunk_size": 512},
    embed={"model": "text-embedding-3-large"},
    store_vectors={"store": "weaviate"},
)

# Extract configured components

vector_store = builder.get_component("store_vectors")
knowledge_graph = builder.get_component("knowledge_graph")

```

The `rag_pipeline` template encapsulates the entire ingestion workflow: document loading, text chunking, embedding generation, and vector storage. You can override defaults like `chunk_size` and embedding models through the configuration dictionary.

## Step 2: Configure the Hybrid Context Retriever

The `ContextRetriever` class in [`semantica/context/context_retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context/context_retriever.py) implements the GraphRAG logic, performing vector search, graph expansion, and result fusion in a single retrieval operation.

```python
from semantica.context import ContextRetriever

# Initialize the hybrid retriever

retriever = ContextRetriever(
    vector_store=vector_store,
    knowledge_graph=knowledge_graph,
    hybrid_alpha=0.6,          # 60% graph weight, 40% vector

    max_expansion_hops=2,      # Multi-hop neighborhood expansion

)

```

Key parameters control the retrieval behavior:

- **`hybrid_alpha`** (float: 0.0 to 1.0) – Blending weight between vector and graph scores (0 = vector only, 1 = graph only)
- **`max_expansion_hops`** (int) – Number of hops for graph neighborhood expansion beyond initial matches
- **`min_relevance_score`** (float) – Minimum threshold for result inclusion

The retriever executes `_retrieve_from_vector()` for dense similarity search and `_retrieve_from_graph()` for structured queries, then normalizes scores per source and applies semantic re-ranking against the query embedding.

## Step 3: Integrate Knowledge Graph Construction

For graph-aware retrieval, instantiate a knowledge graph using the `kg_construction` template or load an existing graph via the Agno adapter.

```python

# Option A: Build a new knowledge graph from source documents

kg_builder = manager.create_pipeline_from_template(
    "kg_construction",
    ingest_sources={"sources": ["data/clinical_trials/"]},
    extract_entities={"entities": True},
    extract_relations={"relationships": True},
)
knowledge_graph = kg_builder.get_component("build_graph")

# Option B: Use existing graph with Agno adapter

from semantica.integrations.agno import AgnoKnowledgeGraph
knowledge_graph = AgnoKnowledgeGraph(connection_params={"uri": "neo4j://localhost"})

```

The `AgnoKnowledgeGraph` class implements the required `GraphStore` interface with `query()` and `get_neighbors()` methods, enabling the retriever to traverse entity relationships during context gathering.

## Step 4: Connect to LangChain Pipelines

For integration with existing LangChain applications, wrap the Semantica retriever with the `SemanticaRetriever` class from [`semantica/integrations/langchain/retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/integrations/langchain/retriever.py).

```python
from semantica.integrations.langchain import SemanticaRetriever
from langchain.chains import RetrievalQA
from langchain.llms import OpenAI

# Create LangChain-compatible retriever

semantica_retriever = SemanticaRetriever(
    vector_store=vector_store,
    knowledge_graph=knowledge_graph,
    hybrid_alpha=0.7,
    search_kwargs={"max_results": 5}
)

# Build QA chain

qa_chain = RetrievalQA.from_chain_type(
    llm=OpenAI(model="gpt-4"),
    retriever=semantica_retriever,
    return_source_documents=True,
)

# Execute query

result = qa_chain({"query": "Explain the mechanism of action of aspirin"})

```

The `SemanticaRetriever` extends LangChain's `BaseRetriever`, making it compatible with `RetrievalQA`, `ConversationalRetrievalChain`, and other LangChain abstractions.

## Step 5: Execute Hybrid Retrieval

Execute the full GraphRAG retrieval pipeline by calling `retrieve()` with your query. The system automatically handles vector search, graph expansion, deduplication, and re-ranking.

```python

# Perform hybrid retrieval

results = retriever.retrieve(
    "What drugs target COX enzymes?",
    max_results=5,
    min_relevance_score=0.2,
)

# Process entity-aware context

for context in results:
    print(f"[{context.source}] {context.content} (score={context.score:.2f})")
    # context.content includes related entities and relationships

```

The retrieval process follows this execution flow:

1. **Vector Search** – Dense similarity search against embedded document chunks
2. **Graph Query** – Structured query against entities and relationships
3. **Multi-hop Expansion** – Traverse up to `max_expansion_hops` from graph matches using [`kg/path_finder.py`](https://github.com/semantica-agi/semantica/blob/main/kg/path_finder.py) utilities
4. **Score Normalization** – Normalize scores per source and apply `hybrid_alpha` weighting
5. **Deduplication** – Remove duplicates by entity ID or content hash
6. **Context Boosting** – Add up to 20% relevance boost for rich graph hits
7. **Re-ranking** – Final semantic similarity scoring against query embedding

## Complete Implementation Example

This end-to-end example demonstrates building a GraphRAG system for biomedical question answering:

```python
from semantica.pipeline import PipelineTemplateManager
from semantica.context import ContextRetriever
from semantica.integrations.langchain import SemanticaRetriever
from langchain.chains import RetrievalQA
from langchain_openai import ChatOpenAI

def build_graphrag_system():
    # 1. Initialize pipeline infrastructure

    manager = PipelineTemplateManager()
    
    # 2. Build vector pipeline

    rag_builder = manager.create_pipeline_from_template(
        "rag_pipeline",
        chunk={"chunk_size": 512, "overlap": 50},
        embed={"model": "text-embedding-3-large"},
        store_vectors={"store": "weaviate", "host": "localhost:8080"},
    )
    vector_store = rag_builder.get_component("store_vectors")
    
    # 3. Build or load knowledge graph

    kg_builder = manager.create_pipeline_from_template(
        "kg_construction",
        ingest_sources={"sources": ["data/drug_interactions/"]},
    )
    knowledge_graph = kg_builder.get_component("build_graph")
    
    # 4. Initialize hybrid retriever

    context_retriever = ContextRetriever(
        vector_store=vector_store,
        knowledge_graph=knowledge_graph,
        hybrid_alpha=0.6,
        max_expansion_hops=2,
    )
    
    # 5. Create LangChain integration

    retriever = SemanticaRetriever.from_context_retriever(
        context_retriever,
        search_kwargs={"max_results": 8}
    )
    
    # 6. Build RAG chain

    llm = ChatOpenAI(model="gpt-4-turbo")
    qa_system = RetrievalQA.from_chain_type(
        llm=llm,
        chain_type="stuff",
        retriever=retriever,
        return_source_documents=True,
    )
    
    return qa_system

# Execute the pipeline

system = build_graphrag_system()
response = system.invoke({"query": "Which NSAIDs interact with anticoagulants?"})
print(response["result"])

```

## Summary

- **Modular Architecture** – Semantica separates ingestion ([`pipeline_templates.py`](https://github.com/semantica-agi/semantica/blob/main/pipeline_templates.py)), retrieval ([`context_retriever.py`](https://github.com/semantica-agi/semantica/blob/main/context_retriever.py)), and integration ([`integrations/langchain/retriever.py`](https://github.com/semantica-agi/semantica/blob/main/integrations/langchain/retriever.py)) concerns for maintainable GraphRAG pipelines.
- **Hybrid Scoring** – The `ContextRetriever` uses `hybrid_alpha` to balance vector similarity against graph structure, with configurable multi-hop expansion via `max_expansion_hops`.
- **GraphStore Interface** – Knowledge graph adapters like `AgnoKnowledgeGraph` provide standardized `query()` and `get_neighbors()` methods for traversal operations.
- **LangChain Compatibility** – The `SemanticaRetriever` wrapper enables drop-in replacement of standard retrievers in existing LangChain applications.
- **Automatic Context Fusion** – Retrieved `RetrievedContext` objects contain entity-aware descriptions that incorporate relationships, reducing hallucination in generated responses.

## Frequently Asked Questions

### How does the ContextRetriever balance vector search versus graph traversal?

The `ContextRetriever` normalizes retrieval scores separately for vector and graph sources, then applies the `hybrid_alpha` parameter (0.0 to 1.0) as a weighted combination. When `hybrid_alpha` is 0.6, the final score consists of 60% graph relevance and 40% vector similarity. The system also applies context boosts up to 20% for results that contain rich graph relationships, favoring structurally connected information.

### What file handles the multi-hop graph expansion in Semantica?

Multi-hop traversal logic resides in [`semantica/kg/path_finder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/path_finder.py), which the `ContextRetriever` imports to expand initial graph matches up to `max_expansion_hops` iterations. This module implements efficient neighborhood discovery without revisiting nodes, ensuring the retrieval set grows to include semantically related entities while maintaining performance.

### Can I use Semantica GraphRAG without LangChain?

Yes. The `ContextRetriever` class operates independently of LangChain. You can call `retriever.retrieve(query)` directly and pass the resulting `RetrievedContext` list to any LLM client. The LangChain integration in [`semantica/integrations/langchain/retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/integrations/langchain/retriever.py) is purely optional and provides compatibility for existing LangChain workflows.

### Which vector stores are compatible with the Semantica GraphRAG pipeline?

Semantica supports any vector store implementing the standard `search` and `embed` interface, including Weaviate and pgvector. Configuration examples for pgvector are documented in [`docs/vector_stores/pgvector.md`](https://github.com/semantica-agi/semantica/blob/main/docs/vector_stores/pgvector.md). The vector store component is injected into `ContextRetriever` during initialization, allowing you to swap backends without modifying retrieval logic.