How to Configure Amazon Neptune as a Graph Database in Cognee
To configure Amazon Neptune as a graph database in Cognee, install the optional neptune dependency, set your GRAPH_ID environment variable, and initialize both the graph and vector providers to neptune_analytics using config.set_graph_db_config() and config.set_vector_db_config().
Cognee is an open-source knowledge graph framework that abstracts graph databases as pluggable providers. When you configure Amazon Neptune as a graph database in Cognee, the framework automatically wires together a graph layer for OpenCypher queries and a hybrid vector layer for embedding-based searches, both targeting your Neptune Analytics endpoint.
Architecture Overview
Cognee’s Neptune integration consists of two specialized adapters that share a common configuration object.
NeptuneGraphDB (Graph Layer)
The NeptuneGraphDB class in cognee/infrastructure/databases/graph/neptune_driver/adapter.py implements the GraphDBInterface. It communicates with Amazon Neptune Analytics through the LangChain-AWS client (NeptuneAnalyticsGraph), translating all graph operations—node creation, edge insertion, sub-graph extraction, and metrics—into OpenCypher queries.
NeptuneAnalyticsAdapter (Hybrid Vector Layer)
The NeptuneAnalyticsAdapter in cognee/infrastructure/databases/hybrid/neptune_analytics/NeptuneAnalyticsAdapter.py mixes the graph driver with the VectorDBInterface. This adapter stores embeddings directly on Neptune graph nodes, enabling hybrid graph-vector searches via the neptune.algo.vectors.topKByEmbeddingWithFiltering procedure.
Prerequisites and Installation
Before configuring the connection, ensure you have a Neptune Analytics graph provisioned in your AWS account and note its Graph ID.
Install the optional Neptune dependency:
pip install "cognee[neptune]"
This command pulls langchain-aws, which provides the underlying NeptuneAnalyticsGraph client.
Configuration Steps
Cognee exposes configuration setters in cognee/api/v1/config/config.py. You must configure both the graph database and the vector database to use Neptune Analytics.
1. Set the Environment Variable
Store your Neptune Graph ID in a .env file at your project root:
GRAPH_ID=your-neptune-graph-id
Load this variable in your Python script using dotenv or export it manually.
2. Configure the Graph Provider
Call config.set_graph_db_config() to specify the Neptune Analytics provider and endpoint:
import os
from cognee import config
graph_endpoint = f"neptune-graph://{os.getenv('GRAPH_ID', '')}"
config.set_graph_db_config({
"graph_database_provider": "neptune_analytics",
"graph_database_url": graph_endpoint,
})
This code is implemented in cognee/api/v1/config/config.py at lines 163-168.
3. Configure the Vector Provider
For hybrid search capabilities, set the vector provider to the same Neptune endpoint:
config.set_vector_db_config({
"vector_db_provider": "neptune_analytics",
"vector_db_url": graph_endpoint,
})
This setter is defined in cognee/api/v1/config/config.py at lines 170-176.
Complete Implementation Example
The following script demonstrates a full workflow: configuring Neptune, ingesting data, extracting knowledge, and running a graph completion search. This mirrors the official example in examples/configurations/database_examples/neptune_analytics_aws_database_configuration.py.
import asyncio
import os
import pathlib
from dotenv import load_dotenv
import cognee
from cognee.modules.search.types import SearchType
load_dotenv()
async def main():
# Configure Neptune Analytics for both graph and vector operations
graph_url = f"neptune-graph://{os.getenv('GRAPH_ID', '')}"
cognee.config.set_graph_db_config({
"graph_database_provider": "neptune_analytics",
"graph_database_url": graph_url,
})
cognee.config.set_vector_db_config({
"vector_db_provider": "neptune_analytics",
"vector_db_url": graph_url,
})
# Set up local storage directories
cwd = pathlib.Path(__file__).parent
cognee.config.data_root_directory(str(cwd / "data_storage"))
cognee.config.system_root_directory(str(cwd / "cognee_system"))
# Clean previous runs
await cognee.prune.prune_data()
await cognee.prune.prune_system(metadata=True)
# Add sample data
text = """Amazon Neptune Analytics is a fast, memory-optimized graph database
designed for analytics workloads and real-time graph processing."""
await cognee.add([text], dataset_name="neptune_demo")
# Extract knowledge graph
await cognee.cognify(["neptune_demo"])
# Search the graph
results = await cognee.search(
query_type=SearchType.GRAPH_COMPLETION,
query_text="Neptune Analytics capabilities"
)
for result in results:
print(f"- {result}")
if __name__ == "__main__":
asyncio.run(main())
Hybrid vs. Graph-Only Mode
You can run Neptune in hybrid mode (graph + vectors) or graph-only mode.
-
Hybrid mode: Set both
graph_database_providerandvector_db_providertoneptune_analytics. This stores embeddings on graph nodes and enables vector similarity searches alongside graph traversals. -
Graph-only mode: Set only the graph provider to
neptune_analytics, and configure the vector provider to a different backend (e.g.,pinecone) or leave it unset.
The NeptuneAnalyticsAdapter is recognized as a hybrid provider in the HYBRID_PROVIDERS constant, which the framework uses in unit tests located at cognee/tests/unit/infrastructure/databases/test_get_unified_engine.py.
Summary
- Install the Neptune extension with
pip install "cognee[neptune]"to pull the required LangChain-AWS dependencies. - Set the
GRAPH_IDenvironment variable to your Neptune Analytics graph identifier. - Configure both graph and vector providers using
config.set_graph_db_config()andconfig.set_vector_db_config()with the provider stringneptune_analytics. - Use the endpoint format
neptune-graph://{GRAPH_ID}for both configurations. - The
NeptuneGraphDBadapter handles OpenCypher graph operations, whileNeptuneAnalyticsAdaptermanages vector embeddings and hybrid searches.
Frequently Asked Questions
What is the difference between NeptuneGraphDB and NeptuneAnalyticsAdapter?
NeptuneGraphDB (cognee/infrastructure/databases/graph/neptune_driver/adapter.py) is the core graph adapter that implements GraphDBInterface and executes OpenCypher queries via the LangChain-AWS client. NeptuneAnalyticsAdapter (cognee/infrastructure/databases/hybrid/neptune_analytics/NeptuneAnalyticsAdapter.py) extends this functionality by implementing VectorDBInterface, allowing you to store and query embeddings directly on Neptune graph nodes for hybrid search capabilities.
Can I use Amazon Neptune for only the graph database without vector storage?
Yes. If you only need graph functionality without vector search, set graph_database_provider to neptune_analytics and configure vector_db_provider to a different backend like pinecone or leave it unset. Cognee will route all graph operations to Neptune while using the specified alternative for vector storage.
What URL format does Cognee expect for Neptune Analytics connections?
Cognee expects the graph_database_url and vector_db_url to follow the format neptune-graph://{GRAPH_ID}, where GRAPH_ID is the unique identifier of your Neptune Analytics graph from the AWS console. The framework parses this URL to initialize the LangChain-AWS NeptuneAnalyticsGraph client.
Which dependency provides the Neptune connectivity in Cognee?
The langchain-aws package provides the underlying connectivity, which is included when you install Cognee with the neptune extra: pip install "cognee[neptune]". This package supplies the NeptuneAnalyticsGraph class that NeptuneGraphDB uses to communicate with your AWS endpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →