How to Integrate OCI Generative AI with Oracle Database for Production AI Applications
Integrate OCI Generative AI with Oracle Database by deploying a lightweight OpenAI-compatible proxy to access OCI LLMs, storing embeddings in Oracle AI Database 26ai using in-database ONNX inference, and querying hybrid vector indexes via Spring AI or LangChain for retrieval-augmented generation.
The oracle-devrel/oracle-ai-developer-hub repository provides production-ready reference implementations that demonstrate how to integrate OCI Generative AI with Oracle Database for enterprise AI applications. This architecture combines OCI's managed LLM service with Oracle AI Database 26ai's unified vector and relational engine to create scalable, secure AI systems. By leveraging in-database embedding generation and hybrid search capabilities, you can build RAG pipelines that keep data within the database while utilizing OCI's generative models.
Architecture Overview
The integration follows a three-layer architecture that separates concerns between the LLM backend, embedding store, and application logic.
LLM Backend Layer: Handles text generation and tool-calling through OCI GenAI endpoints (e.g., xai.grok-4.20) accessed via an OpenAI-compatible proxy located in apps/picooraclaw/oci-genai/README.md.
Embedding Store Layer: Generates and persists dense vectors using Oracle AI Database's built-in ONNX runtime. The hybrid vector index (DBMS_HYBRID_VECTOR.SEARCH) combines vector similarity, keyword search, and reciprocal rank fusion as documented in apps/oracle-database-java-agent-memory/README.md.
Application Logic Layer: Orchestrates chat, memory, and retrieval-augmented generation using Spring AI or LangChain. Configuration resides in apps/oracle-database-java-agent-memory/src/chatserver/src/main/resources/application.yaml, supporting memory tables (SPRING_AI_CHAT_MEMORY) and procedural tool integration.
Exposing OCI GenAI via an OpenAI-Compatible Proxy
The Hub provides a minimal Python proxy that translates standard OpenAI HTTP API calls into OCI-authenticated requests. This approach allows any OpenAI-compatible client (LangChain, Ollama SDKs, or Spring AI) to consume OCI GenAI without code changes.
The proxy implementation in apps/picooraclaw/oci-genai/proxy.py (approximately 70 lines) performs the following:
- Reads your
~/.oci/configprofile using environment variablesOCI_PROFILE,OCI_REGION, andOCI_COMPARTMENT_ID - Forwards
/v1/chat/completionsrequests to the OCI GenAI REST endpoint with required OCI signature authentication - Returns responses unchanged to the client
According to the source code in apps/picooraclaw/oci-genai/README.md (lines 3-5, 22-30, 78-84), the proxy uses the oci-openai library to handle authentication transparently.
Starting the Proxy
# Navigate to the proxy directory
cd apps/picooraclaw/oci-genai
# Install dependencies
pip install -r requirements.txt
# Configure environment variables
export OCI_PROFILE=DEFAULT
export OCI_REGION=us-chicago-1
export OCI_COMPARTMENT_ID=ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID
# Launch the proxy on port 9999
python proxy.py
Once running, any application pointing to http://localhost:9999/v1 can access OCI GenAI models using standard OpenAI client libraries.
In-Database Embedding Generation with ONNX
Oracle AI Database 26ai includes a built-in ONNX runtime that eliminates the need for external embedding services. You can load pre-trained encoders (such as all-MiniLM-L12-v2.onnx) directly into the database and generate vectors via SQL.
The setup script apps/oracle-database-java-agent-memory/setup-hybrid-search.sql automates this process. First, load the ONNX model using DBMS_MODEL.CREATE_MODEL:
BEGIN
DBMS_MODEL.CREATE_MODEL(
model_name => 'EMBEDDING_MODEL',
model_type => DBMS_MODEL.ONNX,
model_location => 'file:/opt/oracle/dumps/all_MiniLM_L12_v2.onnx'
);
END;
/
Then create a table with a vector column and hybrid index:
CREATE TABLE policy_docs (
id NUMBER GENERATED BY DEFAULT AS IDENTITY,
text CLOB,
embedding VECTOR(384) -- Dimension matches model output
);
CREATE INDEX policy_docs_hybrid_idx ON policy_docs(embedding)
USING HYPERLOGLOG
WITH (
VECTOR_SEARCH => 'ON',
TEXT_SEARCH => 'ON',
RRF => 'TRUE',
TOP_K => 5
);
As documented in apps/oracle-database-java-agent-memory/README.md (lines 20-33, 46-53), this hybrid index fuses vector similarity search with Oracle Text keyword search using Reciprocal Rank Fusion (RRF) to return the most relevant results.
Implementing Retrieval-Augmented Generation (RAG)
With the proxy and database configured, implement RAG by connecting your application logic to both services. The typical flow involves:
- Receiving a user prompt through Spring AI or LangChain
- Calling the LLM via the proxy endpoint (
http://localhost:9999/v1) - Executing tool calls that query the hybrid vector index using
DBMS_HYBRID_VECTOR.SEARCH - Feeding retrieved passages back to the LLM for final answer generation
- Persisting conversation history and tool results in Oracle tables for episodic memory
Spring AI Configuration
Add the OCI GenAI starter to your build.gradle:
implementation 'org.springframework.ai:spring-ai-starter-model-oci-genai'
Configure the connection in application.yaml as shown in apps/oracle-database-java-agent-memory/src/chatserver/src/main/resources/application.yaml:
spring:
ai:
oci:
genai:
chat:
model: xai.grok-4.20
compartment-id: ${OCI_COMPARTMENT_ID}
region: ${OCI_REGION}
profile: ${OCI_PROFILE}
Java Hybrid Retriever Implementation
Query the hybrid index from your Spring AI retriever:
@Retriever
public class OracleHybridRetriever implements DocumentRetriever {
@Autowired
private JdbcTemplate jdbc;
@Override
public List<Document> retrieve(String query) {
String sql = """
SELECT text FROM policy_docs
ORDER BY DBMS_HYBRID_VECTOR.SEARCH(
embedding,
DBMS_AI_EMBEDDING.generate('EMBEDDING_MODEL', ?),
5
) DESC
FETCH FIRST 5 ROWS ONLY
""";
return jdbc.query(sql, new Object[]{query},
(rs, rowNum) -> new Document(rs.getString("text")));
}
}
This pattern appears in the Java demo's RetrievalAugmentationAdvisor configuration (lines 29-32 of the README).
Python Client Example
For LangChain applications, point the OpenAI client to the proxy:
from langchain.llms import OpenAI
import os
os.environ["OPENAI_API_BASE"] = "http://localhost:9999/v1"
os.environ["OPENAI_API_KEY"] = "oci-genai" # Placeholder; proxy ignores this
llm = OpenAI(model="xai.grok-4.20", temperature=0.7)
response = llm("Explain the difference between vector search and keyword search.")
print(response)
Production Deployment Considerations
Deploying this architecture requires attention to scalability, security, and observability. The Hub provides specific guidance in apps/oci-generative-ai-jet-ui/README.md (lines 7-11) and apps/oracle-database-java-agent-memory/CLOUD_DEPLOYMENT.md.
Scalability: Deploy the proxy as a lightweight Docker container behind a load balancer. OCI GenAI scales automatically, while Oracle Autonomous Database handles vector index performance.
Security: Use OCI User Principal Authentication rather than API keys. Mount ~/.oci/config as a Kubernetes Secret or use encrypted volumes. Never commit credentials to source control.
Observability: Enable request/response logging in the proxy using Python's logging module. Monitor vector index performance through Oracle Database Autonomous Health Checks.
Failover: Configure fallback LLM providers in application.yaml. Spring AI allows switching between Ollama, OpenAI, and OCI GenAI by changing the spring.ai.<provider> prefix without code modifications.
Infrastructure as Code: Reuse the Terraform scripts from apps/oci-generative-ai-jet-ui to provision the OCI GenAI compartment, database instance, and Kubernetes cluster as a unified stack.
Summary
- Deploy the OpenAI-compatible proxy from
apps/picooraclaw/oci-genaito expose OCI GenAI endpoints to any standard AI client library. - Load ONNX models directly into Oracle AI Database 26ai using
DBMS_MODEL.CREATE_MODELand create hybrid vector indexes withDBMS_HYBRID_VECTOR.SEARCHfor efficient retrieval. - Configure Spring AI or LangChain to route LLM calls through the proxy while querying Oracle Database for embeddings and episodic memory storage.
- Maintain production security by using OCI User Principal Auth, deploying via Terraform/Kubernetes, and enabling comprehensive logging for observability.
- Leverage the unified architecture where Oracle Database handles both vector and relational data while OCI GenAI provides the managed LLM backend.
Frequently Asked Questions
What is the advantage of using an OpenAI-compatible proxy for OCI GenAI?
The proxy eliminates vendor lock-in by allowing any OpenAI-compatible client library to consume OCI GenAI without modification. As implemented in apps/picooraclaw/oci-genai/proxy.py, it handles OCI's signature-based authentication transparently, enabling teams to use existing LangChain or Spring AI codebases while switching between Ollama, OpenAI, and OCI GenAI providers through configuration changes alone.
How does Oracle Database 26ai generate embeddings without external API calls?
Oracle AI Database 26ai includes a built-in ONNX runtime that executes machine learning models directly within the database kernel. By loading an encoder model (such as all-MiniLM-L12-v2.onnx) using DBMS_MODEL.CREATE_MODEL, you can call DBMS_CLOUD.INFERENCE or DBMS_AI_EMBEDDING.generate from SQL to create vectors. This approach reduces latency, eliminates network egress costs, and keeps sensitive data within the database security perimeter.
Can I use LangChain instead of Spring AI with this architecture?
Yes. The OpenAI-compatible proxy works with any client library that supports the OpenAI API format. For LangChain Python applications, set the OPENAI_API_BASE environment variable to point to the proxy URL (e.g., http://localhost:9999/v1). The database components remain identical regardless of whether you use LangChain, Spring AI, or direct HTTP calls, as the hybrid vector index is accessed via standard JDBC or SQL.
What security model should I use for production deployments?
Use OCI User Principal Authentication via the ~/.oci/config file rather than embedding API keys in your application. In Kubernetes environments, mount the OCI configuration as a Secret volume. The proxy reads these credentials at runtime, ensuring no secrets appear in environment variables or source code. Additionally, deploy the database within a private subnet and use Oracle Database Vault to restrict access to the embedding models and vector indexes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →