# How to Integrate OCI Generative AI with Oracle Database for Production AI Applications

> Easily integrate OCI Generative AI with Oracle Database for production AI apps. Deploy a proxy, store embeddings with AI Database 26ai, and query hybrid vector indexes for RAG.

- Repository: [Oracle Developers/oracle-ai-developer-hub](https://github.com/oracle-devrel/oracle-ai-developer-hub)
- Tags: how-to-guide
- Published: 2026-05-10

---

**Integrate OCI Generative AI with Oracle Database by deploying a lightweight OpenAI-compatible proxy to access OCI LLMs, storing embeddings in Oracle AI Database 26ai using in-database ONNX inference, and querying hybrid vector indexes via Spring AI or LangChain for retrieval-augmented generation.**

The oracle-devrel/oracle-ai-developer-hub repository provides production-ready reference implementations that demonstrate how to integrate OCI Generative AI with Oracle Database for enterprise AI applications. This architecture combines OCI's managed LLM service with Oracle AI Database 26ai's unified vector and relational engine to create scalable, secure AI systems. By leveraging in-database embedding generation and hybrid search capabilities, you can build RAG pipelines that keep data within the database while utilizing OCI's generative models.

## Architecture Overview

The integration follows a three-layer architecture that separates concerns between the LLM backend, embedding store, and application logic.

**LLM Backend Layer**: Handles text generation and tool-calling through OCI GenAI endpoints (e.g., `xai.grok-4.20`) accessed via an OpenAI-compatible proxy located in [`apps/picooraclaw/oci-genai/README.md`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/picooraclaw/oci-genai/README.md).

**Embedding Store Layer**: Generates and persists dense vectors using Oracle AI Database's built-in ONNX runtime. The hybrid vector index (`DBMS_HYBRID_VECTOR.SEARCH`) combines vector similarity, keyword search, and reciprocal rank fusion as documented in [`apps/oracle-database-java-agent-memory/README.md`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/oracle-database-java-agent-memory/README.md).

**Application Logic Layer**: Orchestrates chat, memory, and retrieval-augmented generation using Spring AI or LangChain. Configuration resides in [`apps/oracle-database-java-agent-memory/src/chatserver/src/main/resources/application.yaml`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/oracle-database-java-agent-memory/src/chatserver/src/main/resources/application.yaml), supporting memory tables (`SPRING_AI_CHAT_MEMORY`) and procedural tool integration.

## Exposing OCI GenAI via an OpenAI-Compatible Proxy

The Hub provides a minimal Python proxy that translates standard OpenAI HTTP API calls into OCI-authenticated requests. This approach allows any OpenAI-compatible client (LangChain, Ollama SDKs, or Spring AI) to consume OCI GenAI without code changes.

The proxy implementation in [`apps/picooraclaw/oci-genai/proxy.py`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/picooraclaw/oci-genai/proxy.py) (approximately 70 lines) performs the following:

- Reads your `~/.oci/config` profile using environment variables `OCI_PROFILE`, `OCI_REGION`, and `OCI_COMPARTMENT_ID`
- Forwards `/v1/chat/completions` requests to the OCI GenAI REST endpoint with required OCI signature authentication
- Returns responses unchanged to the client

According to the source code in [`apps/picooraclaw/oci-genai/README.md`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/picooraclaw/oci-genai/README.md) (lines 3-5, 22-30, 78-84), the proxy uses the `oci-openai` library to handle authentication transparently.

### Starting the Proxy

```bash

# Navigate to the proxy directory

cd apps/picooraclaw/oci-genai

# Install dependencies

pip install -r requirements.txt

# Configure environment variables

export OCI_PROFILE=DEFAULT
export OCI_REGION=us-chicago-1
export OCI_COMPARTMENT_ID=ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID

# Launch the proxy on port 9999

python proxy.py

```

Once running, any application pointing to `http://localhost:9999/v1` can access OCI GenAI models using standard OpenAI client libraries.

## In-Database Embedding Generation with ONNX

Oracle AI Database 26ai includes a built-in ONNX runtime that eliminates the need for external embedding services. You can load pre-trained encoders (such as `all-MiniLM-L12-v2.onnx`) directly into the database and generate vectors via SQL.

The setup script [`apps/oracle-database-java-agent-memory/setup-hybrid-search.sql`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/oracle-database-java-agent-memory/setup-hybrid-search.sql) automates this process. First, load the ONNX model using `DBMS_MODEL.CREATE_MODEL`:

```sql
BEGIN
  DBMS_MODEL.CREATE_MODEL(
    model_name => 'EMBEDDING_MODEL',
    model_type => DBMS_MODEL.ONNX,
    model_location => 'file:/opt/oracle/dumps/all_MiniLM_L12_v2.onnx'
  );
END;
/

```

Then create a table with a vector column and hybrid index:

```sql
CREATE TABLE policy_docs (
  id   NUMBER GENERATED BY DEFAULT AS IDENTITY,
  text CLOB,
  embedding VECTOR(384)  -- Dimension matches model output
);

CREATE INDEX policy_docs_hybrid_idx ON policy_docs(embedding)
  USING HYPERLOGLOG
  WITH (
    VECTOR_SEARCH => 'ON',
    TEXT_SEARCH   => 'ON',
    RRF           => 'TRUE',
    TOP_K         => 5
  );

```

As documented in [`apps/oracle-database-java-agent-memory/README.md`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/oracle-database-java-agent-memory/README.md) (lines 20-33, 46-53), this hybrid index fuses vector similarity search with Oracle Text keyword search using Reciprocal Rank Fusion (RRF) to return the most relevant results.

## Implementing Retrieval-Augmented Generation (RAG)

With the proxy and database configured, implement RAG by connecting your application logic to both services. The typical flow involves:

1. Receiving a user prompt through Spring AI or LangChain
2. Calling the LLM via the proxy endpoint (`http://localhost:9999/v1`)
3. Executing tool calls that query the hybrid vector index using `DBMS_HYBRID_VECTOR.SEARCH`
4. Feeding retrieved passages back to the LLM for final answer generation
5. Persisting conversation history and tool results in Oracle tables for episodic memory

### Spring AI Configuration

Add the OCI GenAI starter to your `build.gradle`:

```gradle
implementation 'org.springframework.ai:spring-ai-starter-model-oci-genai'

```

Configure the connection in [`application.yaml`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/application.yaml) as shown in [`apps/oracle-database-java-agent-memory/src/chatserver/src/main/resources/application.yaml`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/oracle-database-java-agent-memory/src/chatserver/src/main/resources/application.yaml):

```yaml
spring:
  ai:
    oci:
      genai:
        chat:
          model: xai.grok-4.20
          compartment-id: ${OCI_COMPARTMENT_ID}
          region: ${OCI_REGION}
          profile: ${OCI_PROFILE}

```

### Java Hybrid Retriever Implementation

Query the hybrid index from your Spring AI retriever:

```java
@Retriever
public class OracleHybridRetriever implements DocumentRetriever {

    @Autowired 
    private JdbcTemplate jdbc;

    @Override
    public List<Document> retrieve(String query) {
        String sql = """
            SELECT text FROM policy_docs
            ORDER BY DBMS_HYBRID_VECTOR.SEARCH(
                embedding, 
                DBMS_AI_EMBEDDING.generate('EMBEDDING_MODEL', ?), 
                5
            ) DESC
            FETCH FIRST 5 ROWS ONLY
        """;
        return jdbc.query(sql, new Object[]{query},
            (rs, rowNum) -> new Document(rs.getString("text")));
    }
}

```

This pattern appears in the Java demo's `RetrievalAugmentationAdvisor` configuration (lines 29-32 of the README).

### Python Client Example

For LangChain applications, point the OpenAI client to the proxy:

```python
from langchain.llms import OpenAI
import os

os.environ["OPENAI_API_BASE"] = "http://localhost:9999/v1"
os.environ["OPENAI_API_KEY"] = "oci-genai"  # Placeholder; proxy ignores this

llm = OpenAI(model="xai.grok-4.20", temperature=0.7)
response = llm("Explain the difference between vector search and keyword search.")
print(response)

```

## Production Deployment Considerations

Deploying this architecture requires attention to scalability, security, and observability. The Hub provides specific guidance in [`apps/oci-generative-ai-jet-ui/README.md`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/oci-generative-ai-jet-ui/README.md) (lines 7-11) and [`apps/oracle-database-java-agent-memory/CLOUD_DEPLOYMENT.md`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/oracle-database-java-agent-memory/CLOUD_DEPLOYMENT.md).

**Scalability**: Deploy the proxy as a lightweight Docker container behind a load balancer. OCI GenAI scales automatically, while Oracle Autonomous Database handles vector index performance.

**Security**: Use OCI User Principal Authentication rather than API keys. Mount `~/.oci/config` as a Kubernetes Secret or use encrypted volumes. Never commit credentials to source control.

**Observability**: Enable request/response logging in the proxy using Python's `logging` module. Monitor vector index performance through Oracle Database Autonomous Health Checks.

**Failover**: Configure fallback LLM providers in [`application.yaml`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/application.yaml). Spring AI allows switching between Ollama, OpenAI, and OCI GenAI by changing the `spring.ai.<provider>` prefix without code modifications.

**Infrastructure as Code**: Reuse the Terraform scripts from `apps/oci-generative-ai-jet-ui` to provision the OCI GenAI compartment, database instance, and Kubernetes cluster as a unified stack.

## Summary

- **Deploy the OpenAI-compatible proxy** from `apps/picooraclaw/oci-genai` to expose OCI GenAI endpoints to any standard AI client library.
- **Load ONNX models directly into Oracle AI Database 26ai** using `DBMS_MODEL.CREATE_MODEL` and create hybrid vector indexes with `DBMS_HYBRID_VECTOR.SEARCH` for efficient retrieval.
- **Configure Spring AI or LangChain** to route LLM calls through the proxy while querying Oracle Database for embeddings and episodic memory storage.
- **Maintain production security** by using OCI User Principal Auth, deploying via Terraform/Kubernetes, and enabling comprehensive logging for observability.
- **Leverage the unified architecture** where Oracle Database handles both vector and relational data while OCI GenAI provides the managed LLM backend.

## Frequently Asked Questions

### What is the advantage of using an OpenAI-compatible proxy for OCI GenAI?

The proxy eliminates vendor lock-in by allowing any OpenAI-compatible client library to consume OCI GenAI without modification. As implemented in [`apps/picooraclaw/oci-genai/proxy.py`](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/apps/picooraclaw/oci-genai/proxy.py), it handles OCI's signature-based authentication transparently, enabling teams to use existing LangChain or Spring AI codebases while switching between Ollama, OpenAI, and OCI GenAI providers through configuration changes alone.

### How does Oracle Database 26ai generate embeddings without external API calls?

Oracle AI Database 26ai includes a built-in ONNX runtime that executes machine learning models directly within the database kernel. By loading an encoder model (such as `all-MiniLM-L12-v2.onnx`) using `DBMS_MODEL.CREATE_MODEL`, you can call `DBMS_CLOUD.INFERENCE` or `DBMS_AI_EMBEDDING.generate` from SQL to create vectors. This approach reduces latency, eliminates network egress costs, and keeps sensitive data within the database security perimeter.

### Can I use LangChain instead of Spring AI with this architecture?

Yes. The OpenAI-compatible proxy works with any client library that supports the OpenAI API format. For LangChain Python applications, set the `OPENAI_API_BASE` environment variable to point to the proxy URL (e.g., `http://localhost:9999/v1`). The database components remain identical regardless of whether you use LangChain, Spring AI, or direct HTTP calls, as the hybrid vector index is accessed via standard JDBC or SQL.

### What security model should I use for production deployments?

Use OCI User Principal Authentication via the `~/.oci/config` file rather than embedding API keys in your application. In Kubernetes environments, mount the OCI configuration as a Secret volume. The proxy reads these credentials at runtime, ensuring no secrets appear in environment variables or source code. Additionally, deploy the database within a private subnet and use Oracle Database Vault to restrict access to the embedding models and vector indexes.