How to Build Agentic RAG with Chain of Thought Reasoning Using Oracle AI Database

Build an Agentic RAG system with Chain of Thought reasoning by using Oracle AI Database as a unified vector store, retrieving context via OraDBVectorStore, and orchestrating single-call CoT strategies through the RAGReasoningEnsemble class.

The oracle-devrel/oracle-ai-developer-hub repository provides a complete production framework for implementing agentic retrieval-augmented generation with structured reasoning. This guide explains how to combine Oracle AI Database 26ai vector capabilities with Chain of Thought prompting to create transparent, traceable AI systems that ground LLM outputs in searchable enterprise data.

Architecture Overview

The Agentic RAG application resides in apps/agentic_rag and orchestrates three distinct layers: ingestion, vector storage, and reasoning ensemble. The ingestion layer uses pdf_processor.py (Docling-based PDF extraction) and web_processor.py (HTML scraping) to convert documents into chunked text with metadata. These chunks are embedded and stored in Oracle AI Database through OraDBVectorStore.py, which wraps the 26ai vector API to provide collection-specific query methods like query_pdf_collection and query_web_collection. The reasoning layer, implemented in rag_ensemble.py, runs the RAGReasoningEnsemble class to coordinate retrieval, prompt augmentation, and multi-strategy execution.

The data flow follows nine sequential steps. First, users upload PDFs or URLs to FastAPI endpoints defined in file_routes.py. The processors extract clean text and create chunks preserving source metadata. The OraDBVectorStore.upsert_pdf_chunk method inserts these chunks using vector embeddings. When a query arrives, the ensemble optionally executes _retrieve_context to fetch top-k chunks from the appropriate collection. The _build_augmented_prompt method concatenates retrieved context with the original question. The ensemble then launches selected reasoning strategies (CoT, ReAct, TOT) in parallel, streams ExecutionEvent objects via OraDBEventLogger, and applies similarity-based voting to determine the winning response.

Why Chain of Thought Matters

Chain of Thought (CoT) reasoning improves factuality by forcing the LLM to generate intermediate reasoning steps before producing a final answer. Unlike multi-call agentic approaches like ReAct that require separate API calls for each reasoning step, the CoT implementation in this repository operates as a single-call strategy. The LLM internally works through the problem stepwise within one completion request, significantly reducing latency while maintaining transparency.

When CoT is combined with RAG-augmented context, the model references concrete source excerpts during its reasoning process. This dramatically reduces hallucinations on multi-hop questions—such as comparing hybrid search methodologies across different database implementations—because the LLM grounds its stepwise logic in retrieved chunks rather than parametric knowledge alone.

Implementing the Oracle AI Database Vector Store

The OraDBVectorStore class in apps/agentic_rag/src/OraDBVectorStore.py provides the critical bridge between unstructured enterprise data and the reasoning ensemble. It utilizes Oracle AI Database 26ai's native vector SQL syntax (INSERT INTO ... VALUES (VECTOR(...))) to store embeddings alongside rich metadata including source URLs, page numbers, and collection identifiers.

Each collection (PDF, Web, etc.) exposes dedicated query methods. The query_pdf_collection and query_web_collection methods execute similarity searches returning ranked chunks with associated similarity scores. This enables the RAGReasoningEnsemble._retrieve_context method to calculate average similarity metrics and filter low-relevance content before augmentation.

from apps.agentic_rag.src.OraDBVectorStore import OraDBVectorStore

vector_store = OraDBVectorStore(
    dsn="oracle_ai_db_high",
    user="AI_USER",
    password="secure_password",
    collection="PDF"
)

# Retrieve relevant chunks for CoT augmentation

results = vector_store.query_pdf_collection(
    query_embedding=embedding_vector,
    top_k=5
)

The system also implements OraDBEventLogger for compliance and debugging. This class persists every ExecutionEvent to the database—enabling full audit trails of which chunks were retrieved, which reasoning strategies were invoked, and how the ensemble voted on the final answer.

Building the Reasoning Ensemble with CoT

The RAGReasoningEnsemble class defined in apps/agentic_rag/src/reasoning/rag_ensemble.py orchestrates the entire agentic workflow. When use_rag=True and strategies=["cot"] are specified, the ensemble executes four critical phases: retrieval, augmentation, execution, and voting.

During retrieval, the ensemble calls _retrieve_context to query the vector store and compute aggregated similarity scores. The augmentation phase uses _build_augmented_prompt to prepend retrieved context blocks to the original query, creating a "knowledge-enhanced" prompt. The execution phase launches the CoT strategy (and any others specified) in parallel through the run method. Each strategy runs as a separate LLM call managed by the ReasoningEnsemble base class. Finally, the voting phase clusters similar responses and returns the majority winner through the ReasoningResult dataclass.

import asyncio
from apps.agentic_rag.src.reasoning.rag_ensemble import RAGReasoningEnsemble

async def run_cot_reasoning():
    ensemble = RAGReasoningEnsemble(
        model_name="gemma3:270m",
        vector_store=vector_store,
        event_logger=logger
    )
    
    result = await ensemble.run(
        query="What are the benefits of hybrid vector-keyword search?",
        strategies=["cot"],
        use_rag=True,
        collection="PDF"
    )
    
    return result.winner["response"]

Complete Implementation Example

The following implementation demonstrates end-to-end CoT-enabled Agentic RAG with event logging and context inspection:

import asyncio
from apps.agentic_rag.src.reasoning.rag_ensemble import RAGReasoningEnsemble
from apps.agentic_rag.src.OraDBVectorStore import OraDBVectorStore
from apps.agentic_rag.src.OraDBEventLogger import OraDBEventLogger

async def main():
    # Initialize Oracle AI Database vector store

    vector_store = OraDBVectorStore(
        dsn="oracle_ai_db_high",
        user="AI_USER",
        password="******",
        collection="PDF"
    )

    # Enable persistent event logging for traceability

    logger = OraDBEventLogger(
        dsn="oracle_ai_db_high", 
        user="AI_USER", 
        password="******"
    )

    # Build the reasoning ensemble with CoT strategy

    ensemble = RAGReasoningEnsemble(
        model_name="gemma3:270m",
        vector_store=vector_store,
        event_logger=logger,
    )

    # Execute CoT reasoning with RAG context

    result = await ensemble.run(
        query="What are the benefits of hybrid vector-keyword search in Oracle AI Database?",
        strategies=["cot"],
        use_rag=True,
        collection="PDF",
    )

    # Output final answer

    print("\n=== Final Answer ===")
    print(result.winner["response"])

    # Display retrieved source chunks

    print("\n=== Retrieved Chunks ===")
    for i, chunk in enumerate(result.rag_context["chunks"], 1):
        print(f"\n--- Chunk {i} (score {chunk['score']:.2f}) ---")
        print(chunk["content"])

    # Display execution trace

    print("\n=== Execution Trace ===")
    for ev in result.execution_trace:
        print(f"[{ev.timestamp}] {ev.event_type.upper()}: {ev.message}")

asyncio.run(main())

This example initializes the 26ai vector store connection, enables audit logging, and executes a single-call CoT strategy grounded in retrieved PDF chunks. The result.winner contains the final answer after ensemble voting, while result.rag_context provides the source material for verification.

Enabling CoT in the Web Interface

The Gradio interface in apps/agentic_rag/gradio_app.py exposes a "Chain of Thought" toggle that controls the reasoning pathway. When enabled, the backend automatically sets use_rag=True and includes "cot" in the strategy list passed to RAGReasoningEnsemble.run.


# Excerpt from gradio_app.py

def chat_with_documents(question, enable_cot, collection):
    strategies = ["cot"] if enable_cot else ["standard"]
    result = asyncio.run(
        ensemble.run(
            query=question,
            strategies=strategies,
            use_rag=enable_cot,
            collection=collection,
        )
    )
    return result.winner["response"]

The UI streams ExecutionEvent objects through run_with_streaming, displaying icons (🔗 for CoT) and live timestamps as the reasoning progresses. This provides immediate visibility into which chunks were retrieved and how the ensemble reached its conclusion.

Key Files and References

File Purpose
apps/agentic_rag/src/reasoning/rag_ensemble.py Core RAGReasoningEnsemble class orchestrating retrieval, augmentation, and voting
apps/agentic_rag/src/OraDBVectorStore.py 26ai vector API wrapper with collection-specific query methods
apps/agentic_rag/src/pdf_processor.py Docling-based PDF text extraction and chunking
apps/agentic_rag/src/web_processor.py HTML ingestion and metadata preservation
apps/agentic_rag/src/agents/agent_factory.py Factory constructing CoTAgent and other reasoning agents
notebooks/agentic_rag_langchain_oracledb_demo.ipynb End-to-end tutorial demonstrating CoT workflow

Summary

  • Oracle AI Database 26ai serves as the unified vector and metadata store through OraDBVectorStore, supporting similarity search with SQL-native vector operations.
  • Chain of Thought reasoning operates as a single-call strategy within RAGReasoningEnsemble, reducing latency compared to multi-step approaches while maintaining reasoning transparency.
  • The architecture automatically augments CoT prompts with retrieved context from query_pdf_collection or query_web_collection, grounding answers in enterprise data.
  • Event logging via OraDBEventLogger provides complete audit trails of retrieval, reasoning, and voting phases for compliance and debugging.
  • Toggle CoT reasoning through the UI switch or programmatically via strategies=["cot"] and use_rag=True parameters.

Frequently Asked Questions

What is the difference between CoT and ReAct reasoning in this implementation?

CoT (Chain of Thought) executes step-by-step reasoning within a single LLM completion call, making it faster and more cost-effective for complex queries. ReAct (Reasoning + Acting) requires multiple sequential API calls to interleave reasoning with tool use or retrieval actions. According to the oracle-devrel/oracle-ai-developer-hub source code, CoT is preferred when latency is critical, while ReAct offers more granular control for multi-tool scenarios.

How does Oracle AI Database store vector embeddings?

The database utilizes the 26ai vector API with native SQL syntax. The OraDBVectorStore class inserts embeddings using INSERT INTO ... VALUES (VECTOR(...)) statements, storing vectors alongside metadata fields in unified tables. This enables hybrid queries combining vector similarity with traditional SQL filters on metadata columns.

Can I use multiple reasoning strategies simultaneously?

Yes. The RAGReasoningEnsemble.run method accepts a list of strategies such as strategies=["cot", "tot", "react"]. The ensemble executes these in parallel, then applies similarity-based clustering to group semantically equivalent responses. The cluster with the most votes determines the final winner, providing robustness against individual strategy failures.

How do I trace the reasoning steps for debugging?

The system emits ExecutionEvent objects throughout the pipeline via OraDBEventLogger. When using run_with_streaming, these events stream to the UI in real-time, showing retrieval timestamps, chunk similarity scores, and reasoning phase completions. The events are also persisted to Oracle AI Database tables for historical analysis and compliance auditing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →