How to Build Agentic RAG with Chain of Thought Reasoning Using Oracle AI Database
Build an Agentic RAG system with Chain of Thought reasoning by using Oracle AI Database as a unified vector store, retrieving context via OraDBVectorStore, and orchestrating single-call CoT strategies through the RAGReasoningEnsemble class.
The oracle-devrel/oracle-ai-developer-hub repository provides a complete production framework for implementing agentic retrieval-augmented generation with structured reasoning. This guide explains how to combine Oracle AI Database 26ai vector capabilities with Chain of Thought prompting to create transparent, traceable AI systems that ground LLM outputs in searchable enterprise data.
Architecture Overview
The Agentic RAG application resides in apps/agentic_rag and orchestrates three distinct layers: ingestion, vector storage, and reasoning ensemble. The ingestion layer uses pdf_processor.py (Docling-based PDF extraction) and web_processor.py (HTML scraping) to convert documents into chunked text with metadata. These chunks are embedded and stored in Oracle AI Database through OraDBVectorStore.py, which wraps the 26ai vector API to provide collection-specific query methods like query_pdf_collection and query_web_collection. The reasoning layer, implemented in rag_ensemble.py, runs the RAGReasoningEnsemble class to coordinate retrieval, prompt augmentation, and multi-strategy execution.
The data flow follows nine sequential steps. First, users upload PDFs or URLs to FastAPI endpoints defined in file_routes.py. The processors extract clean text and create chunks preserving source metadata. The OraDBVectorStore.upsert_pdf_chunk method inserts these chunks using vector embeddings. When a query arrives, the ensemble optionally executes _retrieve_context to fetch top-k chunks from the appropriate collection. The _build_augmented_prompt method concatenates retrieved context with the original question. The ensemble then launches selected reasoning strategies (CoT, ReAct, TOT) in parallel, streams ExecutionEvent objects via OraDBEventLogger, and applies similarity-based voting to determine the winning response.
Why Chain of Thought Matters
Chain of Thought (CoT) reasoning improves factuality by forcing the LLM to generate intermediate reasoning steps before producing a final answer. Unlike multi-call agentic approaches like ReAct that require separate API calls for each reasoning step, the CoT implementation in this repository operates as a single-call strategy. The LLM internally works through the problem stepwise within one completion request, significantly reducing latency while maintaining transparency.
When CoT is combined with RAG-augmented context, the model references concrete source excerpts during its reasoning process. This dramatically reduces hallucinations on multi-hop questions—such as comparing hybrid search methodologies across different database implementations—because the LLM grounds its stepwise logic in retrieved chunks rather than parametric knowledge alone.
Implementing the Oracle AI Database Vector Store
The OraDBVectorStore class in apps/agentic_rag/src/OraDBVectorStore.py provides the critical bridge between unstructured enterprise data and the reasoning ensemble. It utilizes Oracle AI Database 26ai's native vector SQL syntax (INSERT INTO ... VALUES (VECTOR(...))) to store embeddings alongside rich metadata including source URLs, page numbers, and collection identifiers.
Each collection (PDF, Web, etc.) exposes dedicated query methods. The query_pdf_collection and query_web_collection methods execute similarity searches returning ranked chunks with associated similarity scores. This enables the RAGReasoningEnsemble._retrieve_context method to calculate average similarity metrics and filter low-relevance content before augmentation.
from apps.agentic_rag.src.OraDBVectorStore import OraDBVectorStore
vector_store = OraDBVectorStore(
dsn="oracle_ai_db_high",
user="AI_USER",
password="secure_password",
collection="PDF"
)
# Retrieve relevant chunks for CoT augmentation
results = vector_store.query_pdf_collection(
query_embedding=embedding_vector,
top_k=5
)
The system also implements OraDBEventLogger for compliance and debugging. This class persists every ExecutionEvent to the database—enabling full audit trails of which chunks were retrieved, which reasoning strategies were invoked, and how the ensemble voted on the final answer.
Building the Reasoning Ensemble with CoT
The RAGReasoningEnsemble class defined in apps/agentic_rag/src/reasoning/rag_ensemble.py orchestrates the entire agentic workflow. When use_rag=True and strategies=["cot"] are specified, the ensemble executes four critical phases: retrieval, augmentation, execution, and voting.
During retrieval, the ensemble calls _retrieve_context to query the vector store and compute aggregated similarity scores. The augmentation phase uses _build_augmented_prompt to prepend retrieved context blocks to the original query, creating a "knowledge-enhanced" prompt. The execution phase launches the CoT strategy (and any others specified) in parallel through the run method. Each strategy runs as a separate LLM call managed by the ReasoningEnsemble base class. Finally, the voting phase clusters similar responses and returns the majority winner through the ReasoningResult dataclass.
import asyncio
from apps.agentic_rag.src.reasoning.rag_ensemble import RAGReasoningEnsemble
async def run_cot_reasoning():
ensemble = RAGReasoningEnsemble(
model_name="gemma3:270m",
vector_store=vector_store,
event_logger=logger
)
result = await ensemble.run(
query="What are the benefits of hybrid vector-keyword search?",
strategies=["cot"],
use_rag=True,
collection="PDF"
)
return result.winner["response"]
Complete Implementation Example
The following implementation demonstrates end-to-end CoT-enabled Agentic RAG with event logging and context inspection:
import asyncio
from apps.agentic_rag.src.reasoning.rag_ensemble import RAGReasoningEnsemble
from apps.agentic_rag.src.OraDBVectorStore import OraDBVectorStore
from apps.agentic_rag.src.OraDBEventLogger import OraDBEventLogger
async def main():
# Initialize Oracle AI Database vector store
vector_store = OraDBVectorStore(
dsn="oracle_ai_db_high",
user="AI_USER",
password="******",
collection="PDF"
)
# Enable persistent event logging for traceability
logger = OraDBEventLogger(
dsn="oracle_ai_db_high",
user="AI_USER",
password="******"
)
# Build the reasoning ensemble with CoT strategy
ensemble = RAGReasoningEnsemble(
model_name="gemma3:270m",
vector_store=vector_store,
event_logger=logger,
)
# Execute CoT reasoning with RAG context
result = await ensemble.run(
query="What are the benefits of hybrid vector-keyword search in Oracle AI Database?",
strategies=["cot"],
use_rag=True,
collection="PDF",
)
# Output final answer
print("\n=== Final Answer ===")
print(result.winner["response"])
# Display retrieved source chunks
print("\n=== Retrieved Chunks ===")
for i, chunk in enumerate(result.rag_context["chunks"], 1):
print(f"\n--- Chunk {i} (score {chunk['score']:.2f}) ---")
print(chunk["content"])
# Display execution trace
print("\n=== Execution Trace ===")
for ev in result.execution_trace:
print(f"[{ev.timestamp}] {ev.event_type.upper()}: {ev.message}")
asyncio.run(main())
This example initializes the 26ai vector store connection, enables audit logging, and executes a single-call CoT strategy grounded in retrieved PDF chunks. The result.winner contains the final answer after ensemble voting, while result.rag_context provides the source material for verification.
Enabling CoT in the Web Interface
The Gradio interface in apps/agentic_rag/gradio_app.py exposes a "Chain of Thought" toggle that controls the reasoning pathway. When enabled, the backend automatically sets use_rag=True and includes "cot" in the strategy list passed to RAGReasoningEnsemble.run.
# Excerpt from gradio_app.py
def chat_with_documents(question, enable_cot, collection):
strategies = ["cot"] if enable_cot else ["standard"]
result = asyncio.run(
ensemble.run(
query=question,
strategies=strategies,
use_rag=enable_cot,
collection=collection,
)
)
return result.winner["response"]
The UI streams ExecutionEvent objects through run_with_streaming, displaying icons (🔗 for CoT) and live timestamps as the reasoning progresses. This provides immediate visibility into which chunks were retrieved and how the ensemble reached its conclusion.
Key Files and References
| File | Purpose |
|---|---|
apps/agentic_rag/src/reasoning/rag_ensemble.py |
Core RAGReasoningEnsemble class orchestrating retrieval, augmentation, and voting |
apps/agentic_rag/src/OraDBVectorStore.py |
26ai vector API wrapper with collection-specific query methods |
apps/agentic_rag/src/pdf_processor.py |
Docling-based PDF text extraction and chunking |
apps/agentic_rag/src/web_processor.py |
HTML ingestion and metadata preservation |
apps/agentic_rag/src/agents/agent_factory.py |
Factory constructing CoTAgent and other reasoning agents |
notebooks/agentic_rag_langchain_oracledb_demo.ipynb |
End-to-end tutorial demonstrating CoT workflow |
Summary
- Oracle AI Database 26ai serves as the unified vector and metadata store through
OraDBVectorStore, supporting similarity search with SQL-native vector operations. - Chain of Thought reasoning operates as a single-call strategy within
RAGReasoningEnsemble, reducing latency compared to multi-step approaches while maintaining reasoning transparency. - The architecture automatically augments CoT prompts with retrieved context from
query_pdf_collectionorquery_web_collection, grounding answers in enterprise data. - Event logging via
OraDBEventLoggerprovides complete audit trails of retrieval, reasoning, and voting phases for compliance and debugging. - Toggle CoT reasoning through the UI switch or programmatically via
strategies=["cot"]anduse_rag=Trueparameters.
Frequently Asked Questions
What is the difference between CoT and ReAct reasoning in this implementation?
CoT (Chain of Thought) executes step-by-step reasoning within a single LLM completion call, making it faster and more cost-effective for complex queries. ReAct (Reasoning + Acting) requires multiple sequential API calls to interleave reasoning with tool use or retrieval actions. According to the oracle-devrel/oracle-ai-developer-hub source code, CoT is preferred when latency is critical, while ReAct offers more granular control for multi-tool scenarios.
How does Oracle AI Database store vector embeddings?
The database utilizes the 26ai vector API with native SQL syntax. The OraDBVectorStore class inserts embeddings using INSERT INTO ... VALUES (VECTOR(...)) statements, storing vectors alongside metadata fields in unified tables. This enables hybrid queries combining vector similarity with traditional SQL filters on metadata columns.
Can I use multiple reasoning strategies simultaneously?
Yes. The RAGReasoningEnsemble.run method accepts a list of strategies such as strategies=["cot", "tot", "react"]. The ensemble executes these in parallel, then applies similarity-based clustering to group semantically equivalent responses. The cluster with the most votes determines the final winner, providing robustness against individual strategy failures.
How do I trace the reasoning steps for debugging?
The system emits ExecutionEvent objects throughout the pipeline via OraDBEventLogger. When using run_with_streaming, these events stream to the UI in real-time, showing retrieval timestamps, chunk similarity scores, and reasoning phase completions. The events are also persisted to Oracle AI Database tables for historical analysis and compliance auditing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →