How to Configure Jina AI Embeddings with Fallback Strategies in Production RAG Systems
Configure Jina AI embeddings by initializing the async client with your API key through the embeddings factory, then set the max_retrieval_attempts threshold in AgentConfig to trigger graceful fallback responses when retrieval fails repeatedly.
The production-agentic-rag-course repository demonstrates a robust implementation of Jina AI embeddings within an agentic retrieval-augmented generation (RAG) architecture. This system converts text passages and queries into 1024-dimensional dense vectors using Jina's /embeddings endpoint, while implementing fallback strategies that prevent silent failures when external APIs timeout or return no relevant hits. Below is the complete technical guide to wiring the Jina client and tuning the degradation behavior.
Setting Up the Jina AI Embeddings Client
The embedding functionality lives in src/services/embeddings/jina_client.py as a thin async wrapper around Jina's REST API. The JinaEmbeddingsClient class handles batching, error propagation, and connection management.
Initialization and Configuration
The client requires an API key and accepts an optional base URL (defaulting to https://api.jina.ai/v1). Instantiate it through the factory to ensure environment variables are properly loaded:
from src.services.embeddings.factory import get_jina_embeddings_client
jina = get_jina_embeddings_client() # Reads settings.jina_api_key from env
The factory function in src/services/embeddings/factory.py extracts credentials from the global settings:
def get_jina_embeddings_client() -> JinaEmbeddingsClient:
api_key = settings.jina_api_key
base_url = getattr(settings, "jina_base_url", "https://api.jina.ai/v1")
return JinaEmbeddingsClient(api_key=api_key, base_url=base_url)
Store your credentials in src/config.py or via environment variables:
# .env
JINA_API_KEY=your-jina-api-key
Embedding Operations
The client exposes two primary methods. embed_passages batch-processes documents (default batch size 100), while embed_query handles single search strings:
# Batch embedding for document indexing
passages = [
"Large language models have transformed AI research.",
"Vector databases enable efficient similarity search."
]
vectors = await jina.embed_passages(passages, batch_size=100)
# Single query embedding for retrieval
query_vec = await jina.embed_query("How do vector databases work?")
All network calls raise httpx.HTTPError on failure, which the retrieval node catches to increment retry counters.
Implementing Fallback Strategies in Retrieval Nodes
The fallback logic resides in src/services/agents/nodes/retrieve_node.py. This node orchestrates the search loop and determines when to abandon retrieval attempts and return a user-friendly message.
The Fallback Trigger Mechanism
The node compares the current attempt counter against max_retrieval_attempts from the runtime context (lines 66-85):
if current_attempts >= max_retrieval_attempts:
fallback_msg = (
f"I apologize, but I couldn't find relevant research papers after {max_retrieval_attempts} attempts."
"\nThis may be because:\n"
"1. No papers in the database contain relevant information\n"
"2. The query terms don't match the indexed content\n\n"
"Please try rephrasing your question with more specific technical terms."
)
return {**updates, "messages": [AIMessage(content=fallback_msg)]}
If the limit is not reached, the node creates a retrieve_papers tool call that invokes the OpenSearch pipeline using the Jina query embedding to fetch nearest neighbors.
Configuring Fallback Parameters
Tune the degradation behavior via src/services/agents/config.py. The AgentConfig Pydantic model exposes max_retrieval_attempts with a default value of 3:
class AgentConfig(BaseModel):
"""
:param max_retrieval_attempts: Maximum number of retrieval attempts before fallback
"""
max_retrieval_attempts: int = 3
top_k: int = 5
# … other options …
Override this default through environment variables without modifying source code:
AGENT_MAX_RETRIEVAL_ATTEMPTS=5
The configuration instance injects into the LangGraph runtime, making the value accessible to the retrieval node via runtime.context.max_retrieval_attempts.
Complete Configuration Examples
Initializing the Jina Client with Custom Settings
from src.services.embeddings.factory import get_jina_embeddings_client
from src.config import settings
async def setup_embeddings():
# Ensure JINA_API_KEY is set in environment
assert settings.jina_api_key, "JINA_API_KEY not configured"
client = get_jina_embeddings_client()
# Verify connectivity
test_vec = await client.embed_query("test")
assert len(test_vec) == 1024, "Unexpected embedding dimension"
return client
Adjusting the Fallback Threshold
from src.services.agents.config import AgentConfig
# Explicit configuration
cfg = AgentConfig(max_retrieval_attempts=2, top_k=10)
# Or via environment-aware instantiation
import os
os.environ["AGENT_MAX_RETRIEVAL_ATTEMPTS"] = "4"
cfg = AgentConfig() # Automatically reads from env
End-to-End Retrieval with Fallback Handling
from src.services.agents.factory import AgenticRAGFactory
from src.services.agents.config import AgentConfig
async def execute_rag_query(user_question: str):
# Configure strict fallback (fail fast after 2 attempts)
config = AgentConfig(max_retrieval_attempts=2)
rag = await AgenticRAGFactory.create(config)
result = await rag.run({
"messages": [{"role": "user", "content": user_question}]
})
final_message = result["messages"][-1]["content"]
# final_message contains either retrieved context or the fallback apology
return final_message
Summary
- Initialize the
JinaEmbeddingsClientthroughsrc/services/embeddings/factory.pyusing theJINA_API_KEYenvironment variable. - Embed content using
embed_passagesfor batches orembed_queryfor single strings, both returning 1024-dimensional vectors. - Configure the fallback threshold via
AgentConfig.max_retrieval_attemptsinsrc/services/agents/config.py, or override with theAGENT_MAX_RETRIEVAL_ATTEMPTSenvironment variable. - Handle failures gracefully as the retrieval node in
src/services/agents/nodes/retrieve_node.pyautomatically emits explanatory messages when exceeding the configured attempt limit.
Frequently Asked Questions
How do I change the number of retrieval attempts before the fallback triggers?
Set the AGENT_MAX_RETRIEVAL_ATTEMPTS environment variable or instantiate AgentConfig with max_retrieval_attempts=N. The default is 3 attempts as defined in src/services/agents/config.py. The retrieval node checks this value against the current attempt counter before deciding whether to query OpenSearch or return the fallback message.
What happens when the Jina AI API returns an error or times out?
The JinaEmbeddingsClient raises httpx.HTTPError for network or API failures. The retrieval node catches these exceptions, increments the attempt counter, and either retries (if under the limit) or returns the fallback message defined in src/services/agents/nodes/retrieve_node.py lines 66-85. This ensures user-facing stability even when external services degrade.
Can I use a custom Jina AI base URL or self-hosted endpoint?
Yes. While the default is https://api.jina.ai/v1, you can specify jina_base_url in your settings or pass it to the JinaEmbeddingsClient constructor. The factory in src/services/embeddings/factory.py checks for this attribute using getattr(settings, "jina_base_url", default_url), allowing seamless integration with enterprise proxies or self-hosted Jina instances.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →