How to Use Query Transformations for Better RAG: A Complete Implementation Guide

Query transformations use LLM-driven techniques to rewrite, broaden, or decompose user questions before retrieval, significantly improving document relevance and reducing hallucinations in RAG pipelines.

The NirDiamant/RAG_Techniques repository provides a production-ready implementation of query transformations that you can integrate into any retrieval-augmented generation system. By reshaping queries before they hit your vector store, you can overcome lexical mismatches and capture broader contextual information that standard embedding searches often miss.

What Are Query Transformations in RAG?

Query transformations are preprocessing steps that modify the original user query using lightweight LLM calls before the retrieval phase. According to the implementation in all_rag_techniques_runnable_scripts/query_transformations.py, these transformations address three critical failure modes in standard RAG:

  • Vague or overly broad queries that return irrelevant chunks
  • Overly specific queries that miss necessary background context
  • Multi-part questions that require information from disparate document sections

The RAGQueryProcessor class orchestrates these transformations using gpt-4o by default, though you can configure any LangChain-compatible LLM.

The Three Core Query Transformation Techniques

The repository implements three complementary strategies that you can use individually or in combination.

Query Rewriting for Specificity

Query rewriting reformulates vague user inputs into detailed, retrieval-friendly versions. The transformation uses the prompt: "Rewrite the original query to be more detailed and likely to retrieve relevant information."

This technique improves lexical matching with indexed chunks by injecting domain-specific terminology and expanding abbreviations. In query_transformations.py, the re_write_llm chain handles this transformation, outputting a single enriched query string that typically yields higher similarity scores in vector searches.

Step-Back Prompting for Contextual Breadth

Step-back prompting generates a broader, more general query that captures background information the original question might miss. The prompt template specifies: "Generate a step-back query that is more general and can help retrieve background information."

This technique is particularly effective for ambiguous or multi-part questions where the user assumes implicit knowledge. The step_back_llm chain in the implementation produces a higher-level query that retrieves surrounding context, ensuring the final RAG answer has sufficient grounding in prerequisite concepts.

Sub-Query Decomposition for Complex Questions

Sub-query decomposition splits complex questions into 2-4 simpler, independent queries that can be retrieved separately. Using the prompt "Decompose the original query into simpler sub-queries," the subquery_decomposer_chain generates a list of focused questions.

This approach allows the retriever to target distinct aspects of a multi-faceted problem. The implementation parses the LLM output into a list of sub-queries, each of which can be run against the vector store independently. Results are then aggregated and deduplicated before being passed to the generation LLM, ensuring comprehensive coverage of complex user intents.

Implementation Deep Dive: RAGQueryProcessor

The RAGQueryProcessor class in all_rag_techniques_runnable_scripts/query_transformations.py encapsulates the transformation logic. When instantiated, it initializes three ChatOpenAI instances (defaulting to gpt-4o) and builds PromptTemplate objects for each transformation strategy.

The class structure follows this execution flow:

  1. Initialization: Loads environment variables (OpenAI API key) and instantiates LLM chains
  2. Template Binding: Wraps each PromptTemplate with its respective LLM using LangChain's chain syntax
  3. Execution: The run() method (or manual invocation) processes the original query through all three transformations sequentially

Here is the core implementation pattern:

from langchain_openai import ChatOpenAI
from langchain.prompts import PromptTemplate

# Initialize LLMs for each transformation

rewrite_llm = ChatOpenAI(model="gpt-4o", temperature=0)
step_back_llm = ChatOpenAI(model="gpt-4o", temperature=0)
subquery_llm = ChatOpenAI(model="gpt-4o", temperature=0)

# Define prompts

rewrite_template = """Rewrite the original query to be more detailed and likely to retrieve relevant information.
Original query: {original_query}
Rewritten query:"""

step_back_template = """Generate a step-back query that is more general and can help retrieve background information.
Original query: {original_query}
Step-back query:"""

subquery_template = """Decompose the original query into simpler sub-queries.
Original query: {original_query}
Sub-queries:"""

Practical Integration Example

To use query transformations in your RAG pipeline, instantiate the RAGQueryProcessor and feed the transformed outputs into your retrieval component. The repository provides a runnable script at all_rag_techniques_runnable_scripts/query_transformations.py that demonstrates the complete workflow.

Here is how to integrate transformations into a custom retrieval pipeline:

from all_rag_techniques_runnable_scripts.query_transformations import RAGQueryProcessor
from helper_functions import get_vector_store  # Utility from the repo

def enhanced_rag_pipeline(user_query: str):
    # Initialize processor

    processor = RAGQueryProcessor()
    
    # Generate transformed queries

    rewritten = processor.re_write_llm.invoke(
        processor.query_rewriter.format(original_query=user_query)
    ).content
    
    step_back = processor.step_back_llm.invoke(
        processor.step_back_chain.format(original_query=user_query)
    ).content
    
    # Parse sub-queries from decomposer output

    raw_subqueries = processor.subquery_decomposer_chain.invoke(
        {"original_query": user_query}
    ).content
    
    # Split into list (handles various formatting)

    sub_queries = [
        q.strip().lstrip("0123456789.- ") 
        for q in raw_subqueries.split("\n") 
        if q.strip() and not q.lower().startswith("sub-queries")
    ]
    
    # Aggregate all query variants

    all_queries = [rewritten, step_back] + sub_queries
    
    # Retrieve using all variants (example with vector store)

    vector_store = get_vector_store()
    all_results = []
    
    for query in all_queries:
        results = vector_store.similarity_search(query, k=3)
        all_results.extend(results)
    
    # Deduplicate and rank results

    seen = set()
    unique_results = []
    for doc in all_results:
        if doc.page_content not in seen:
            seen.add(doc.page_content)
            unique_results.append(doc)
    
    return unique_results

# Execute pipeline

if __name__ == "__main__":
    query = "What are the impacts of climate change on the environment?"
    documents = enhanced_rag_pipeline(query)
    print(f"Retrieved {len(documents)} unique documents")

You can also run the standalone demonstration script provided in the repository:

python all_rag_techniques_runnable_scripts/query_transformations.py --query "Your question here"

This executes the RAGQueryProcessor and prints the three transformed query variants, which you can then manually inspect or pipe into your retrieval system.

Summary

Query transformations enhance RAG pipelines by restructuring user questions before retrieval, addressing three common failure modes:

  • Query rewriting enriches vague questions with specific terminology, improving lexical matches against vector stores
  • Step-back prompting generates broader background queries that capture prerequisite context often missed by narrow questions
  • Sub-query decomposition splits complex multi-part questions into retrievable units, ensuring comprehensive coverage

The NirDiamant/RAG_Techniques repository provides a complete implementation in all_rag_techniques_runnable_scripts/query_transformations.py, featuring the RAGQueryProcessor class that orchestrates these transformations using LangChain and OpenAI models.

Frequently Asked Questions

What is the difference between query rewriting and sub-query decomposition?

Query rewriting produces a single, more detailed version of your original question, while sub-query decomposition splits the question into multiple independent queries. Rewriting improves retrieval when your original query is too vague or uses imprecise terminology, whereas decomposition handles complex questions with multiple distinct components that require separate evidence gathering.

Do query transformations add significant latency to RAG pipelines?

Yes, query transformations introduce additional LLM calls before retrieval, typically adding 500ms to 2 seconds depending on model choice and query complexity. However, the latency cost is often justified by significant improvements in retrieval accuracy and answer quality. For production systems, you can cache common query patterns or use faster models like gpt-3.5-turbo for the transformation step while reserving stronger models for final answer generation.

Can I use query transformations with non-OpenAI models?

Absolutely. The RAGQueryProcessor class uses LangChain's generic ChatOpenAI interface by default, but you can substitute any LangChain-compatible chat model including Anthropic's Claude, Google's Gemini, or local models via Ollama or HuggingFace. Simply replace the ChatOpenAI instantiations in all_rag_techniques_runnable_scripts/query_transformations.py with your preferred model class while maintaining the same prompt template structure.

How do I handle the multiple result sets from transformed queries?

When using step-back prompting or sub-query decomposition, you will retrieve documents for each query variant. The recommended approach is to aggregate all results, deduplicate based on document content or metadata IDs, and optionally rerank using a cross-encoder or reciprocal rank fusion. The example integration in this guide demonstrates a simple deduplication strategy using Python sets, but production systems should implement semantic deduplication to avoid near-duplicate chunks that share meaning but differ in wording.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →