# How to Use Query Transformations for Better RAG: A Complete Implementation Guide

> Master Query Transformations in RAG Improve document relevance and reduce hallucinations in your RAG pipelines with LLM-driven query rewriting. Implement advanced RAG techniques now.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: how-to-guide
- Published: 2026-02-19

---

**Query transformations use LLM-driven techniques to rewrite, broaden, or decompose user questions before retrieval, significantly improving document relevance and reducing hallucinations in RAG pipelines.**

The **NirDiamant/RAG_Techniques** repository provides a production-ready implementation of query transformations that you can integrate into any retrieval-augmented generation system. By reshaping queries before they hit your vector store, you can overcome lexical mismatches and capture broader contextual information that standard embedding searches often miss.

## What Are Query Transformations in RAG?

Query transformations are preprocessing steps that modify the original user query using lightweight LLM calls before the retrieval phase. According to the implementation in [`all_rag_techniques_runnable_scripts/query_transformations.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/query_transformations.py), these transformations address three critical failure modes in standard RAG:

- **Vague or overly broad queries** that return irrelevant chunks
- **Overly specific queries** that miss necessary background context
- **Multi-part questions** that require information from disparate document sections

The `RAGQueryProcessor` class orchestrates these transformations using `gpt-4o` by default, though you can configure any LangChain-compatible LLM.

## The Three Core Query Transformation Techniques

The repository implements three complementary strategies that you can use individually or in combination.

### Query Rewriting for Specificity

**Query rewriting** reformulates vague user inputs into detailed, retrieval-friendly versions. The transformation uses the prompt: *"Rewrite the original query to be more detailed and likely to retrieve relevant information."*

This technique improves lexical matching with indexed chunks by injecting domain-specific terminology and expanding abbreviations. In [`query_transformations.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/query_transformations.py), the `re_write_llm` chain handles this transformation, outputting a single enriched query string that typically yields higher similarity scores in vector searches.

### Step-Back Prompting for Contextual Breadth

**Step-back prompting** generates a broader, more general query that captures background information the original question might miss. The prompt template specifies: *"Generate a step-back query that is more general and can help retrieve background information."*

This technique is particularly effective for ambiguous or multi-part questions where the user assumes implicit knowledge. The `step_back_llm` chain in the implementation produces a higher-level query that retrieves surrounding context, ensuring the final RAG answer has sufficient grounding in prerequisite concepts.

### Sub-Query Decomposition for Complex Questions

**Sub-query decomposition** splits complex questions into 2-4 simpler, independent queries that can be retrieved separately. Using the prompt *"Decompose the original query into simpler sub-queries,"* the `subquery_decomposer_chain` generates a list of focused questions.

This approach allows the retriever to target distinct aspects of a multi-faceted problem. The implementation parses the LLM output into a list of sub-queries, each of which can be run against the vector store independently. Results are then aggregated and deduplicated before being passed to the generation LLM, ensuring comprehensive coverage of complex user intents.

## Implementation Deep Dive: RAGQueryProcessor

The `RAGQueryProcessor` class in [`all_rag_techniques_runnable_scripts/query_transformations.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/query_transformations.py) encapsulates the transformation logic. When instantiated, it initializes three `ChatOpenAI` instances (defaulting to `gpt-4o`) and builds `PromptTemplate` objects for each transformation strategy.

The class structure follows this execution flow:

1. **Initialization**: Loads environment variables (OpenAI API key) and instantiates LLM chains
2. **Template Binding**: Wraps each `PromptTemplate` with its respective LLM using LangChain's chain syntax
3. **Execution**: The `run()` method (or manual invocation) processes the original query through all three transformations sequentially

Here is the core implementation pattern:

```python
from langchain_openai import ChatOpenAI
from langchain.prompts import PromptTemplate

# Initialize LLMs for each transformation

rewrite_llm = ChatOpenAI(model="gpt-4o", temperature=0)
step_back_llm = ChatOpenAI(model="gpt-4o", temperature=0)
subquery_llm = ChatOpenAI(model="gpt-4o", temperature=0)

# Define prompts

rewrite_template = """Rewrite the original query to be more detailed and likely to retrieve relevant information.
Original query: {original_query}
Rewritten query:"""

step_back_template = """Generate a step-back query that is more general and can help retrieve background information.
Original query: {original_query}
Step-back query:"""

subquery_template = """Decompose the original query into simpler sub-queries.
Original query: {original_query}
Sub-queries:"""

```

## Practical Integration Example

To use query transformations in your RAG pipeline, instantiate the `RAGQueryProcessor` and feed the transformed outputs into your retrieval component. The repository provides a runnable script at [`all_rag_techniques_runnable_scripts/query_transformations.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/query_transformations.py) that demonstrates the complete workflow.

Here is how to integrate transformations into a custom retrieval pipeline:

```python
from all_rag_techniques_runnable_scripts.query_transformations import RAGQueryProcessor
from helper_functions import get_vector_store  # Utility from the repo

def enhanced_rag_pipeline(user_query: str):
    # Initialize processor

    processor = RAGQueryProcessor()
    
    # Generate transformed queries

    rewritten = processor.re_write_llm.invoke(
        processor.query_rewriter.format(original_query=user_query)
    ).content
    
    step_back = processor.step_back_llm.invoke(
        processor.step_back_chain.format(original_query=user_query)
    ).content
    
    # Parse sub-queries from decomposer output

    raw_subqueries = processor.subquery_decomposer_chain.invoke(
        {"original_query": user_query}
    ).content
    
    # Split into list (handles various formatting)

    sub_queries = [
        q.strip().lstrip("0123456789.- ") 
        for q in raw_subqueries.split("\n") 
        if q.strip() and not q.lower().startswith("sub-queries")
    ]
    
    # Aggregate all query variants

    all_queries = [rewritten, step_back] + sub_queries
    
    # Retrieve using all variants (example with vector store)

    vector_store = get_vector_store()
    all_results = []
    
    for query in all_queries:
        results = vector_store.similarity_search(query, k=3)
        all_results.extend(results)
    
    # Deduplicate and rank results

    seen = set()
    unique_results = []
    for doc in all_results:
        if doc.page_content not in seen:
            seen.add(doc.page_content)
            unique_results.append(doc)
    
    return unique_results

# Execute pipeline

if __name__ == "__main__":
    query = "What are the impacts of climate change on the environment?"
    documents = enhanced_rag_pipeline(query)
    print(f"Retrieved {len(documents)} unique documents")

```

You can also run the standalone demonstration script provided in the repository:

```bash
python all_rag_techniques_runnable_scripts/query_transformations.py --query "Your question here"

```

This executes the `RAGQueryProcessor` and prints the three transformed query variants, which you can then manually inspect or pipe into your retrieval system.

## Summary

Query transformations enhance RAG pipelines by restructuring user questions before retrieval, addressing three common failure modes:

- **Query rewriting** enriches vague questions with specific terminology, improving lexical matches against vector stores
- **Step-back prompting** generates broader background queries that capture prerequisite context often missed by narrow questions
- **Sub-query decomposition** splits complex multi-part questions into retrievable units, ensuring comprehensive coverage

The **NirDiamant/RAG_Techniques** repository provides a complete implementation in [`all_rag_techniques_runnable_scripts/query_transformations.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/query_transformations.py), featuring the `RAGQueryProcessor` class that orchestrates these transformations using LangChain and OpenAI models.

## Frequently Asked Questions

### What is the difference between query rewriting and sub-query decomposition?

Query rewriting produces a single, more detailed version of your original question, while sub-query decomposition splits the question into multiple independent queries. Rewriting improves retrieval when your original query is too vague or uses imprecise terminology, whereas decomposition handles complex questions with multiple distinct components that require separate evidence gathering.

### Do query transformations add significant latency to RAG pipelines?

Yes, query transformations introduce additional LLM calls before retrieval, typically adding 500ms to 2 seconds depending on model choice and query complexity. However, the latency cost is often justified by significant improvements in retrieval accuracy and answer quality. For production systems, you can cache common query patterns or use faster models like `gpt-3.5-turbo` for the transformation step while reserving stronger models for final answer generation.

### Can I use query transformations with non-OpenAI models?

Absolutely. The `RAGQueryProcessor` class uses LangChain's generic `ChatOpenAI` interface by default, but you can substitute any LangChain-compatible chat model including Anthropic's Claude, Google's Gemini, or local models via Ollama or HuggingFace. Simply replace the `ChatOpenAI` instantiations in [`all_rag_techniques_runnable_scripts/query_transformations.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/query_transformations.py) with your preferred model class while maintaining the same prompt template structure.

### How do I handle the multiple result sets from transformed queries?

When using step-back prompting or sub-query decomposition, you will retrieve documents for each query variant. The recommended approach is to aggregate all results, deduplicate based on document content or metadata IDs, and optionally rerank using a cross-encoder or reciprocal rank fusion. The example integration in this guide demonstrates a simple deduplication strategy using Python sets, but production systems should implement semantic deduplication to avoid near-duplicate chunks that share meaning but differ in wording.