# How Asynchronous Context Retrieval Works in Jupyter Notebooks in llama-github

> Discover how asynchronous context retrieval in Jupyter notebooks prevents event loop errors with GithubRAG and nest_asyncio for seamless parallel RAG pipelines.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: internals
- Published: 2026-03-04

---

**The `GithubRAG` class automatically detects Jupyter notebook environments and applies `nest_asyncio` to execute parallel RAG pipelines without triggering "event loop is already running" errors.**

Asynchronous context retrieval enables the `llama-github` library to perform parallel searches across GitHub code, issues, and repositories, but Jupyter notebooks present unique challenges due to their persistent IPython event loop. The library handles this through runtime detection logic in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py), allowing researchers to leverage full async performance without manual loop management.

## Dual API Design: Async and Sync Entry Points

The `GithubRAG` class provides two distinct methods for context retrieval to accommodate different execution environments.

**`async_retrieve_context`** is the core coroutine that executes the complete RAG pipeline—including question analysis, Google searches, and GitHub API queries—as a true asynchronous operation using `asyncio.create_task` for parallel execution.

**`retrieve_context`** acts as a universal wrapper that inspects the current runtime and executes the async pipeline appropriately. This method handles the complexity of Jupyter notebooks, existing event loops, and plain Python scripts transparently.

## Runtime Detection and Jupyter Handling

When `retrieve_context` is invoked, the library first determines the execution context before scheduling the async work. In [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) (lines 24-28), the code checks for an IPython kernel trait via `get_ipython()` to identify Jupyter notebook environments.

If the detection returns a valid kernel, the library imports `nest_asyncio` and calls `nest_asyncio.apply()` to patch the running event loop. This critical step allows nested execution of `asyncio.run()` within Jupyter's already-active loop, preventing the standard library from raising a runtime error.

## Execution Flow in retrieve_context

The synchronous wrapper follows a deterministic decision tree based on the current event loop state:

1. **Determine effective `simple_mode`** (line 20) to decide whether to use the full RAG pipeline or a simplified retrieval path.

2. **Obtain the current event loop** via `self.loop = asyncio.get_event_loop()` to inspect the runtime state.

3. **Detect Jupyter notebooks** by verifying the presence of an IPython kernel trait (lines 24-26).

4. **If running in Jupyter**: Import `nest_asyncio`, apply the patch (line 27), and execute the coroutine with `asyncio.run(self.async_retrieve_context(...))` (line 28).

5. **If an external asyncio loop is already running** (e.g., when called inside another async function), return a `Future` via `asyncio.ensure_future` (line 30) to allow awaiting the result elsewhere.

6. **For plain Python scripts**, run the coroutine to completion using `self.loop.run_until_complete` (line 31).

## Parallel Async Pipeline Architecture

The `async_retrieve_context` method orchestrates multiple I/O-bound operations concurrently. Rather than executing Google searches, code searches, issue lookups, and repository queries sequentially, the implementation creates separate tasks using `asyncio.create_task` for each search domain.

These tasks are then awaited simultaneously using `await asyncio.gather(...)`, dramatically reducing overall latency compared to synchronous execution. The asynchronous HTTP client defined in [`llama_github/utils.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/utils.py) supports these parallel operations, while [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) handles the subsequent ranking and context assembly.

## Implementation Examples

The following pattern works seamlessly in Jupyter notebooks without requiring manual async syntax:

```python
from llama_github.github_rag import GithubRAG

# Initialize the RAG object

rag = GithubRAG(github_access_token="YOUR_TOKEN", openai_api_key="YOUR_OPENAI_KEY")

# Synchronous call - automatically handles Jupyter's event loop

contexts = rag.retrieve_context(
    query="How does the async context retrieval work in Jupyter?",
    simple_mode=False
)

for i, ctx in enumerate(contexts, 1):
    print(f"--- Context {i} ---")
    print(ctx)

```

For use inside existing async functions or standalone scripts, await the coroutine directly:

```python
async def fetch_context():
    rag = GithubRAG(github_access_token="TOKEN")
    contexts = await rag.async_retrieve_context("Explain async retrieval")
    return contexts

```

## Key Source Files

- **[`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py)**: Contains the `GithubRAG` class with `retrieve_context` (lines 20-31) and `async_retrieve_context` implementations.
- **[`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py)**: Orchestrates question analysis, search criteria generation, and context ranking.
- **[`llama_github/utils.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/utils.py)**: Provides `AsyncHTTPClient` for non-blocking HTTP requests to GitHub and Google APIs.
- **[`llama_github/logger.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/logger.py)**: Centralized logging infrastructure used throughout the async execution flow.

## Summary

- **Dual API design**: `async_retrieve_context` provides the raw coroutine, while `retrieve_context` offers a runtime-aware synchronous wrapper.
- **Jupyter detection**: The library inspects `get_ipython()` to identify notebook kernels and applies `nest_asyncio` automatically.
- **Event loop handling**: Three execution paths support Jupyter notebooks (patched `asyncio.run`), nested async contexts (`asyncio.ensure_future`), and plain scripts (`run_until_complete`).
- **Parallel execution**: The async pipeline uses `asyncio.gather` to run Google, code, issue, and repo searches concurrently, minimizing retrieval latency.

## Frequently Asked Questions

### Why does Jupyter need special handling for asynchronous context retrieval?

Jupyter notebooks run an IPython kernel that maintains a persistent event loop to support interactive widgets and magic commands. Standard `asyncio.run()` calls fail in this environment because the library cannot create a new event loop when one is already active. The `llama-github` source code detects this condition and applies `nest_asyncio` to enable nested loop execution.

### What is nest_asyncio and is it safe to use?

`nest_asyncio` is a library that patches Python's asyncio to allow nested execution of event loops. When `retrieve_context` detects a Jupyter environment in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) (line 27), it calls `nest_asyncio.apply()` to modify the loop policy temporarily. This is safe for RAG retrieval workflows and prevents the "event loop is already running" RuntimeError without disrupting the notebook's existing async infrastructure.

### Can I use async_retrieve_context directly in a Jupyter notebook?

While you can technically define `await rag.async_retrieve_context()` in a notebook cell, you risk encountering event loop conflicts if the cell is executed multiple times or if other async operations are pending. The recommended approach is using `rag.retrieve_context()`, which handles all edge cases—including automatic `nest_asyncio` application—ensuring consistent behavior across execution contexts.

### How does the library handle performance in different environments?

Regardless of whether you call `retrieve_context` in Jupyter, a standalone script, or an async web server, the underlying `async_retrieve_context` pipeline maintains the same parallelism through `asyncio.gather`. The wrapper method merely adapts the entry point to the host environment's event loop requirements, ensuring optimal I/O concurrency without blocking the main thread.