# How to Configure Repository Pool Cleanup Interval and Max Idle Time in Llama-GitHub

> Learn to configure repository pool cleanup interval and max idle time in llama-github. Optimize your setup easily with these essential settings. Get clear instructions now.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: how-to-guide
- Published: 2026-03-04

---

**To configure the repository pool cleanup interval and max idle time in llama-github, pass `cleanup_interval` and `max_idle_time` directly to the `RepositoryPool` constructor, or use the `repo_cleanup_interval` and `repo_max_idle_time` keyword arguments when initializing the `GithubRAG` class.**

The `RepositoryPool` singleton in `jetxu-llm/llama-github` caches `Repository` objects to minimize API calls, using a background thread that periodically purges stale entries. You can tune how aggressively this cleanup runs by adjusting two key parameters at initialization, either through the low-level pool API or via the high-level `GithubRAG` wrapper.

## Understanding the Repository Pool Parameters

The cleanup behavior is controlled by two integer values representing seconds:

- **`cleanup_interval`**: Specifies how often the background thread wakes to scan for idle repositories. Default is `3600` seconds (1 hour).
- **`max_idle_time`**: Defines the maximum duration a repository can remain unused before its cached data is evicted. Default is `86400` seconds (24 hours).

These parameters are defined in the `RepositoryPool` class within [`llama_github/data_retrieval/github_entities.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/data_retrieval/github_entities.py). The constructor stores them as instance attributes at lines 812–822, and the private `_cleanup` loop uses them to determine when to purge caches. When the `GithubRAG` wrapper instantiates the pool, it forwards any keyword arguments named `repo_cleanup_interval` and `repo_max_idle_time` to these parameters according to the mapping logic at lines 87–96 in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py).

## Direct RepositoryPool Configuration

For applications using the pool directly without the `GithubRAG` wrapper, instantiate `RepositoryPool` with your desired timing values:

```python
from llama_github.data_retrieval.github_entities import RepositoryPool
from llama_github.github_integration.github_auth_manager import ExtendedGithub

# Authenticated Github client

github_client: ExtendedGithub = ...

# Configure 5-minute cleanup checks and 30-minute idle timeout

pool = RepositoryPool(
    github_instance=github_client,
    cleanup_interval=300,      # 5 minutes

    max_idle_time=1800         # 30 minutes

)

repo = pool.get_repository("owner/repo")

```

This approach gives you precise control over memory usage versus re-fetch latency. Lower values reduce memory footprint but increase GitHub API calls, while higher values improve performance for repeated accesses at the cost of retained memory.

## Configuration via GithubRAG

When using the high-level `GithubRAG` interface, pass the parameters with the `repo_` prefix. The wrapper automatically maps these to the underlying pool constructor:

```python
from llama_github.github_rag import GithubRAG

rag = GithubRAG(
    github_access_token="ghp_XXXXXXXXXXXXXXXXXXXX",
    repo_cleanup_interval=600,   # Cleanup every 10 minutes

    repo_max_idle_time=7200      # Evict after 2 hours idle

)

# Access the configured pool

repo = rag.RepositoryPool.get_repository("owner/repo")

```

This indirection allows you to configure pool behavior without directly importing `RepositoryPool`, keeping your application code focused on retrieval-augmented generation tasks while still controlling resource management.

## Testing with Aggressive Cleanup

The test suite demonstrates how to verify cleanup logic with accelerated timing. In [`tests/test_data_retrieval.py`](https://github.com/jetxu-llm/llama-github/blob/main/tests/test_data_retrieval.py) at lines 65–70, the tests instantiate the pool with minimal delays to ensure the eviction logic triggers quickly:

```python
from unittest.mock import MagicMock
from llama_github.data_retrieval.github_entities import RepositoryPool

mock_github = MagicMock()
pool = RepositoryPool(
    mock_github,
    cleanup_interval=0.1,   # 100ms between checks

    max_idle_time=0.1       # 100ms idle threshold

)

```

Use this pattern in your own integration tests when you need to validate that repository data refreshes correctly without waiting for the default 24-hour cycle.

## Summary

- The `RepositoryPool` in [`llama_github/data_retrieval/github_entities.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/data_retrieval/github_entities.py) accepts `cleanup_interval` and `max_idle_time` parameters to control cache eviction.
- Default values are 3600 seconds (1 hour) for cleanup frequency and 86400 seconds (24 hours) for maximum idle time.
- The `GithubRAG` class in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) forwards `repo_cleanup_interval` and `repo_max_idle_time` kwargs to the pool constructor.
- Adjust these values based on your memory constraints and acceptable latency for re-fetching repository metadata.

## Frequently Asked Questions

### What are the default cleanup interval and max idle time values?

The default `cleanup_interval` is 3600 seconds (1 hour), meaning the background thread scans for stale entries once per hour. The default `max_idle_time` is 86400 seconds (24 hours), so repositories unused for a full day are evicted from the cache. These defaults balance memory efficiency against the cost of re-fetching repository data from the GitHub API.

### How do I completely disable repository pool cleanup?

There is no explicit "disable" flag in the source code. To effectively disable cleanup, set `cleanup_interval` to a very large value (e.g., `31536000` for one year) and `max_idle_time` to an equally large value. This prevents the `_cleanup` loop from purging entries during your application's lifetime, though this will increase memory usage proportionally to the number of unique repositories accessed.

### Can I change the cleanup interval after the pool is created?

No, the `cleanup_interval` and `max_idle_time` are set during `RepositoryPool` initialization and stored as instance attributes. To change these values, you must create a new `RepositoryPool` instance with the desired parameters. Note that `RepositoryPool` is typically managed as a singleton by `GithubRAG`, so you would need to re-initialize the `GithubRAG` object to apply new timing values.

### Where is the cleanup logic implemented in the source code?

The cleanup loop is implemented in the private `_cleanup` method of the `RepositoryPool` class in [`llama_github/data_retrieval/github_entities.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/data_retrieval/github_entities.py). This method runs in a background thread, sleeping for `self.cleanup_interval` seconds between iterations, and removes repository entries whose last access time exceeds `self.max_idle_time`. The mapping of user-facing parameters to these instance variables occurs in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) at lines 87–96.