# How interviewstreet/hiring-agent Implements GitHub API Rate Limiting with Proactive Waiting

> Learn how interviewstreet/hiring-agent prevents HTTP 403 errors by proactively waiting for GitHub API rate limits. Discover the `X-RateLimit-Remaining` header strategy.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-08

---

**The hiring-agent repository prevents HTTP 403 errors by implementing a proactive waiting strategy in `_fetch_github_api` that monitors `X-RateLimit-Remaining` headers and automatically sleeps until the reset time when fewer than 10 requests remain.**

The `interviewstreet/hiring-agent` project retrieves GitHub user profiles and repositories through the REST API, requiring robust handling of GitHub's strict hourly quotas. This article examines how the repository implements **GitHub API rate limiting with proactive waiting** to ensure uninterrupted data fetching without triggering rate limit violations.

## Understanding GitHub API Rate Limits

GitHub enforces hourly request quotas that vary by authentication status. **Unauthenticated requests** are limited to **60 calls per hour**, while **authenticated requests** using a personal access token receive a **5,000-request quota**. Exceeding these limits results in HTTP 403 Forbidden responses, making proactive monitoring essential for production reliability.

## The Proactive Waiting Strategy in `_fetch_github_api`

The core rate-limiting logic resides in the private helper `_fetch_github_api` within [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 55-98). This function inspects response headers after every API call to determine whether the system can safely proceed or must wait for the quota to refresh.

### Reading Rate Limit Headers

After each HTTP request, the code extracts three critical headers from GitHub's response at lines 55-62:

- **`X-RateLimit-Limit`**: Total requests allowed in the current window (60 or 5000)
- **`X-RateLimit-Remaining`**: Requests remaining before the next reset
- **`X-RateLimit-Reset`**: UNIX timestamp indicating when the quota refreshes

These values are logged via `logger.info` to provide real-time visibility into API consumption.

### The Safety Threshold Logic

The implementation employs a two-tier threshold system to balance safety with performance:

1. **Critical threshold (10 calls)**: When `X-RateLimit-Remaining` drops below 10, the code triggers a **proactive sleep** to guarantee the next request executes after the reset window.
2. **Warning threshold (100 calls)**: When remaining calls fall between 10 and 100, the system logs the low quota without pausing execution, allowing developers to see warnings while continuing operation.

### Calculating Sleep Duration

When the critical threshold triggers at line 68, the code calculates wait time using the formula: `reset_timestamp - current_timestamp + 5_seconds_buffer`. This compensates for local clock drift. The duration is **capped at one hour** to prevent indefinite blocking, then executed via `time.sleep(wait_seconds)` before the response processing continues at line 96.

## Implementation Details in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)

The proactive waiting block occupies lines 68-96 of [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py). The workflow follows this precise sequence:

1. Execute the HTTP request with optional `Authorization: token ...` header for authentication
2. Parse rate limit headers from the response metadata
3. Compare remaining calls against the 10-request safety threshold
4. If below threshold, compute seconds until reset plus a 5-second buffer
5. Sleep for the calculated duration (maximum 3,600 seconds)
6. Process the response body, caching to local file when `DEVELOPMENT_MODE` is enabled in [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py)

## Practical Usage Examples

### Basic Profile Fetching

The public function `fetch_and_display_github_info` wraps the rate-limited internals:

```python
from github import fetch_and_display_github_info

# Fetches profile and repositories with automatic rate limiting

result = fetch_and_display_github_info("https://github.com/PavitKaur05")
print(result["profile"])
print(f"Found {result['total_projects']} repositories")

```

When the quota drops below 10 requests, the console displays:

```

⚠️  GitHub API rate limit low: 8/5000 requests remaining. Resets at 2026‑07‑08 14:12:30
⏳ Proactively sleeping for 212 seconds until rate limit resets...
✅ Rate limit should be reset now. Continuing...

```

### Authenticated Requests with Higher Quotas

Set the `GITHUB_TOKEN` environment variable to upgrade from 60 to 5,000 hourly requests:

```python
import os
os.environ["GITHUB_TOKEN"] = "ghp_XXXXXXXXXXXXXXXXXXXX"

# Now uses authenticated quota with same proactive waiting protection

result = fetch_and_display_github_info("https://github.com/PavitKaur05")

```

Authentication is passed through the `Authorization: token ...` header within `_fetch_github_api`.

## Configuration and Caching

The [`config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/config.py) file defines `DEVELOPMENT_MODE`, which enables local file caching of API responses. When active, successful requests persist to disk, allowing subsequent runs to bypass the API entirely and avoid consuming rate limit quota during development cycles.

## Summary

- The `_fetch_github_api` function in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 55-98) monitors `X-RateLimit-Remaining` headers on every request to track quota consumption.
- **Proactive waiting** triggers when fewer than 10 requests remain, sleeping until the `X-RateLimit-Reset` timestamp plus a 5-second buffer to account for clock skew.
- Sleep duration is **capped at one hour** to prevent excessive delays from blocking execution indefinitely.
- Warnings are logged when the quota drops below 100 calls, providing visibility without blocking execution.
- Authentication via `GITHUB_TOKEN` increases the quota from 60 to 5,000 requests per hour while maintaining the same protective logic.

## Frequently Asked Questions

### What happens when the rate limit is exactly at zero?

When `X-RateLimit-Remaining` reaches zero, the proactive waiting logic still activates because the value is below the 10-request threshold. The code calculates the time until the next reset window using the `X-RateLimit-Reset` timestamp and sleeps accordingly, ensuring the subsequent request succeeds with a fresh quota.

### How does proactive waiting differ from reactive error handling?

Reactive approaches catch HTTP 403 errors after they occur and implement exponential backoff. The hiring-agent's proactive strategy inspects headers before processing responses, preventing 403 errors entirely by pausing execution when the quota is nearly exhausted rather than recovering from failures.

### Why is there a 5-second buffer added to the reset time?

The 5-second buffer compensates for clock skew between the local system and GitHub's servers. This ensures that when the sleep completes, GitHub has actually reset the quota window, avoiding edge cases where a premature request would still trigger a rate limit violation due to timestamp mismatches.

### Can the sleep duration exceed one hour?

No. The implementation explicitly caps the wait time at one hour (3,600 seconds) within the calculation logic. If extreme clock discrepancies or calculation errors suggest a longer wait, the code limits the sleep to one hour and proceeds, allowing any remaining rate limit errors to surface naturally rather than blocking indefinitely.