How interviewstreet/hiring-agent Implements GitHub API Rate Limiting with Proactive Waiting

The hiring-agent repository prevents HTTP 403 errors by implementing a proactive waiting strategy in _fetch_github_api that monitors X-RateLimit-Remaining headers and automatically sleeps until the reset time when fewer than 10 requests remain.

The interviewstreet/hiring-agent project retrieves GitHub user profiles and repositories through the REST API, requiring robust handling of GitHub's strict hourly quotas. This article examines how the repository implements GitHub API rate limiting with proactive waiting to ensure uninterrupted data fetching without triggering rate limit violations.

Understanding GitHub API Rate Limits

GitHub enforces hourly request quotas that vary by authentication status. Unauthenticated requests are limited to 60 calls per hour, while authenticated requests using a personal access token receive a 5,000-request quota. Exceeding these limits results in HTTP 403 Forbidden responses, making proactive monitoring essential for production reliability.

The Proactive Waiting Strategy in _fetch_github_api

The core rate-limiting logic resides in the private helper _fetch_github_api within github.py (lines 55-98). This function inspects response headers after every API call to determine whether the system can safely proceed or must wait for the quota to refresh.

Reading Rate Limit Headers

After each HTTP request, the code extracts three critical headers from GitHub's response at lines 55-62:

  • X-RateLimit-Limit: Total requests allowed in the current window (60 or 5000)
  • X-RateLimit-Remaining: Requests remaining before the next reset
  • X-RateLimit-Reset: UNIX timestamp indicating when the quota refreshes

These values are logged via logger.info to provide real-time visibility into API consumption.

The Safety Threshold Logic

The implementation employs a two-tier threshold system to balance safety with performance:

  1. Critical threshold (10 calls): When X-RateLimit-Remaining drops below 10, the code triggers a proactive sleep to guarantee the next request executes after the reset window.
  2. Warning threshold (100 calls): When remaining calls fall between 10 and 100, the system logs the low quota without pausing execution, allowing developers to see warnings while continuing operation.

Calculating Sleep Duration

When the critical threshold triggers at line 68, the code calculates wait time using the formula: reset_timestamp - current_timestamp + 5_seconds_buffer. This compensates for local clock drift. The duration is capped at one hour to prevent indefinite blocking, then executed via time.sleep(wait_seconds) before the response processing continues at line 96.

Implementation Details in github.py

The proactive waiting block occupies lines 68-96 of github.py. The workflow follows this precise sequence:

  1. Execute the HTTP request with optional Authorization: token ... header for authentication
  2. Parse rate limit headers from the response metadata
  3. Compare remaining calls against the 10-request safety threshold
  4. If below threshold, compute seconds until reset plus a 5-second buffer
  5. Sleep for the calculated duration (maximum 3,600 seconds)
  6. Process the response body, caching to local file when DEVELOPMENT_MODE is enabled in config.py

Practical Usage Examples

Basic Profile Fetching

The public function fetch_and_display_github_info wraps the rate-limited internals:

from github import fetch_and_display_github_info

# Fetches profile and repositories with automatic rate limiting

result = fetch_and_display_github_info("https://github.com/PavitKaur05")
print(result["profile"])
print(f"Found {result['total_projects']} repositories")

When the quota drops below 10 requests, the console displays:


⚠️  GitHub API rate limit low: 8/5000 requests remaining. Resets at 2026‑07‑08 14:12:30
⏳ Proactively sleeping for 212 seconds until rate limit resets...
✅ Rate limit should be reset now. Continuing...

Authenticated Requests with Higher Quotas

Set the GITHUB_TOKEN environment variable to upgrade from 60 to 5,000 hourly requests:

import os
os.environ["GITHUB_TOKEN"] = "ghp_XXXXXXXXXXXXXXXXXXXX"

# Now uses authenticated quota with same proactive waiting protection

result = fetch_and_display_github_info("https://github.com/PavitKaur05")

Authentication is passed through the Authorization: token ... header within _fetch_github_api.

Configuration and Caching

The config.py file defines DEVELOPMENT_MODE, which enables local file caching of API responses. When active, successful requests persist to disk, allowing subsequent runs to bypass the API entirely and avoid consuming rate limit quota during development cycles.

Summary

  • The _fetch_github_api function in github.py (lines 55-98) monitors X-RateLimit-Remaining headers on every request to track quota consumption.
  • Proactive waiting triggers when fewer than 10 requests remain, sleeping until the X-RateLimit-Reset timestamp plus a 5-second buffer to account for clock skew.
  • Sleep duration is capped at one hour to prevent excessive delays from blocking execution indefinitely.
  • Warnings are logged when the quota drops below 100 calls, providing visibility without blocking execution.
  • Authentication via GITHUB_TOKEN increases the quota from 60 to 5,000 requests per hour while maintaining the same protective logic.

Frequently Asked Questions

What happens when the rate limit is exactly at zero?

When X-RateLimit-Remaining reaches zero, the proactive waiting logic still activates because the value is below the 10-request threshold. The code calculates the time until the next reset window using the X-RateLimit-Reset timestamp and sleeps accordingly, ensuring the subsequent request succeeds with a fresh quota.

How does proactive waiting differ from reactive error handling?

Reactive approaches catch HTTP 403 errors after they occur and implement exponential backoff. The hiring-agent's proactive strategy inspects headers before processing responses, preventing 403 errors entirely by pausing execution when the quota is nearly exhausted rather than recovering from failures.

Why is there a 5-second buffer added to the reset time?

The 5-second buffer compensates for clock skew between the local system and GitHub's servers. This ensures that when the sleep completes, GitHub has actually reset the quota window, avoiding edge cases where a premature request would still trigger a rate limit violation due to timestamp mismatches.

Can the sleep duration exceed one hour?

No. The implementation explicitly caps the wait time at one hour (3,600 seconds) within the calculation logic. If extreme clock discrepancies or calculation errors suggest a longer wait, the code limits the sleep to one hour and proceeds, allowing any remaining rate limit errors to surface naturally rather than blocking indefinitely.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →