Copilot SDK API Rate Limits: A Complete Guide to User Quotas, Provider Constraints, and Error Handling
The Copilot SDK enforces strict request quotas at both the GitHub user level and third-party provider level, returning structured rate_limit errors with specific codes like user_weekly_rate_limited and optional secondsUntilReset timers to facilitate intelligent retry logic.
The github/copilot-sdk implements a multi-layered rate limiting architecture designed to prevent abuse while maintaining service stability across Node.js, Python, Go, Java, Rust, and .NET bindings. When applications exceed these limits—whether through individual user quotas or upstream provider constraints—the SDK returns standardized error payloads defined in nodejs/src/generated/session-events.ts. Understanding these Copilot SDK API rate limits is essential for building resilient integrations that handle backoff periods gracefully and leverage automatic failover mechanisms.
Types of Rate Limits in the Copilot SDK
The SDK manages multiple concurrent limitation tiers that affect request throughput.
Per-User Copilot Limits
GitHub imposes specific quotas on individual user accounts that restrict weekly and global usage. When a user exceeds these thresholds, the API returns a rate_limit error containing codes such as user_weekly_rate_limited or user_global_rate_limited. These limits are documented in docs/setup/github-oauth.md at line 460, which defines the user-level rate limit table governing access across the entire Copilot ecosystem.
Third-Provider Rate Limits
The SDK acts as a proxy to underlying AI providers such as OpenAI, which enforce their own independent quotas. When these upstream providers reject a request due to quota exhaustion, the SDK surfaces the restriction through the same rate_limit error type while preserving provider-specific context. This architecture is implemented in nodejs/src/generated/session-events.ts between lines 1042 and 1045, ensuring consistent error handling regardless of which provider ultimately services the request.
Rate Limit Error Structure and Response Fields
Every rate limit response follows a standardized schema defined in the generated SessionEvents type, part of the public API surface for all supported languages.
When triggered, responses include:
errorType: Always set to"rate_limit"to identify the categoryerrorCode: Enumerates the specific cause, such as"user_weekly_rate_limited","user_global_rate_limited", or"integration_rate_limited"secondsUntilReset: Optional integer indicating remaining time before quota refreshes
The TypeScript definition in nodejs/src/generated/session-events.ts specifies these fields:
/** The rate limit error code that triggered this request */
errorCode?: string;
/** Seconds until the rate limit resets, when known. */
secondsUntilReset?: number;
Automatic Recovery and Mitigation Strategies
The Copilot SDK provides built-in mechanisms to handle quota exhaustion gracefully without requiring manual intervention.
Auto-Mode Switch Behavior
When the runtime detects a rate limit error, it can automatically transition the session to auto mode and emit an auto_mode_switch.requested event. This mechanism, documented in nodejs/src/generated/session-events.ts (lines 1037–1041), captures the original error code and reset timer to inform downstream UI components. Developers can enable the continueOnAutoMode flag to suppress duplicate error notifications while the SDK handles the transition silently.
Client-Side Caching to Optimize Usage
To prevent unnecessary API calls that count against rate limits, the SDK implements aggressive result caching for expensive operations such as model listings. According to nodejs/src/client.ts at lines 1896–1898, successful responses are cached after the first request, with subsequent calls returning cached data rather than issuing new network requests. The Python implementation in python/copilot/client.py and Rust handler in rust/src/handler.rs confirm this behavior across language bindings.
Handling Rate Limits in Production Code
When building production applications, implement exponential backoff strategies that respect the secondsUntilReset field when provided. If the reset timer is absent, fall back to standard exponential backoff starting at 1 second.
Example error handling pattern in TypeScript:
try {
const response = await copilotClient.generateCompletions(params);
} catch (error) {
if (error.errorType === "rate_limit") {
const delayMs = (error.secondsUntilReset || 60) * 1000;
await sleep(delayMs);
// Implement retry logic here
}
}
For long-running sessions, configure the continueOnAutoMode option to allow seamless transitions during quota exhaustion without interrupting the user experience.
Summary
- The Copilot SDK API enforces user-specific limits (weekly and global) and provider-specific limits (OpenAI quotas) that return standardized
rate_limiterrors. - Error responses include precise codes such as
user_weekly_rate_limitedand optionalsecondsUntilResettimers defined innodejs/src/generated/session-events.ts. - The SDK supports automatic mode switching via
auto_mode_switch.requestedevents when limits are hit, configurable through thecontinueOnAutoModeflag. - Built-in caching in
nodejs/src/client.tsandpython/copilot/client.pyprevents redundant requests by storing successful responses for model listings and context retrieval. - Always implement backoff logic that respects reset timers to avoid hard failures and ensure compliance with GitHub's usage policies.
Frequently Asked Questions
What error code does the Copilot SDK return when a user exceeds their weekly limit?
When a user exceeds their weekly Copilot quota, the SDK returns a rate_limit error with the specific error code user_weekly_rate_limited. This code is defined in the SessionEvents type schema at nodejs/src/generated/session-events.ts and is surfaced consistently across all language bindings including Python, Go, Java, and Rust.
Does the Copilot SDK provide information about when a rate limit will reset?
Yes, when the upstream service provides reset timing information, the SDK includes a secondsUntilReset field in the error response. This integer value indicates the number of seconds remaining until the quota refreshes, allowing client applications to display countdown timers or schedule precise retry attempts. This behavior is documented in both nodejs/src/generated/session-events.ts and the Rust handler at rust/src/handler.rs.
Can the Copilot SDK automatically handle rate limit errors without manual intervention?
Yes, the SDK supports automatic session management through the auto_mode_switch.requested event, which triggers when rate limits are encountered. By enabling the continueOnAutoMode configuration flag, applications can allow the SDK to silently transition to auto mode without surfacing duplicate error UIs, maintaining session continuity while respecting provider quotas.
How does the Copilot SDK prevent applications from unnecessarily hitting rate limits?
The SDK implements result caching mechanisms that store successful responses for expensive operations like model listings. As implemented in nodejs/src/client.ts (lines 1896–1898) and confirmed in python/copilot/client.py, the client caches results after the first successful call, ensuring subsequent requests return cached data rather than consuming additional quota through redundant API calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →