# Copilot SDK API Rate Limits: A Complete Guide to User Quotas, Provider Constraints, and Error Handling

> Master Copilot SDK API rate limits. Understand user quotas, provider constraints, and implement error handling for `rate_limit` errors using `secondsUntilReset` for smart retries.

- Repository: [GitHub/copilot-sdk](https://github.com/github/copilot-sdk)
- Tags: api-reference
- Published: 2026-07-20

---

**The Copilot SDK enforces strict request quotas at both the GitHub user level and third-party provider level, returning structured `rate_limit` errors with specific codes like `user_weekly_rate_limited` and optional `secondsUntilReset` timers to facilitate intelligent retry logic.**

The `github/copilot-sdk` implements a multi-layered rate limiting architecture designed to prevent abuse while maintaining service stability across Node.js, Python, Go, Java, Rust, and .NET bindings. When applications exceed these limits—whether through individual user quotas or upstream provider constraints—the SDK returns standardized error payloads defined in [`nodejs/src/generated/session-events.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/generated/session-events.ts). Understanding these Copilot SDK API rate limits is essential for building resilient integrations that handle backoff periods gracefully and leverage automatic failover mechanisms.

## Types of Rate Limits in the Copilot SDK

The SDK manages multiple concurrent limitation tiers that affect request throughput.

### Per-User Copilot Limits

GitHub imposes specific quotas on individual user accounts that restrict weekly and global usage. When a user exceeds these thresholds, the API returns a `rate_limit` error containing codes such as `user_weekly_rate_limited` or `user_global_rate_limited`. These limits are documented in [`docs/setup/github-oauth.md`](https://github.com/github/copilot-sdk/blob/main/docs/setup/github-oauth.md) at line 460, which defines the user-level rate limit table governing access across the entire Copilot ecosystem.

### Third-Provider Rate Limits

The SDK acts as a proxy to underlying AI providers such as OpenAI, which enforce their own independent quotas. When these upstream providers reject a request due to quota exhaustion, the SDK surfaces the restriction through the same `rate_limit` error type while preserving provider-specific context. This architecture is implemented in [`nodejs/src/generated/session-events.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/generated/session-events.ts) between lines 1042 and 1045, ensuring consistent error handling regardless of which provider ultimately services the request.

## Rate Limit Error Structure and Response Fields

Every rate limit response follows a standardized schema defined in the generated `SessionEvents` type, part of the public API surface for all supported languages.

When triggered, responses include:

- **`errorType`**: Always set to `"rate_limit"` to identify the category
- **`errorCode`**: Enumerates the specific cause, such as `"user_weekly_rate_limited"`, `"user_global_rate_limited"`, or `"integration_rate_limited"`
- **`secondsUntilReset`**: Optional integer indicating remaining time before quota refreshes

The TypeScript definition in [`nodejs/src/generated/session-events.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/generated/session-events.ts) specifies these fields:

```typescript
/** The rate limit error code that triggered this request */
errorCode?: string;
/** Seconds until the rate limit resets, when known. */
secondsUntilReset?: number;

```

## Automatic Recovery and Mitigation Strategies

The Copilot SDK provides built-in mechanisms to handle quota exhaustion gracefully without requiring manual intervention.

### Auto-Mode Switch Behavior

When the runtime detects a rate limit error, it can automatically transition the session to auto mode and emit an `auto_mode_switch.requested` event. This mechanism, documented in [`nodejs/src/generated/session-events.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/generated/session-events.ts) (lines 1037–1041), captures the original error code and reset timer to inform downstream UI components. Developers can enable the `continueOnAutoMode` flag to suppress duplicate error notifications while the SDK handles the transition silently.

### Client-Side Caching to Optimize Usage

To prevent unnecessary API calls that count against rate limits, the SDK implements aggressive result caching for expensive operations such as model listings. According to [`nodejs/src/client.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/client.ts) at lines 1896–1898, successful responses are cached after the first request, with subsequent calls returning cached data rather than issuing new network requests. The Python implementation in [`python/copilot/client.py`](https://github.com/github/copilot-sdk/blob/main/python/copilot/client.py) and Rust handler in [`rust/src/handler.rs`](https://github.com/github/copilot-sdk/blob/main/rust/src/handler.rs) confirm this behavior across language bindings.

## Handling Rate Limits in Production Code

When building production applications, implement exponential backoff strategies that respect the `secondsUntilReset` field when provided. If the reset timer is absent, fall back to standard exponential backoff starting at 1 second.

Example error handling pattern in TypeScript:

```typescript
try {
  const response = await copilotClient.generateCompletions(params);
} catch (error) {
  if (error.errorType === "rate_limit") {
    const delayMs = (error.secondsUntilReset || 60) * 1000;
    await sleep(delayMs);
    // Implement retry logic here
  }
}

```

For long-running sessions, configure the `continueOnAutoMode` option to allow seamless transitions during quota exhaustion without interrupting the user experience.

## Summary

- The Copilot SDK API enforces **user-specific limits** (weekly and global) and **provider-specific limits** (OpenAI quotas) that return standardized `rate_limit` errors.
- Error responses include precise codes such as `user_weekly_rate_limited` and optional `secondsUntilReset` timers defined in [`nodejs/src/generated/session-events.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/generated/session-events.ts).
- The SDK supports **automatic mode switching** via `auto_mode_switch.requested` events when limits are hit, configurable through the `continueOnAutoMode` flag.
- **Built-in caching** in [`nodejs/src/client.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/client.ts) and [`python/copilot/client.py`](https://github.com/github/copilot-sdk/blob/main/python/copilot/client.py) prevents redundant requests by storing successful responses for model listings and context retrieval.
- Always implement backoff logic that respects reset timers to avoid hard failures and ensure compliance with GitHub's usage policies.

## Frequently Asked Questions

### What error code does the Copilot SDK return when a user exceeds their weekly limit?

When a user exceeds their weekly Copilot quota, the SDK returns a `rate_limit` error with the specific error code `user_weekly_rate_limited`. This code is defined in the `SessionEvents` type schema at [`nodejs/src/generated/session-events.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/generated/session-events.ts) and is surfaced consistently across all language bindings including Python, Go, Java, and Rust.

### Does the Copilot SDK provide information about when a rate limit will reset?

Yes, when the upstream service provides reset timing information, the SDK includes a `secondsUntilReset` field in the error response. This integer value indicates the number of seconds remaining until the quota refreshes, allowing client applications to display countdown timers or schedule precise retry attempts. This behavior is documented in both [`nodejs/src/generated/session-events.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/generated/session-events.ts) and the Rust handler at [`rust/src/handler.rs`](https://github.com/github/copilot-sdk/blob/main/rust/src/handler.rs).

### Can the Copilot SDK automatically handle rate limit errors without manual intervention?

Yes, the SDK supports automatic session management through the `auto_mode_switch.requested` event, which triggers when rate limits are encountered. By enabling the `continueOnAutoMode` configuration flag, applications can allow the SDK to silently transition to auto mode without surfacing duplicate error UIs, maintaining session continuity while respecting provider quotas.

### How does the Copilot SDK prevent applications from unnecessarily hitting rate limits?

The SDK implements result caching mechanisms that store successful responses for expensive operations like model listings. As implemented in [`nodejs/src/client.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/client.ts) (lines 1896–1898) and confirmed in [`python/copilot/client.py`](https://github.com/github/copilot-sdk/blob/main/python/copilot/client.py), the client caches results after the first successful call, ensuring subsequent requests return cached data rather than consuming additional quota through redundant API calls.