How to Configure Traffic Throttling at Index, Node, and User Levels in INFINI Gateway
Configure traffic throttling in INFINI Gateway by enabling traffic_control for node-level limits, or by adding request_path_limiter and request_user_limiter filters to your flows for index-level and user-level rate limiting.
INFINI Gateway provides a flexible, multi-layered traffic throttling configuration that allows you to control request rates at different granularities. Whether you need to protect individual Elasticsearch nodes from overload, enforce different quality-of-service tiers per index, or apply usage caps per user account, the gateway implements these limits through a shared GenericLimiter framework located in proxy/filters/throttle/request_limiter_base.go.
Understanding the Rate-Limiting Framework
All throttling mechanisms in INFINI Gateway rely on the GenericLimiter implementation in proxy/filters/throttle/request_limiter_base.go. This component provides consistent rate-limiting logic across node, index, and user levels with the following configurable parameters:
max_requests/burst_requests– Control request-count limits per time windowmax_bytes/burst_bytes– Enforce bandwidth limitsaction– Choose betweenretry(default) ordropbehavior when limits are exceededmax_retry_times,retry_delay_in_ms,status,message– Fine-tune retry behavior and client responses
Node-Level Traffic Throttling
Node-level traffic throttling controls the overall request volume, bandwidth, and connection count for each individual Elasticsearch node in your cluster. This is the coarsest granularity and serves as your first line of defense against cluster overload.
Configuration Parameters
Enable node-level throttling by adding the traffic_control section to your Elasticsearch connection configuration in gateway.yml:
es:
endpoints:
- "http://es-node1:9200"
- "http://es-node2:9200"
discovery:
enabled: true
refresh:
enabled: true
interval: "30s"
traffic_control:
enabled: true # Enable throttling
max_qps_per_node: 5000 # Max requests per second per node
max_bytes_per_node: 10485760 # 10 MiB/s per node
max_connection_per_node: 200 # Max concurrent TCP connections
max_wait_time_in_ms: 10000 # Max wait time before rejection
Runtime Enforcement
The gateway enforces node-level limits at runtime through the metadata.CheckNodeTrafficThrottle function, which is invoked in two critical locations:
proxy/output/elastic/reverseproxy.go(line 558) – Applies throttling to normal HTTP proxy trafficpipeline/queue_consumer/diskqueue_consumer.go(line 339) – Enforces limits on queued bulk requests
These calls forward the node host identifier to the GenericLimiter, ensuring that traffic to each backend node stays within configured bounds.
Index-Level Traffic Throttling
Index-level throttling allows you to enforce different rate limits for specific Elasticsearch indices, enabling quality-of-service tiers where critical indices receive dedicated bandwidth while others share a common pool.
Configuration Example
Create a flow that includes the request_path_limiter filter to implement index-level throttling:
flow:
- name: index_rate_limit
filter:
- request_path_limiter:
message: "Rate limit exceeded for this index"
rules:
- pattern: "/(?P<index_name>medcl)/_search"
max_qps: 3
group: index_name
- pattern: "/(?P<index_name>.*?)/_search"
max_qps: 100
group: index_name
The filter implementation resides in proxy/filters/throttle/request_path_limiter.go.
How It Works
The request_path_limiter filter processes requests through the following steps:
- Path Matching – The request URL is matched against each rule in order using regular expressions
- Token Extraction – When a rule matches, the captured group (specified by
group, such asindex_name) becomes the token for the generic limiter - Rate Enforcement – The GenericLimiter checks the QPS (
max_qps) for that specific index token - Action Execution – If the limit is exceeded, the filter returns HTTP 429 with the configured
message, or retries according to theactionfield
User-Level Traffic Throttling
User-level throttling enables you to enforce request quotas based on HTTP Basic Authentication credentials, allowing you to differentiate between premium users, standard accounts, and potential abusers.
Configuration Example
Add the request_user_limiter filter to your flow to implement per-user rate limiting:
flow:
- name: user_rate_limit
filter:
- request_user_limiter:
user:
- elastic
- medcl
max_requests: 256
max_bytes: 102400 # 100 KiB/s per user
action: retry # or "drop"
message: "you reached our limit"
The filter implementation is located in proxy/filters/throttle/request_user_limiter.go.
Processing Steps
The request_user_limiter filter executes the following logic:
- Credential Extraction – Parses the HTTP Basic-Auth header from the request using
ctx.Request.ParseBasicAuth() - User Matching – Checks if the extracted username exists in the configured
userlist - Token Generation – If matched, creates a limiter token with type
userand the username as the token value - Limit Enforcement – Invokes the GenericLimiter to enforce
max_requestsandmax_byteslimits - Response Handling – Returns the configured
messagewith appropriate HTTP status when limits are exceeded, or applies retry logic based on theactionsetting
Additional Granularities
Beyond the three primary levels, INFINI Gateway provides several specialized throttling filters that use the same GenericLimiter infrastructure:
request_host_limiter– Throttles by HTTPHostheader, useful for multi-tenant deployments with virtual hostsrequest_client_ip_limiter– Limits requests per client IP address to prevent abuse from specific sourcesrequest_api_key_limiter– Enforces quotas based on API key headers for third-party integrationsbulk_request_throttle– Controls bulk write throughput per index by parsing bulk request payloads
All filters reside in proxy/filters/throttle/ and share the configuration schema defined in request_limiter_base.go.
Summary
-
Node-level throttling is configured via the
traffic_controlsection in your Elasticsearch connection settings and enforced bymetadata.CheckNodeTrafficThrottleinproxy/output/elastic/reverseproxy.goandpipeline/queue_consumer/diskqueue_consumer.go. -
Index-level throttling uses the
request_path_limiterfilter to extract index names from URL paths via regular expressions and apply per-index QPS and bandwidth limits. -
User-level throttling leverages the
request_user_limiterfilter to parse HTTP Basic-Auth credentials and enforce request quotas per username. -
All granularities rely on the GenericLimiter in
proxy/filters/throttle/request_limiter_base.go, providing consistent behavior for request counts, byte limits, retry logic, and error responses.
Frequently Asked Questions
How do I combine multiple throttling levels in a single flow?
You can chain multiple limiter filters in sequence within your flow definition. For example, place request_user_limiter first to enforce per-user caps, followed by request_path_limiter to apply index-specific limits. Each filter operates independently using the shared GenericLimiter, so a request must pass all configured limits to proceed.
What happens when a throttling limit is exceeded?
By default, the GenericLimiter returns HTTP 429 (Too Many Requests) with your configured message. If you set action: retry, the gateway will automatically retry the request up to max_retry_times with retry_delay_in_ms intervals before finally rejecting it. For queued bulk operations in diskqueue_consumer.go, the retry behavior ensures data durability while respecting node-level bandwidth constraints.
Can I throttle by API key instead of username?
Yes, use the request_api_key_limiter filter instead of request_user_limiter. This filter extracts tokens from a configurable header (typically Authorization or X-API-Key) and applies the same GenericLimiter logic. This is ideal for third-party API access management where Basic Auth is not used.
Where is the node-level throttling enforced in the codebase?
Node-level throttling is enforced in two critical locations: proxy/output/elastic/reverseproxy.go at line 558 for standard HTTP proxy traffic, and pipeline/queue_consumer/diskqueue_consumer.go at line 339 for queued bulk ingestion. Both locations invoke metadata.CheckNodeTrafficThrottle, which delegates to the GenericLimiter to check against the traffic_control configuration parameters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →