# How to Configure Query Caching in INFINI Gateway to Accelerate Search Requests

> Accelerate search requests with INFINI Gateway query caching. Learn to configure get cache and set cache filter plugins for faster query results.

- Repository: [INFINI Labs/gateway](https://github.com/infinilabs/gateway)
- Tags: how-to-guide
- Published: 2026-03-04

---

**INFINI Gateway accelerates repeated search requests by caching full HTTP responses using the `get_cache` and `set_cache` filter plugins, which store and retrieve query results based on deterministic request fingerprints.**

INFINI Gateway (infinilabs/gateway) provides a high-performance caching layer that sits between clients and Elasticsearch clusters. By implementing query caching in INFINI Gateway, you eliminate redundant backend processing and reduce latency to sub-millisecond levels for frequently accessed data. The mechanism is implemented in [`proxy/filters/cache/request_cache.go`](https://github.com/infinilabs/gateway/blob/main/proxy/filters/cache/request_cache.go) and relies on coordinated filter plugins to intercept, store, and serve cached responses.

## Understanding the Query Caching Architecture

The caching mechanism relies on two specialized filter plugins. The `get_cache` filter intercepts incoming requests to check for cached entries, while `set_cache` stores responses after successful backend queries.

When a request arrives, the gateway computes a deterministic fingerprint by hashing the HTTP method, URL, query-string arguments, and selected headers (lines 355-369 in [`request_cache.go`](https://github.com/infinilabs/gateway/blob/main/request_cache.go)). This hash serves as the cache key, ensuring that identical requests yield identical keys regardless of which gateway instance processes them.

### Cache Lookup and Storage Flow

A typical implementation uses a **cache-first** flow:

- `get_cache` checks for existing entries using the computed fingerprint
- If found, the cached response (headers, body, and status) returns immediately
- If missed, the request proceeds to the `elasticsearch` filter
- `set_cache` captures the successful response and stores it when `ctx.Set(common.CACHEABLE, true)` is present

## Supported Cache Backends

The `cache_type` parameter (default: `ristretto`) determines the storage backend, as implemented in the constructor logic (lines 88-104 of [`request_cache.go`](https://github.com/infinilabs/gateway/blob/main/request_cache.go)):

- **Ristretto** – An in-process, high-performance LRU cache optimized for concurrent access. Configure memory limits using `max_cached_size` (default 1 GB).
- **Ccache** – A layered in-memory cache supporting item count limits via `max_cached_item` (default 1,000,000).
- **Redis** – An external Redis instance for distributed caching across multiple gateway nodes, instantiated via `redis.NewClient`.

## Configuring Cache Parameters

All default values are defined in [`proxy/filters/cache/request_cache.go`](https://github.com/infinilabs/gateway/blob/main/proxy/filters/cache/request_cache.go) (lines 71-81). The following parameters control cache behavior:

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `cache_type` | string | `ristretto` | Backend selection (`ristretto`, `ccache`, `redis`) |
| `cache_ttl` | duration | `10s` | Time-to-live for standard query results |
| `async_search_cache_ttl` | duration | `10m` | Extended TTL for asynchronous search results |
| `max_cached_item` | int | `1000000` | Maximum item count (ccache only) |
| `max_cached_size` | int64 | `1GB` | Memory limit in bytes (ristretto only) |
| `pass_patterns` | []string | `["_bulk","_cat","scroll",...]` | URL patterns excluded from caching |
| `validated_status` | []int | `[200,201,404,403,413,400,301]` | HTTP status codes eligible for caching |
| `min_response_size` | int | `-1` | Minimum body size to cache |
| `max_response_size` | int | `int(^uint(0)>>1)` | Maximum body size to cache |

TTL values are parsed using `util.GetDurationOrDefault`, supporting standard duration strings like `30s`, `5m`, or `1h`.

## Implementing a Cache-First Flow

Configure the `cache_first` flow in your [`gateway.yml`](https://github.com/infinilabs/gateway/blob/main/gateway.yml) to enable end-to-end query caching:

```yaml
flow:
  - name: cache_first
    filter:
      - get_cache:
          pass_patterns: ["_cat","scroll","_refresh"]
      - elasticsearch:
          elasticsearch: prod
          max_connection_per_node: 1000
      - set_cache:
          cache_ttl: 30s
          max_cache_items: 200000

```

In this configuration, `get_cache` attempts to serve responses from memory before forwarding requests to Elasticsearch. Successful responses are then stored by `set_cache` with a 30-second TTL. The `pass_patterns` list ensures that administrative endpoints like `_cat` and `scroll` operations never cache results.

## Bypassing and Fine-Tuning the Cache

For ad-hoc queries requiring fresh data, append `no_cache=true` to the query string:

```bash
curl "http://localhost:8000/_search?q=hostname:myhost&no_cache=true"

```

This parameter instructs both `get_cache` and `set_cache` to skip processing for that specific request, as documented in [`docs/content.en/docs/references/filters/cache.md`](https://github.com/infinilabs/gateway/blob/main/docs/content.en/docs/references/filters/cache.md).

To optimize cache efficiency, tune `min_response_size` and `max_response_size` to avoid caching very small or extremely large payloads that provide diminishing returns. The `validated_status` array ensures that error responses (like 404s) can be cached to prevent unnecessary backend load from invalid queries.

## Programmatic Configuration

For dynamic setups, instantiate filters directly in Go using the constructor pattern found in [`request_cache.go`](https://github.com/infinilabs/gateway/blob/main/request_cache.go) (lines 88-104):

```go
cfg := cache.Config{
    CacheType:   "redis",
    CacheTTL:    "15s",
    MaxCachedSize: 2<<30, // 2 GB
    PassPatterns: []string{"_cat","_bulk"},
}
filter, err := cache.NewGet(&config.Config{Raw: cfg})
if err != nil {
    log.Fatal(err)
}

```

This approach allows runtime cache backend selection and is particularly useful when deploying INFINI Gateway as a library within larger applications.

## Summary

- **INFINI Gateway** implements query caching through `get_cache` and `set_cache` filter plugins that intercept HTTP requests to Elasticsearch.
- **Request fingerprinting** (lines 355-369 in [`request_cache.go`](https://github.com/infinilabs/gateway/blob/main/request_cache.go)) creates cache keys from method, URL, and headers, ensuring deterministic lookups across distributed instances.
- **Three backends** are supported: Ristretto (default in-process LRU), Ccache (layered memory), and Redis (distributed external storage).
- **Configuration parameters** including `cache_ttl`, `pass_patterns`, and `validated_status` allow precise control over what gets cached and for how long.
- **Cache bypass** via `no_cache=true` query parameter ensures fresh data when needed without modifying configuration.

## Frequently Asked Questions

### How does request fingerprinting work in INFINI Gateway?

The gateway generates cache keys by computing a hash of the HTTP method, full URL, query-string arguments, and selected headers. This implementation in [`proxy/filters/cache/request_cache.go`](https://github.com/infinilabs/gateway/blob/main/proxy/filters/cache/request_cache.go) (lines 355-369) ensures that identical requests produce identical fingerprints regardless of which gateway instance processes them, enabling consistent cache hits across distributed deployments.

### What is the difference between `get_cache` and `set_cache` filters?

`get_cache` acts as an interceptor that checks for existing cached responses before forwarding requests to Elasticsearch. If found, it immediately returns the stored HTTP response (body, headers, and status). `set_cache` operates after the `elasticsearch` filter, storing successful responses only when the request context indicates cacheability via `ctx.Set(common.CACHEABLE, true)`.

### How can I invalidate or bypass the query cache for specific requests?

Append `no_cache=true` to any request's query string to force both filters to skip cache lookup and storage for that specific call. Alternatively, modify `pass_patterns` in your configuration to exclude specific URL patterns from caching entirely, or adjust the `cache_ttl` duration to control how long entries persist before automatic expiration.

### Which cache backend should I choose for production deployments?

Choose **Ristretto** for single-instance deployments requiring ultra-low latency and sub-millisecond access times. Select **Redis** when running multiple gateway instances that require shared cache state across nodes. Use **Ccache** only if you need specific layered caching semantics with explicit item counting rather than memory-based limits.