How to configure a shared rate-limit store (Redis)
Problem: you run mcpgw as multiple replicas behind a load balancer. Each replica keeps its own in-memory token buckets, so a policy rule of “10 req/sec” or an input_rate_limit of “50 req/sec” is actually enforced at up to N × that — one bucket per replica.
Solution: point all replicas at a shared Redis. Buckets then live in Redis and every replica enforces one global limit. A single rate_limit_store block covers both input_rate_limit and policy.rules[].action: rate_limit — you do not configure two stores.
Recipe
1. Provision a Redis (managed — ElastiCache, Memorystore — or self-hosted). A single instance is fine for modest deployments (< ~1000 RPS gateway-wide); use Redis Cluster above that.
2. Add the rate_limit_store block to every replica’s mcpgw.yaml:
rate_limit_store:
type: redis
url: redis://:${REDIS_PASSWORD}@redis.internal:6379/0
key_prefix: "mcpgw:rl:"
operation_timeout: 50ms
fail_open: false
Use ${ENV_VAR} references for credentials — do not inline the password. Use a rediss:// URL (or tls.enabled: true with tls.ca_file) for TLS.
3. Restart the gateway (not SIGHUP). Switching the store type requires a restart so live buckets are not silently dropped. On boot the gateway logs:
INFO rate-limit store type=redis
4. Verify the global limit. With burst set low, hammer a single session across replicas and confirm the aggregate caps at tokens_per_second + burst, not N × it.
Failure behavior
If Redis becomes unreachable or a call exceeds operation_timeout, mcpgw degrades per fail_open:
fail_open: false(default) — deny the request (treat the bucket as empty), preserving the configured security control during an outage.fail_open: true— allow the request. Set this explicitly when availability matters more than rate enforcement. The gateway logsWARN rate-limit store unreachable; degrading per fail_open(rate-limited to once/minute) andINFO rate-limit store recoveredwhen Redis returns.
After any failure a 1-second window short-circuits further Redis calls to the fail-open/close result, so a sick Redis does not stack up timeouts on the request hot path.
No-Redis alternatives
If you do not want another stateful dependency:
- Sticky session affinity at the LB so a session always lands on the same replica. Works, but breaks on replica scale-down and constrains your LB.
- Accept divergence — set
tokens_per_secondwith enough headroom thatN ×is still acceptable.
Pitfalls
operation_timeoutis on the hot path. Every rate-limited request makes one Redis round-trip. Keep Redis close (same region/VPC) and the timeout tight (50ms default).- Share carefully. If other apps use the same Redis, keep a distinct
key_prefix. typechanges need a restart, notSIGHUP.url/TLS changes also currently require a restart.