Skip to content

Runbook: rate-limit exhaustion

OHIP limits calls per application key (≈ 50 req/s documented). StayFn’s governor keeps a token bucket per app key (default 40 tokens, 40/s; Ohip:RateLimit:*, per tenant/environment/hotel through the variables ohip.ratelimit.capacity / ohip.ratelimit.refillPerSecond) and reserves each function’s [OhipBudget] before it runs. An OHIP 429 honours Retry-After.

  • OHIP answering 429: every attempt honours its Retry-After; after the in-process attempts (3 for HTTP triggers) an HTTP invocation answers 502 upstream_error with upstreamStatus 429 and Retry-After, ends dead and leaves a dead letter; event and schedule invocations are re-queued through the outbox instead.
  • StayFn’s own governor refusing (the function’s [OhipBudget] is not available within its timeout): HTTP invocations answer 429 rate_limited with Retry-After before any OHIP call.
  • Metrics: stayfn.ohip.rate_limited rising, stayfn.ohip.ratelimit.waits / stayfn.ohip.ratelimit.wait_ms high, stayfn.ohip.ratelimit.tokens near 0 (Prometheus: stayfn_ohip_rate_limited_total, …).
  • The dashboard overview shows rate-limited calls and a budget with no available tokens or a blockedUntil.
  • Distinguish: 429 quota_exceeded is the tenant quota (quota.*), not OHIP; 503 overloaded is load shedding (docs/limits.md).
  1. GET /api/admin/overview?tenantId=<id>&minutes=15 → ohip.rateLimited, and ohip.budgets[] (capacity, refill, available tokens, blockedUntil when OHIP sent Retry-After).
  2. Was it OHIP or our own bucket? GET /api/admin/invocations/{id} → the OHIP call timeline marks rateLimited calls and their status (429 from OHIP) — our bucket alone produces waits without 429 calls.
  3. Who consumes the budget: GET /api/admin/usage?tenantId=<id>&groupBy=function (OHIP calls per function) and the invocation list filtered by function; a schedule or an event storm is the usual cause.
  4. Several tenants sharing one app key share one bucket (the limit is per app key).
  • A burst that will pass: nothing — events and schedules retry through the outbox; HTTP callers honour Retry-After.
  • Our bucket is set above what OHIP grants: lower ohip.ratelimit.capacity / refillPerSecond for the environment (admin variables) so we queue before OHIP answers 429; the governor rebuilds the bucket on the change.
  • A function spends more than it should: reduce its [OhipBudget]/page sizes, move bulk work to a schedule, or disable the schedule with the variable schedule.<name>.enabled = false without a deploy.
  • A noisy tenant: quota.invocationsPerDay / quota.concurrentInvocations (system-scope owner).
  • Sustained need: ask Oracle for a higher limit or a second application key and split tenants across keys.

WireMock answered the reservation search with 429 Retry-After: 2; the invocation failed with the rate-limited OHIP call recorded, the overview counted it, and after the mapping reset the next invocation succeeded.