Runbook: rate-limit exhaustion
OHIP limits calls per application key (≈ 50 req/s documented). StayFn’s governor keeps a token bucket per app key (default 40
tokens, 40/s; Ohip:RateLimit:*, per tenant/environment/hotel through the variables ohip.ratelimit.capacity / ohip.ratelimit.refillPerSecond)
and reserves each function’s [OhipBudget] before it runs. An OHIP 429 honours Retry-After.
Symptoms
Section titled “Symptoms”- OHIP answering 429: every attempt honours its
Retry-After; after the in-process attempts (3 for HTTP triggers) an HTTP invocation answers 502upstream_errorwithupstreamStatus429 andRetry-After, endsdeadand leaves a dead letter; event and schedule invocations are re-queued through the outbox instead. - StayFn’s own governor refusing (the function’s
[OhipBudget]is not available within its timeout): HTTP invocations answer 429rate_limitedwithRetry-Afterbefore any OHIP call. - Metrics:
stayfn.ohip.rate_limitedrising,stayfn.ohip.ratelimit.waits/stayfn.ohip.ratelimit.wait_mshigh,stayfn.ohip.ratelimit.tokensnear 0 (Prometheus:stayfn_ohip_rate_limited_total, …). - The dashboard overview shows rate-limited calls and a budget with no available tokens or a
blockedUntil. - Distinguish: 429
quota_exceededis the tenant quota (quota.*), not OHIP; 503overloadedis load shedding (docs/limits.md).
Diagnose
Section titled “Diagnose”GET /api/admin/overview?tenantId=<id>&minutes=15→ohip.rateLimited, andohip.budgets[](capacity, refill, available tokens,blockedUntilwhen OHIP sentRetry-After).- Was it OHIP or our own bucket?
GET /api/admin/invocations/{id}→ the OHIP call timeline marksrateLimitedcalls and their status (429 from OHIP) — our bucket alone produces waits without 429 calls. - Who consumes the budget:
GET /api/admin/usage?tenantId=<id>&groupBy=function(OHIP calls per function) and the invocation list filtered by function; a schedule or an event storm is the usual cause. - Several tenants sharing one app key share one bucket (the limit is per app key).
Remediate
Section titled “Remediate”- A burst that will pass: nothing — events and schedules retry through the outbox; HTTP callers honour
Retry-After. - Our bucket is set above what OHIP grants: lower
ohip.ratelimit.capacity/refillPerSecondfor the environment (admin variables) so we queue before OHIP answers 429; the governor rebuilds the bucket on the change. - A function spends more than it should: reduce its
[OhipBudget]/page sizes, move bulk work to a schedule, or disable the schedule with the variableschedule.<name>.enabled = falsewithout a deploy. - A noisy tenant:
quota.invocationsPerDay/quota.concurrentInvocations(system-scope owner). - Sustained need: ask Oracle for a higher limit or a second application key and split tenants across keys.
WireMock answered the reservation search with 429 Retry-After: 2; the invocation failed with the rate-limited OHIP call recorded, the overview
counted it, and after the mapping reset the next invocation succeeded.