Skip to content

Limits and load tests

What one StayFn host replica sustains, measured with k6 against the compose perf stack (tools/perf/, WireMock as OHIP), and the limits that matter more than StayFn’s own throughput. The numbers come from one developer machine that another build agent was using at the same time; treat them as orders of magnitude and re-measure on your hardware (.github/workflows/perf.yml, or tools/perf/README.md locally).

Item Value
Host Apple M4 (Mac16,10), 10 cores, 16 GB, macOS 26.5.2
Docker Colima VM (aarch64, 4 vCPU, 6 GiB), Docker 29.5.2, kernel 6.8 — host, PostgreSQL and WireMock share the 4 vCPU
Stack stayfn/host:0.8.0-local (Release, ASPNETCORE_ENVIRONMENT=Production), PostgreSQL 16-alpine (defaults, max_connections 100), WireMock 3.13.1; k6 2.3.0 on macOS
Background the Phase 7b agent was building and testing on the same machine (load average 3–5 before each run)
Date 2026-09-25

HTTP invoke — rsv/get-reservation, target 200 req/s

Section titled “HTTP invoke — rsv/get-reservation, target 200 req/s”

Each request: API-key lookup, quota check, invocation row, one OHIP search through the full handler chain (token cache, governor, retry, logging, call recording), mapping, row completion and usage accounting. tools/perf/invoke.js, constant arrival rate, 20 warm-up requests.

Final configuration (load shedding 64 in flight / 128 queued, framework Information logs raised to Warning, nofile 65536):

Run Offered Served Failed p50 p95 p99 Notes
fresh host, 60 s 200 req/s 198.3 req/s 0 % 3.6 ms 7.7 ms 134 ms thresholds met (p95 < 500 ms, errors < 1 %, 0 dropped)
warm, 60 s 200 req/s 200.0 req/s 0 % 3.6 ms 5.2 ms 14.4 ms
warm, 30 s 400 req/s 399.0 req/s 0 % 3.2 ms 7.3 ms 95.6 ms
fresh host, 30 s 800 req/s 779.5 req/s 0.6 % 42.8 ms 261 ms 297 ms the few failures are 503 overloaded while the fresh process warmed up
warm, 30 s 1,200 req/s 1,190 req/s 0 % 29.5 ms 95.9 ms 144 ms
warm, 30 s 2,000 req/s ≈ 1,050 req/s served, 47 % shed 47 % (all 503 overloaded) 158 ms 211 ms 263 ms host 138 % CPU, PostgreSQL 130 %: the VM’s 4 vCPU are the ceiling; the host stays up and answers probes

Knee on this machine: ≈ 1,000–1,200 req/s of a one-OHIP-call function per replica, bounded by CPU shared with PostgreSQL. Beyond it the replica sheds with 503 overloaded + Retry-After: 1 within a millisecond instead of queueing.

How we got there (runs before the fixes, same stack). The first runs used the Phase 7 behaviour: every request logged ~13 lines (EF Core “Executed DbCommand”, HttpClient request lines), nothing bounded concurrency, and the container had Docker’s default nofile of 1024.

Run Offered Served Failed p50 p95 Outcome
fresh host, 60 s 200 req/s 193.8 req/s 0 % 4.4 ms 2.15 s 365 iterations dropped, thresholds missed
warm, 60 s 200 req/s 199.5 req/s 0 % 3.6 ms 6.2 ms
warm, 30 s 400 req/s 398.9 req/s 0 % 4.2 ms 8.4 ms
30 s 800 req/s 624 req/s 75 % 442 ms 2.0 s process crashed (exit 139): 1,024 open descriptors, then DNS, assembly-load and “Out of memory” failures
30 s, limiter 256/512, nofile 65536 800 req/s ≈ 115 req/s successful 83 % 5.3 ms 6.9 s survived, but 256 requests in flight contended for the 100 database connections

Fixes: a global concurrency limiter with 503 overloaded (Security:MaxConcurrentRequests 64, Security:MaxQueuedRequests 128; /health, /metrics, /api/admin/stream and /mcp exempt), EF Core and HttpClient categories at Warning, and nofile 65536 in both compose files.

tools/perf/outbox.js: signed internal reservation.changed webhook events over 50 aggregates (partition keys), each with its own confirmation number. Per event: webhook ingress (HMAC, dedupe, inbox row) → outbox row → dispatcher (per-partition FIFO lease) → handler invocation → child invocation rsv/sync-reservation-to-crm-stub → one OHIP search. Drain = time after the last event until no inbox row is received and no outbox row is pending or processing.

Run Offered Accepted Ingress p95 Drain after the load Invocations OHIP calls Failures / dead letters
2 min 1,000 events/min 2,001 10.5 ms 1.2 s 4,002 succeeded 2,001 0 / 0
1 min 6,000 events/min 6,001 6.7 ms 1.2 s 12,002 succeeded 6,001 0 / 0

The dispatcher keeps up at six times the plan’s target with the default Events:Outbox settings (poll 1 s, batch 50, parallelism 4); the drain time is one poll interval. During the 1,000/min run the host used ≈ 30 % of one vCPU and PostgreSQL ≈ 13 %.

  • OHIP is the real ceiling. OHIP limits calls per application key (≈ 50 req/s documented); the governor defaults to 40 req/s per app key and queues (events, schedules) or answers 429 (HTTP) beyond it. One replica handles 20× that, so add replicas for availability, not for upstream throughput. The perf stack opens the governor to 5,000/s only because the stand-in has no limit.
  • Concurrency per replica: 64 requests in flight, 128 queued (Security:MaxConcurrentRequests, Security:MaxQueuedRequests); more is shed with 503 overloaded. Raise both only together with the database pool (Maximum Pool Size, default 100) and PostgreSQL max_connections (every replica holds up to its pool size plus two listener connections).
  • Request bodies: 1 MB globally (Kestrel:Limits:MaxRequestBodySize) and per endpoint (functions, webhook, admin).
  • File descriptors: run the container with nofile ≥ 65536 (both compose files set it; Kubernetes runtimes usually default higher). With 1,024 an overloaded replica failed hard.
  • Outbox: a leased row blocks later rows of its partition (per-aggregate FIFO); throughput across partitions scales with Events:Outbox:MaxParallelism × replicas. A future retry row holds its partition until due.
  • Streaming ingest: one consumer per (tenant, environment, chain) across all replicas (advisory lease); not load-tested.
  • Metering and retention run in the background with batched statements (docs/runbooks/backup-restore.md for retention windows).