Limits and load tests
What one StayFn host replica sustains, measured with k6 against the compose perf stack (tools/perf/, WireMock as OHIP), and the limits that
matter more than StayFn’s own throughput. The numbers come from one developer machine that another build agent was using at the same time;
treat them as orders of magnitude and re-measure on your hardware (.github/workflows/perf.yml, or tools/perf/README.md locally).
Machine
Section titled “Machine”| Item | Value |
|---|---|
| Host | Apple M4 (Mac16,10), 10 cores, 16 GB, macOS 26.5.2 |
| Docker | Colima VM (aarch64, 4 vCPU, 6 GiB), Docker 29.5.2, kernel 6.8 — host, PostgreSQL and WireMock share the 4 vCPU |
| Stack | stayfn/host:0.8.0-local (Release, ASPNETCORE_ENVIRONMENT=Production), PostgreSQL 16-alpine (defaults, max_connections 100), WireMock 3.13.1; k6 2.3.0 on macOS |
| Background | the Phase 7b agent was building and testing on the same machine (load average 3–5 before each run) |
| Date | 2026-09-25 |
Measured
Section titled “Measured”HTTP invoke — rsv/get-reservation, target 200 req/s
Section titled “HTTP invoke — rsv/get-reservation, target 200 req/s”Each request: API-key lookup, quota check, invocation row, one OHIP search through the full handler chain (token cache, governor, retry,
logging, call recording), mapping, row completion and usage accounting. tools/perf/invoke.js, constant arrival rate, 20 warm-up requests.
Final configuration (load shedding 64 in flight / 128 queued, framework Information logs raised to Warning, nofile 65536):
| Run | Offered | Served | Failed | p50 | p95 | p99 | Notes |
|---|---|---|---|---|---|---|---|
| fresh host, 60 s | 200 req/s | 198.3 req/s | 0 % | 3.6 ms | 7.7 ms | 134 ms | thresholds met (p95 < 500 ms, errors < 1 %, 0 dropped) |
| warm, 60 s | 200 req/s | 200.0 req/s | 0 % | 3.6 ms | 5.2 ms | 14.4 ms | |
| warm, 30 s | 400 req/s | 399.0 req/s | 0 % | 3.2 ms | 7.3 ms | 95.6 ms | |
| fresh host, 30 s | 800 req/s | 779.5 req/s | 0.6 % | 42.8 ms | 261 ms | 297 ms | the few failures are 503 overloaded while the fresh process warmed up |
| warm, 30 s | 1,200 req/s | 1,190 req/s | 0 % | 29.5 ms | 95.9 ms | 144 ms | |
| warm, 30 s | 2,000 req/s | ≈ 1,050 req/s served, 47 % shed | 47 % (all 503 overloaded) |
158 ms | 211 ms | 263 ms | host 138 % CPU, PostgreSQL 130 %: the VM’s 4 vCPU are the ceiling; the host stays up and answers probes |
Knee on this machine: ≈ 1,000–1,200 req/s of a one-OHIP-call function per replica, bounded by CPU shared with PostgreSQL. Beyond it the
replica sheds with 503 overloaded + Retry-After: 1 within a millisecond instead of queueing.
How we got there (runs before the fixes, same stack). The first runs used the Phase 7 behaviour: every request logged ~13 lines (EF Core
“Executed DbCommand”, HttpClient request lines), nothing bounded concurrency, and the container had Docker’s default nofile of 1024.
| Run | Offered | Served | Failed | p50 | p95 | Outcome |
|---|---|---|---|---|---|---|
| fresh host, 60 s | 200 req/s | 193.8 req/s | 0 % | 4.4 ms | 2.15 s | 365 iterations dropped, thresholds missed |
| warm, 60 s | 200 req/s | 199.5 req/s | 0 % | 3.6 ms | 6.2 ms | |
| warm, 30 s | 400 req/s | 398.9 req/s | 0 % | 4.2 ms | 8.4 ms | |
| 30 s | 800 req/s | 624 req/s | 75 % | 442 ms | 2.0 s | process crashed (exit 139): 1,024 open descriptors, then DNS, assembly-load and “Out of memory” failures |
30 s, limiter 256/512, nofile 65536 |
800 req/s | ≈ 115 req/s successful | 83 % | 5.3 ms | 6.9 s | survived, but 256 requests in flight contended for the 100 database connections |
Fixes: a global concurrency limiter with 503 overloaded (Security:MaxConcurrentRequests 64,
Security:MaxQueuedRequests 128; /health, /metrics, /api/admin/stream and /mcp exempt), EF Core and HttpClient categories at Warning,
and nofile 65536 in both compose files.
Outbox — 1,000 events/min
Section titled “Outbox — 1,000 events/min”tools/perf/outbox.js: signed internal reservation.changed webhook events over 50 aggregates (partition keys), each with its own confirmation
number. Per event: webhook ingress (HMAC, dedupe, inbox row) → outbox row → dispatcher (per-partition FIFO lease) → handler invocation →
child invocation rsv/sync-reservation-to-crm-stub → one OHIP search. Drain = time after the last event until no inbox row is received
and no outbox row is pending or processing.
| Run | Offered | Accepted | Ingress p95 | Drain after the load | Invocations | OHIP calls | Failures / dead letters |
|---|---|---|---|---|---|---|---|
| 2 min | 1,000 events/min | 2,001 | 10.5 ms | 1.2 s | 4,002 succeeded | 2,001 | 0 / 0 |
| 1 min | 6,000 events/min | 6,001 | 6.7 ms | 1.2 s | 12,002 succeeded | 6,001 | 0 / 0 |
The dispatcher keeps up at six times the plan’s target with the default Events:Outbox settings (poll 1 s, batch 50, parallelism 4); the
drain time is one poll interval. During the 1,000/min run the host used ≈ 30 % of one vCPU and PostgreSQL ≈ 13 %.
Limits
Section titled “Limits”- OHIP is the real ceiling. OHIP limits calls per application key (≈ 50 req/s documented); the governor defaults to 40 req/s per app key and queues (events, schedules) or answers 429 (HTTP) beyond it. One replica handles 20× that, so add replicas for availability, not for upstream throughput. The perf stack opens the governor to 5,000/s only because the stand-in has no limit.
- Concurrency per replica: 64 requests in flight, 128 queued (
Security:MaxConcurrentRequests,Security:MaxQueuedRequests); more is shed with 503overloaded. Raise both only together with the database pool (Maximum Pool Size, default 100) and PostgreSQLmax_connections(every replica holds up to its pool size plus two listener connections). - Request bodies: 1 MB globally (
Kestrel:Limits:MaxRequestBodySize) and per endpoint (functions, webhook, admin). - File descriptors: run the container with
nofile≥ 65536 (both compose files set it; Kubernetes runtimes usually default higher). With 1,024 an overloaded replica failed hard. - Outbox: a leased row blocks later rows of its partition (per-aggregate FIFO); throughput across partitions scales with
Events:Outbox:MaxParallelism× replicas. A future retry row holds its partition until due. - Streaming ingest: one consumer per (tenant, environment, chain) across all replicas (advisory lease); not load-tested.
- Metering and retention run in the background with batched statements (
docs/runbooks/backup-restore.mdfor retention windows).