Runbook: stream stalled
The streaming ingest (OHIP Streaming API, GraphQL over WebSocket) writes each event to the inbox and then its checkpoint; one consumer per
(tenant, environment, chain) across replicas. It is off by default (Events:Streaming:Enabled=false) and switched on per
deployment (docs/deployment.md “Streaming”; the sandbox dev loop uses the sandbox-streaming launch profile). Oracle allows one consumer per
chain and app key. Live protocol facts: the URL carries ?key=<sha256 hex of the app key>; an invalid token closes with
4401; a second consumer is refused with 4409; the gateway drops a socket without pings after ~60 s; a client complete is answered with
complete and a server close 1000.
Symptoms
Section titled “Symptoms”GET /api/admin/events/checkpointsshows a checkpoint whoseupdatedAtstops moving (streaming on).stayfn.events.outbox.lag_seconds(age of the oldest pending outbox row) orstayfn.events.outbox.pendinggrows;stayfn.events.stream.reconnectsclimbs;stayfn.events.ingested{outcome}flat./health/readyreportsstreamUnhealthy (“Streaming is enabled but the ingest loop is not running”).- Handlers not running although webhooks answer 202.
Diagnose
Section titled “Diagnose”curl $BASE/health/ready→streamsays disabled / running and how many streams this replica consumes; with several replicas only the lease holder consumes a given chain.GET /api/admin/events/checkpoints?tenantId=<id>→ checkpoint and lag per environment/chain;GET /api/admin/overview→ inboxreceived(not yet dispatched) and outboxpending/processing.- Logs:
grep -E 'Streaming|stream|4401|4403|4409|StreamingProtocolException'— token expiry (4401/4403) forces a token refetch; 4409 means another consumer holds the stream (another deployment or a probe with the same app key) and is retried afterConflictBackoff; a malformed event is dead-lettered as kindinboxand skipped. - Stop the host’s stream (or scale to zero), then run
stayfn ohip stream-probe --env <env> --offset-type highest --seconds 60: it prints the upgrade status, close codes, message types and event names only. Never run it while the host consumes the same chain and app key. - Outbox side: a partition blocked by a future retry row holds later rows of the same aggregate (by design) — check
stayfn.events.outbox.partition_depthand the oldest pending rows’partitionKeyinGET /api/admin/events?status=received. - Dispatcher alive?
Events:Outbox:Enabledmust be true on at least one replica; its poll log line “Outbox poll: leased …” appears at Debug.
Remediate
Section titled “Remediate”- Token/credential problem:
token-failures.md. - The ingest loop stopped: restart the pod/container (the lease moves to another replica within seconds; ingestion resumes from the checkpoint, duplicates are absorbed by the inbox unique key).
- A poison event: it is already dead-lettered and skipped; replay it after fixing the mapping (
dead-letter-storm.md). - Outbox backlog from an upstream outage: it drains by itself when OHIP recovers; raise
Events:Outbox:MaxParallelismor add replicas only if the lag keeps growing with a healthy upstream. - Streaming never connects (every upgrade a bare
400 Bad Request): the URL lacks?key=<sha256 hex of the app key>(a customWebSocketUrlthat strips it, or a proxy) — the probe with--no-key-hashreproduces the 400. - Connected but no events at all (the probe shows
connection_ackand pongs, nevernext, even from--offset 0): the application is not subscribed to any business events in the OHIP Developer Portal; nothing is enqueued until it is.
On the compose stack (streaming disabled): readiness reported stream healthy/disabled, the checkpoints endpoint was empty, a webhook event was
accepted and dispatched and the outbox metrics returned to zero lag. The live-stream part (probe, host against the sandbox, restart mid-stream) was exercised against the OHIP sandbox.