Skip to content

Runbook: stream stalled

The streaming ingest (OHIP Streaming API, GraphQL over WebSocket) writes each event to the inbox and then its checkpoint; one consumer per (tenant, environment, chain) across replicas. It is off by default (Events:Streaming:Enabled=false) and switched on per deployment (docs/deployment.md “Streaming”; the sandbox dev loop uses the sandbox-streaming launch profile). Oracle allows one consumer per chain and app key. Live protocol facts: the URL carries ?key=<sha256 hex of the app key>; an invalid token closes with 4401; a second consumer is refused with 4409; the gateway drops a socket without pings after ~60 s; a client complete is answered with complete and a server close 1000.

  • GET /api/admin/events/checkpoints shows a checkpoint whose updatedAt stops moving (streaming on).
  • stayfn.events.outbox.lag_seconds (age of the oldest pending outbox row) or stayfn.events.outbox.pending grows; stayfn.events.stream.reconnects climbs; stayfn.events.ingested{outcome} flat.
  • /health/ready reports stream Unhealthy (“Streaming is enabled but the ingest loop is not running”).
  • Handlers not running although webhooks answer 202.
  1. curl $BASE/health/ready → stream says disabled / running and how many streams this replica consumes; with several replicas only the lease holder consumes a given chain.
  2. GET /api/admin/events/checkpoints?tenantId=<id> → checkpoint and lag per environment/chain; GET /api/admin/overview → inbox received (not yet dispatched) and outbox pending/processing.
  3. Logs: grep -E 'Streaming|stream|4401|4403|4409|StreamingProtocolException' — token expiry (4401/4403) forces a token refetch; 4409 means another consumer holds the stream (another deployment or a probe with the same app key) and is retried after ConflictBackoff; a malformed event is dead-lettered as kind inbox and skipped.
  4. Stop the host’s stream (or scale to zero), then run stayfn ohip stream-probe --env <env> --offset-type highest --seconds 60: it prints the upgrade status, close codes, message types and event names only. Never run it while the host consumes the same chain and app key.
  5. Outbox side: a partition blocked by a future retry row holds later rows of the same aggregate (by design) — check stayfn.events.outbox.partition_depth and the oldest pending rows’ partitionKey in GET /api/admin/events?status=received.
  6. Dispatcher alive? Events:Outbox:Enabled must be true on at least one replica; its poll log line “Outbox poll: leased …” appears at Debug.
  • Token/credential problem: token-failures.md.
  • The ingest loop stopped: restart the pod/container (the lease moves to another replica within seconds; ingestion resumes from the checkpoint, duplicates are absorbed by the inbox unique key).
  • A poison event: it is already dead-lettered and skipped; replay it after fixing the mapping (dead-letter-storm.md).
  • Outbox backlog from an upstream outage: it drains by itself when OHIP recovers; raise Events:Outbox:MaxParallelism or add replicas only if the lag keeps growing with a healthy upstream.
  • Streaming never connects (every upgrade a bare 400 Bad Request): the URL lacks ?key=<sha256 hex of the app key> (a custom WebSocketUrl that strips it, or a proxy) — the probe with --no-key-hash reproduces the 400.
  • Connected but no events at all (the probe shows connection_ack and pongs, never next, even from --offset 0): the application is not subscribed to any business events in the OHIP Developer Portal; nothing is enqueued until it is.

On the compose stack (streaming disabled): readiness reported stream healthy/disabled, the checkpoints endpoint was empty, a webhook event was accepted and dispatched and the outbox metrics returned to zero lag. The live-stream part (probe, host against the sandbox, restart mid-stream) was exercised against the OHIP sandbox.