Runbook: backup and restore
Everything durable is in PostgreSQL (stayfn database): tenants, environments (secret references only), hotels, variables and their history,
API key digests, invocations, events, dead letters, audit, usage. Secret values live in the secret store and are backed up there; export files
under Runtime:Export:Root are transient (retention below).
Symptoms
Section titled “Symptoms”- Planned: before an upgrade (the host migrates on start), before changing retention windows, on a schedule.
- Unplanned: a lost volume/instance, a destructive mistake (tenant deleted — soft delete, data retained — or rows removed by hand).
Diagnose
Section titled “Diagnose”- What is lost and since when:
GET /api/admin/audit?tenantId=<id>(who changed what), the last good backup’s timestamp. - Is the schema current?
/health/ready→schema. A restore of an older backup is migrated forward by the next start or the migration Job. - Roles: a restore needs the two roles of
docker/postgres/init-prod.sh(stayfn_migratorowner,stayfn_appDML) to exist first.
Remediate
Section titled “Remediate”Backup (compose; managed PostgreSQL: use its snapshots plus a logical dump for portability):
docker compose -f docker-compose.prod.yml exec -T postgres pg_dump -U postgres -d stayfn -Fc --no-owner --no-privileges > stayfn-$(date -u +%Y%m%dT%H%M%SZ).dumpRestore into an empty database (same or new server; the roles must exist):
docker compose -f docker-compose.prod.yml stop hostdocker compose -f docker-compose.prod.yml exec -T postgres psql -U postgres -c "DROP DATABASE IF EXISTS stayfn WITH (FORCE)" -c "CREATE DATABASE stayfn OWNER stayfn_migrator"docker compose -f docker-compose.prod.yml exec -T postgres psql -U postgres -d stayfn -c "GRANT CONNECT ON DATABASE stayfn TO stayfn_app" -c "REVOKE CREATE ON SCHEMA public FROM PUBLIC" -c "GRANT USAGE ON SCHEMA public TO stayfn_app" -c "ALTER DEFAULT PRIVILEGES FOR ROLE stayfn_migrator IN SCHEMA public GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO stayfn_app" -c "ALTER DEFAULT PRIVILEGES FOR ROLE stayfn_migrator IN SCHEMA public GRANT USAGE, SELECT ON SEQUENCES TO stayfn_app"docker compose -f docker-compose.prod.yml exec -T postgres pg_restore -U postgres -d stayfn --no-owner --role=stayfn_migrator < stayfn-<stamp>.dumpdocker compose -f docker-compose.prod.yml start host--no-owner --role=stayfn_migrator makes the migrator own every restored object, so the default privileges give stayfn_app its DML grants
and row-level security stays enabled and forced (the policies are part of the dump). Verify: /health/ready is 200 and
SELECT relname FROM pg_class WHERE relkind='r' AND relnamespace='public'::regnamespace AND NOT (relrowsecurity AND relforcerowsecurity) returns
only __EFMigrationsHistory and the host-level tables (functions, metering_watermarks).
Retention (T8.5): the retention job runs hourly on one replica (advisory lock) and deletes, per tenant under RLS and in batches
(Retention:BatchSize, 1,000), rows older than: completed invocations with their OHIP calls and idempotency keys Retention:Invocations (90 d),
terminal inbox events Retention:InboxEvents (90 d), done/dead outbox rows Retention:OutboxMessages (30 d), dead letters Retention:DeadLetters
(90 d), audit Retention:AuditLog (365 d), usage Retention:UsageHourly (400 d), schedule claims Retention:ScheduleRuns (30 d), export files
Retention:Exports (30 d). Windows are at least one day; Retention:Enabled=false stops it. Deleted rows are gone — keep backups for as
long as you must keep history (and no longer, for guest data). Counts: log line “Retention sweep over …” and stayfn.retention.deleted{kind}.
A dump of the running prod-compose database was taken, restored into a fresh PostgreSQL container with the roles from init-prod.sh, the row
counts matched and every tenant table still had RLS enabled and forced.