Skip to content

Runbook: backup and restore

Everything durable is in PostgreSQL (stayfn database): tenants, environments (secret references only), hotels, variables and their history, API key digests, invocations, events, dead letters, audit, usage. Secret values live in the secret store and are backed up there; export files under Runtime:Export:Root are transient (retention below).

  • Planned: before an upgrade (the host migrates on start), before changing retention windows, on a schedule.
  • Unplanned: a lost volume/instance, a destructive mistake (tenant deleted — soft delete, data retained — or rows removed by hand).
  1. What is lost and since when: GET /api/admin/audit?tenantId=<id> (who changed what), the last good backup’s timestamp.
  2. Is the schema current? /health/ready → schema. A restore of an older backup is migrated forward by the next start or the migration Job.
  3. Roles: a restore needs the two roles of docker/postgres/init-prod.sh (stayfn_migrator owner, stayfn_app DML) to exist first.

Backup (compose; managed PostgreSQL: use its snapshots plus a logical dump for portability):

Terminal window
docker compose -f docker-compose.prod.yml exec -T postgres pg_dump -U postgres -d stayfn -Fc --no-owner --no-privileges > stayfn-$(date -u +%Y%m%dT%H%M%SZ).dump

Restore into an empty database (same or new server; the roles must exist):

Terminal window
docker compose -f docker-compose.prod.yml stop host
docker compose -f docker-compose.prod.yml exec -T postgres psql -U postgres -c "DROP DATABASE IF EXISTS stayfn WITH (FORCE)" -c "CREATE DATABASE stayfn OWNER stayfn_migrator"
docker compose -f docker-compose.prod.yml exec -T postgres psql -U postgres -d stayfn -c "GRANT CONNECT ON DATABASE stayfn TO stayfn_app" -c "REVOKE CREATE ON SCHEMA public FROM PUBLIC" -c "GRANT USAGE ON SCHEMA public TO stayfn_app" -c "ALTER DEFAULT PRIVILEGES FOR ROLE stayfn_migrator IN SCHEMA public GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO stayfn_app" -c "ALTER DEFAULT PRIVILEGES FOR ROLE stayfn_migrator IN SCHEMA public GRANT USAGE, SELECT ON SEQUENCES TO stayfn_app"
docker compose -f docker-compose.prod.yml exec -T postgres pg_restore -U postgres -d stayfn --no-owner --role=stayfn_migrator < stayfn-<stamp>.dump
docker compose -f docker-compose.prod.yml start host

--no-owner --role=stayfn_migrator makes the migrator own every restored object, so the default privileges give stayfn_app its DML grants and row-level security stays enabled and forced (the policies are part of the dump). Verify: /health/ready is 200 and SELECT relname FROM pg_class WHERE relkind='r' AND relnamespace='public'::regnamespace AND NOT (relrowsecurity AND relforcerowsecurity) returns only __EFMigrationsHistory and the host-level tables (functions, metering_watermarks).

Retention (T8.5): the retention job runs hourly on one replica (advisory lock) and deletes, per tenant under RLS and in batches (Retention:BatchSize, 1,000), rows older than: completed invocations with their OHIP calls and idempotency keys Retention:Invocations (90 d), terminal inbox events Retention:InboxEvents (90 d), done/dead outbox rows Retention:OutboxMessages (30 d), dead letters Retention:DeadLetters (90 d), audit Retention:AuditLog (365 d), usage Retention:UsageHourly (400 d), schedule claims Retention:ScheduleRuns (30 d), export files Retention:Exports (30 d). Windows are at least one day; Retention:Enabled=false stops it. Deleted rows are gone — keep backups for as long as you must keep history (and no longer, for guest data). Counts: log line “Retention sweep over …” and stayfn.retention.deleted{kind}.

A dump of the running prod-compose database was taken, restored into a fresh PostgreSQL container with the roles from init-prod.sh, the row counts matched and every tenant table still had RLS enabled and forced.