pf_app/setup_sql/README.md
Paul Trowbridge d1197df7d5 Replace idempotent schema script with tracked forward-only migrations
setup_sql/01_schema.sql was an idempotent bootstrap script — CREATE TABLE
IF NOT EXISTS plus a tail of ALTER ... ADD COLUMN IF NOT EXISTS. It had no
record of what any given database had applied, which is how a branch could
declare col_meta.in_grain while the running database lacked it, with
nothing able to detect the mismatch. The symptom would have been a
confusing 'column "in_grain" does not exist' inside an unrelated request.

Migrations-only, no hand-maintained current-state file to drift:

- setup_sql/migrations/*.sql applied in filename order, recorded in
  pf.schema_version with a checksum. Split along the schema's actual
  evolution, so each column is declared exactly once — 01_schema.sql had
  grown to declare dim_group, dim_period_col and in_grain twice each.
- lib/migrations.js holds the bookkeeping, shared by the CLI and the boot
  check. scripts/migrate.js provides up | status | baseline.
- server.js refuses to start when the database is behind, listing what is
  pending. This converts silent drift into a clear boot message, which was
  the whole point. PF_SKIP_MIGRATION_CHECK=1 bypasses.
- Four integrity guards, each verified to fire: a migration modified after
  being applied, one recorded as applied but missing from disk, one that
  would apply out of order, and a re-run when already current.
- No IF NOT EXISTS on new migrations. The bookkeeping already guarantees
  one run each, and the guards hide ordering mistakes — that is exactly why
  01_schema.sql had ALTERs sitting above the CREATE TABLE they depended on,
  broken for anyone installing from scratch. 0004 keeps the guard only
  because it was applied by hand before migrations existed.

Verified: replaying all four migrations into a throwaway schema reproduces
the live pf schema exactly, 41 columns, column for column.

schema.generated.sql is a pg_dump snapshot for reading, refreshed by
npm run schema:dump. It excludes the runtime fc_* tables and strips
pg_dump's random \restrict token and version banner, so regenerating an
unchanged schema yields an identical file rather than a spurious diff.

pf.dim_period stays out of migrations — it is a parameterised data load
(fiscal year start month), not a schema change.

The dev database (ubm) has been baselined at all four migrations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 09:07:39 -04:00

2.4 KiB

Database setup

Migrations

migrations/*.sql are applied in filename order and recorded in pf.schema_version. They are the source of truth for the pf schema — there is no hand-maintained current-state file to drift out of sync.

npm run migrate            # apply pending migrations
npm run migrate:status     # what is applied, what is pending
npm run migrate:baseline   # record pending as applied WITHOUT running them

server.js refuses to start when the database is behind, so drift surfaces at boot rather than as column "x" does not exist inside an unrelated request. Set PF_SKIP_MIGRATION_CHECK=1 to bypass.

Fresh database

npm run migrate
psql -d <db> -f setup_sql/gen_dim_period.sql

Existing database that already matches

Use baseline so the runner does not try to re-create tables that exist:

npm run migrate:baseline

To baseline only part of the way — the schema matches through 0003 but not 0004 — pass --up-to and then migrate the rest:

node scripts/migrate.js baseline --up-to=0003_col_meta_dim_group_period.sql
npm run migrate

Writing a migration

  • Name it NNNN_short_description.sql, numbered after the highest existing file.
  • One concern per file. Keep it forward-only; there are no down migrations.
  • Applied migrations are immutable. The runner stores a checksum and refuses to proceed if a file changes after being applied, because editing one means databases silently disagree about what the schema is. To fix a mistake, add a new migration.
  • No IF NOT EXISTS guards on new migrations. The bookkeeping already guarantees each runs once, and the guards hide ordering mistakes — the reason the old 01_schema.sql had ALTERs sitting above the CREATE TABLE they depended on, broken for anyone installing from scratch. 0004 is the one exception, since it was applied by hand before migrations existed.

Not migrations

  • gen_dim_period.sql — creates and populates pf.dim_period. It is a parameterised data load (configurable fiscal year start month), not a schema change, so it stays a script you run deliberately.
  • pf.fc_{tname}_{version_id} — per-version forecast tables, created and dropped at runtime by routes/versions.js from col_meta. Never migrated.
  • schema.generated.sql — a pg_dump snapshot for reading, refreshed with npm run schema:dump. Generated, never edited, never applied.