# Pivot Forecast — CLAUDE.md ## What this app is A web app for building named forecast scenarios against any PostgreSQL table. The workflow: load historical actuals as a baseline (optionally date-shifted into the forecast period), then apply incremental adjustments (scale, recode, clone) to build a plan. All changes are append-only, fully audited, and reversible by log entry. Full spec: `pf_spec.md` Data transport architecture options: `pf_perspective_options.md` --- ## Tech stack - **Backend:** Node.js / Express (`server.js`) - **Database:** PostgreSQL — isolated `pf` schema - **Frontend:** React + Vite + Tailwind CSS in `ui/`; built output lands in `public/app/` - **Pivot:** [Perspective](https://github.com/perspective-dev/perspective) (`@perspective-dev/*` distribution, **not** FINOS `@finos/perspective`) 5.2.0, **bundled inline via the `/inline` entrypoints — never from a CDN** (the 4.x CDN bundle resolves its server WASM to an unversioned path and silently pulls whatever is newest). See `PERSPECTIVE.md`. - **Dev:** `npm run dev` (nodemon) in root; `npm run build` in `ui/` --- ## Project layout ``` server.js Express entry point; pg pool; session; type parsers for bigint/numeric routes/ auth.js POST /api/login, /api/logout, GET /api/me; login throttle tables.js GET /api/tables, /api/tables/:schema/:tname/preview sources.js Source registration, col_meta, SQL generation versions.js Version CRUD, baseline/reference load, data stream operations.js scale, recode, clone, undo — the core forecast ops log.js GET /api/versions/:id/log, DELETE /api/log/:logid lib/ sql_generator.js buildFilterClause, token substitution helpers auth.js scrypt hash/verify, requireAuth, sessionUser; `node lib/auth.js hash` CLI utils.js setup_sql/ 01_schema.sql pf schema DDL — run once to install 02_auth.sql pf.app_user + pf.session ui/src/ observerShim.js Wraps window.Intersection/ResizeObserver; MUST be main.jsx's first import auth.jsx AuthProvider/useAuth; wraps fetch so any 401 returns to login views/ Login.jsx Sign-in form Setup.jsx DB browser, source registration, col_meta editor Baseline.jsx Version management, baseline workbench, reference load Forecast.jsx Perspective pivot, selection handling, operation dispatch components/ OperationPanel.jsx The adjustment workbench — ledger + scale/recode/clone forms BridgeView.jsx Baseline → current waterfall by tag (exports buildSteps/layoutSteps) Sidebar.jsx 3-step collapsible nav StatusBar.jsx Source · version · write target · row counts · theme Timeline.jsx Date-range preview bar for baseline segments ``` --- ## Database schema (`pf`) - **`pf.source`** — registered source tables - **`pf.col_meta`** — column roles: `dimension` | `value` | `units` | `date` | `filter` | `ignore`; `dim_group` groups functionally dependent columns (e.g. a part and its attributes, or a date and its derived year/month dimensions); `dim_period_col` maps a dimension to a `pf.dim_period` column so date-adjacent values are derived at load time rather than copied raw; `in_grain` flags dimension/date columns that define the **display grain** (see below); `is_key` is described under §`is_key` below - **`pf.version`** — named forecast scenarios; `exclude_iters` (default `["reference"]`) blocks those iter values from all operations - **`pf.fc_{tname}_{version_id}`** — one forecast table per version; contains both operational rows (`pf_iter = baseline|scale|recode|clone`) and reference rows (`pf_iter = reference`). Indexed on `pf_logid`, which is how undo, the change-log aggregate and the grain key all find their rows — without it each is a sequential scan of the whole table. Tables created before that index was added do not have it. - **`pf.log`** — audit log; every write gets one entry; `slice` + `params` stored as jsonb - **`pf.sql`** — generated SQL templates per source/operation; tokens substituted at request time - **`pf.app_user`** — login accounts; scrypt `pass_hash`, `is_active`, `last_login_at` - **`pf.session`** — express-session store (connect-pg-simple layout) - **`pf.dim_period`** — calendar lookup table (2018–2035); one row per month keyed on `sdat` (month start date); provides cal/fiscal year, quarter, and month columns; populated by `setup_sql/gen_dim_period.sql` with a configurable fiscal year start month ### `is_key` Read in four places, and **not** the one it sounds like — slices are validated against `filterCols`, which is every `dimension` plus every `date` column regardless of `is_key`. 1. **The key of a `dim_group`** — `resolveGroup()` in `routes/sources.js` takes `members.find(c => c.is_key)`, the column every other member is keyed on for `pf.dim_member` 2. **Value completion** — only `is_key` columns get a dropdown, and `/sources/:id/values/:col` refuses anything else 3. **Sibling autofill** — fires on blur only when `is_key && dim_group` 4. **The `dim_period` anchor** — `role === 'date' && is_key && dim_group` picks the date whose siblings are derived from the calendar Uses 2 and 3 want several columns flagged; use 1 needs exactly one per group. **When a group has more than one, `.find()` silently takes the lowest `opos`.** That is not hypothetical: `segment_new` (opos 12) outranked `part` (opos 15) in the `part` group, so a refresh keyed on a column that is null throughout, matched nothing, and reported success with zero members. `customer`'s group has four keys and picks the right one only by `opos` luck. The two meanings want separating — a per-group key choice, or a rule that the group key is the column named by the group (which these groups nearly follow already, except `sdate` → `sdate_e`). Until then, a refresh that finds two keys should refuse rather than guess. ### Key token substitution tokens `{{fc_table}}`, `{{where_clause}}`, `{{exclude_clause}}`, `{{logid}}`, `{{pf_user}}`, `{{value_incr}}`, `{{units_incr}}`, `{{pct}}`, `{{set_clause}}`, `{{scale_factor}}`, `{{date_offset}}`, `{{filter_clause}}` --- ## Core data flow ### Initial load (Forecast view) `Forecast.jsx` fetches col_meta first, then picks the endpoint: - **grain mode** (any `in_grain` column) — `GET /api/versions/:id/agg`, rows pre-aggregated to the grain, table indexed on `pf_gkey` - **raw mode** (no grain) — `GET /api/versions/:id/data`, raw forecast rows, table indexed on `pf_id` Either way: Arrow IPC binary stream → `worker.table(buffer)` in Perspective WASM. `fetchArrow()` handles both. **Why one batch (not streaming):** pg returns `bigint`/`numeric` as strings by default — type parsers in `server.js` coerce them to numbers. Per-batch Arrow encoding creates independent dictionaries that cause Perspective WASM to crash on dictionary replacement messages. Server accumulates all rows, emits one record batch. ### Display grain Aggregating to the grain the pivot actually displays is the load-time fix — measured 534,902 → 6,154 rows on `osm_stack`. It keeps the **native** Perspective engine, so expand/collapse/depth/sort/filter all still work. Set the grain in Setup (`in_grain` per column); it is baked into `pf.sql` at Generate SQL time so load and operations agree. `grainOf()` in `lib/sql_generator.js` is the single definition of what the grain is — `Setup.jsx` and `routes/log.js` mirror it. Full design: `pf_spec.md` → §Display-grain pre-aggregation. Why not a DuckDB virtual server: `pf_perspective_options.md` → §Spike findings. ### Segment and note columns `/data` and `/agg` both LEFT JOIN `pf.log` and emit two columns the forecast table does not itself carry: - **`pf_segment`** — for a baseline or reference row, that load's label (`tag`, else `note`); `'(adjustment)'` for everything else - **`pf_note`** — the free text on a scale/recode/clone; null on loads They are deliberately separate: commingling a segment name with an adjustment note makes neither pivotable. In grain mode `pf_logid` is part of the grain, so the join adds no rows. The operation routes stamp the same two fields onto the rows they push back incrementally, since those come from `RETURNING *` and would otherwise arrive without them. ### Forecast operations POST to `/api/versions/:id/{scale|recode|clone}` → SQL executed with `RETURNING *` → new rows returned as JSON → `pspTable.update(rows)` — no full reload. In grain mode the operation's final CTE aggregates its own new rows to grain first; since `pf_logid` is part of `pf_gkey` those keys are always new, so `update()` **appends** and the view re-sums. ### Price or volume (`plug`) A sales figure alone does not say which of price or volume moved, so scale takes `plug` — `'price'` (default, volume holds) or `'volume'` (price holds, units scale in proportion). Resolved in `resolveIncrs()` in `routes/operations.js`; the panel only offers it when the edit is dollars-only, because naming units or price has already answered it. The semantics come from the predecessor Excel model, `/opt/forecast_api/VBA/fpvt.frm` → `calc_val` / `calc_price`: ``` plug volume: pchange = fVal/(pVal+bVal); fVol = (pVol+bVol)*pchange plug price: fVol = pVol + bVol ``` A `target_price` with a `target_units` alongside is that form's Edit Price mode, where both are inputs and dollars fall out. The ledger's **Result** line previews value, units and price together using the same rules, so what you see is what gets written. ### Undo `DELETE /api/log/:logid` → removes rows by logid → `table.remove()` of the affected index values (`pf_gkeys` in grain mode, `pf_ids` in raw mode); the view re-sums. No full reload. --- ## Row depth and the observer shim `set_depth()` and per-node expand/collapse live on the **view**, not in `ViewConfig`. So every view rebuild starts fully expanded, and nothing in Perspective restores it — the app has to re-apply depth itself. The engine has no notion of focus or tab visibility. It re-renders because `IntersectionObserver` or `ResizeObserver` told it the element is visible again, and *that* is what discards the view. `ui/src/observerShim.js` wraps both constructors so the re-apply fires on the actual callback rather than on a window `focus` guess. **It must stay the first import in `main.jsx`.** Perspective's viewer captures the constructors at module-evaluation time (`var it=window.ResizeObserver; var st=window.IntersectionObserver`), so any import that reaches `perspective-viewer` first leaves the shim with nothing to intercept. `focus`/`visibilitychange`/`pageshow` remain as a backstop for when that happens. Two things learned the hard way, both easy to repeat: - Around a rebuild the viewer throws `No table set` from `getView()` while `getTable()` resolves happily. Gating on the table does not help; `applyDepthWhenReady()` retries `applyDepth` until it stops throwing. - `getView()` returns a **fresh wrapper object every call**, so object identity cannot be used to detect a rebuild. Two consecutive calls always look different. Tracing lives behind `localStorage.pf_debug = '1'`, prefixed `[pf-depth]` and `[pf-obs]`. **Not solved:** per-node `+/−` is still lost on rebuild. The view exposes `expand(row)` / `collapse(row)` with no getter, so the state cannot be read back — it would have to be shadowed by intercepting the grid's own calls. ## Slice mechanics When the user clicks a pivot cell, `perspective-click` fires. The handler in `Forecast.jsx` extracts `[col, '==', value]` filters from `detail.config.filter` — only `role = dimension` and `role = date` columns are kept as the slice. A plain click replaces the selection; ctrl/⌘/shift-click toggles a slice in or out of it, so the panel holds a **list** of slices sent as `slices` in operation POST bodies (the single `slice` object is still accepted server-side). Dragging across a block of cells selects a region. The datagrid runs in `edit_mode: SELECT_REGION` (forced on restore, so a saved layout can't switch it off) and reports the region as a `perspective-select` event carrying a Perspective **ViewWindow** — `{ start_row, end_row, start_col, end_col }`, *not* the per-row `insertConfigs` payload an older API used. It fires on every mouseover as the region grows, so the handler only records the latest window and a window-level `mouseup` commits it. A single-cell region is ignored there: `perspective-click` already owns plain and modifier clicks, and handling it in both places would undo a ctrl-click toggle. Turning a region back into slices re-derives, per cell, the same filters Perspective attaches to a click — row dimensions from the view's `__ROW_PATH__` (raw values, so dates stay epoch millis rather than whatever the grid formatted them as), column dimensions from the split_by segments of the column name. The grand-total row resolves to no dimension at all and is skipped; that would mean "the whole version". **Selection highlight.** The datagrid highlights whatever sits in its own `model._selection_state.selected_areas`, and wipes that list on every mousedown — so a multi-slice selection built up over several ctrl-clicks would only ever show the last cell. `Forecast.jsx` keeps `areasRef`, a `sliceKey -> rectangles` map parallel to `slices`, and an effect pushes the full set back and redraws after every change. Deselecting anywhere (ctrl-click, the panel's ×, Clear selection) prunes the map by live slice key, so the grid and the panel can't disagree. `pf_iter` is not a col_meta column, so it is stripped when a slice is built: two cells differing only by iter band produce the same effective slice. Duplicates are collapsed before the request — without that, `apply_mode: each` would apply the same change twice. **Limitation:** computed columns created by Perspective's split_by (e.g. Month, YearDate) don't map back to raw rows — only native dimension columns work for slice extraction. --- ## Column hierarchy (collapse / expand) The two pivot axes collapse by completely different mechanisms, and the asymmetry is a Perspective constraint, not a choice: - **Rows.** The `GROUP BY ROLLUP` view holds every level at once; `view.set_depth()` — which lives on the view, not the config — hides the deeper ones. That is what the `EXPAND 0 1 2 3` buttons drive, via `applyDepth()`. - **Columns.** There is no equivalent. `expand()` / `collapse()` take a **row index**, `ViewConfig` has `group_by_depth` but no `split_by_depth`, and `split_rollup_mode` (`'flat' | 'rollup'`) only chooses whether subtotal column groups are *emitted* — it is a view shape, not an interaction. So `applySplitDepth(n)` collapses by restoring a **truncated `split_by`**, which rebuilds the view. Three things follow from the rebuild, and each is handled: 1. The full hierarchy has to be remembered separately — once collapsed, `viewer.save()` only reports the short `split_by`. `splitFullRef` / `splitFull` hold it, and it is persisted into the saved layout as `split_full` so a reload while collapsed can still expand back. `adoptSplit()` is the single place it is set. 2. `perspective-config-update` fires for our own restore as well as the user rearranging the pivot. `collapsingRef` distinguishes them — without it, a collapse would overwrite the full hierarchy with the truncated one and the deeper levels would be unreachable. 3. Row depth lives on the discarded view, so `applyDepth(expandDepthRef.current)` is re-applied afterwards — the same wart as the refocus re-apply. The selection is cleared on every change: slices name the split_by dimensions they were cut from, and the highlight is keyed on grid coordinates. Neither survives a column axis that just changed shape. **Limitation:** this is whole-axis, not per-branch. Excel can collapse 2025 while 2026 stays expanded; truncating `split_by` collapses every column group at that level together. Per-branch is not reachable — `columns` selects which *measures* appear, not individual split combinations. --- ## Operation SQL patterns All three operations follow the same structure: insert a `pf.log` row in a CTE, then insert forecast rows referencing its id. `{{where_clause}}` is built from the slice; `{{exclude_clause}}` blocks `exclude_iters` rows. - **Scale** — distributes `value_incr`/`units_incr` proportionally across rows in the slice using window functions - **Recode** — inserts negative rows (zero out original) + positive rows with `{{set_clause}}` dimension overrides; both share the same logid - **Clone** — copies the slice with `{{set_clause}}` overrides and `{{scale_factor}}` multiplier; original untouched `build_where()` validates every slice key against col_meta (only `role = dimension` allowed). Values are escaped but not parameterized — consistent with existing patterns, debuggable in pg logs. --- ## Authentication Everything under `/api` except the auth routes sits behind a session; the React app is only mounted once there is one (`Gate` in `main.jsx`), because its load effects call the API immediately. - **Accounts:** `pf.app_user` — scrypt hashes from `lib/auth.js`, never plaintext. Managed with `./pf.sh add-user | passwd | list-users | disable-user | enable-user`; the password is read on stdin and hashed before it reaches psql. - **Sessions:** `express-session` + `connect-pg-simple` in `pf.session`, so a restart doesn't sign everyone out and a session can be revoked by deleting its row (`disable-user` does exactly that). Cookie `pf.sid`: httpOnly, SameSite=Lax, Secure unless `COOKIE_SECURE=false`, 12h rolling. - **Config:** `SESSION_SECRET` is required — the server exits at boot without one. `TRUST_PROXY` (default 1) makes `req.ip` and secure-cookie detection correct behind the TLS proxy. `CORS_ORIGIN` is the only way CORS is enabled at all; a wildcard origin plus a session cookie would be cross-site request forgery by construction. - **Login hardening:** `routes/auth.js` throttles to 10 failures per IP per 15 minutes (in-memory), returns one message for unknown/wrong/disabled alike, and regenerates the session id on success. **Identity is server-side.** `pf_user`, `created_by` and `closed_by` come from `sessionUser(req)`, never from the request body — the UI used to send a hardcoded `pf_user: 'admin'`, which any client could have set to anything. The audit log now names the account that made the change. ## Light / dark mode Theme state lives in `ui/src/theme.jsx` — a React context (`ThemeContext`) with a `ThemeProvider` that wraps the app in `main.jsx`. - **Storage key:** `pf_dark` in `localStorage`; falls back to `window.matchMedia('(prefers-color-scheme: dark)')` on first visit - **Toggle:** `setDark(d => !d)` in `StatusBar.jsx`; effect writes `localStorage` and toggles the `.dark` class on `` - **CSS:** `ui/src/index.css` defines CSS custom properties under `:root` (light) and `.dark`. All Tailwind color overrides are written as `.dark .bg-white { ... }` etc. — no Tailwind dark-mode config needed - **Palette:** dark mode uses Perspective's "Pro Dark" colours (`--bg-primary: #242526`, panels `#2a2c2f`, gridlines `#3b3f46`, text `#c5c9d0`) - **Perspective viewer:** `Forecast.jsx` calls `viewer.setAttribute('theme', dark ? 'Pro Dark' : 'Pro Light')` both on initial load and in a `useEffect([dark, versionId])` so the viewer stays in sync when the toggle fires - **Consuming the theme:** `import useTheme from '../theme.jsx'` then `const { dark, setDark } = useTheme()` ## Known issues / active work - **Load time is dominated by row count, not payload size.** On `fc_osm_skinny_29` (2.56M raw rows) a 24-column grain still yields 285,685 rows: ~2s to aggregate in pg, but ~15s to serialise those rows out of Postgres and parse them into JS, then ~3s to build Arrow. Halving the payload (the `pf_gkey` md5) barely moved it. The remaining lever is **dynamic grain** — group by the fields the current pivot actually uses rather than every `in_grain` column; see `pf_spec.md` → §Display-grain pre-aggregation, "the dynamic variant" - Per-node row expand/collapse is lost whenever the view rebuilds; see §Row depth and the observer shim - Default pivot layout should be configurable per source (currently hardcodes first 2 dimensions) - Source/version selection persists in `localStorage` (`pf_sourceId` / `pf_versionId`, `App.jsx`). It is re-validated against the live list whenever that list changes, so a deregistered source or deleted version re-points at the first remaining one instead of leaving a dead id that 404s every call - Col_meta / version schema drift: if col_meta roles change after a version's forecast table is created, SQL and DDL go out of sync — workaround is to delete and recreate the version - Grain drift: changing `in_grain` after a load requires Generate SQL + a page reload, since the loaded table's index and columns are fixed at load time. `routes/log.js` derives the grain from live col_meta, so a grain changed mid-session yields `pf_gkeys` that don't match the loaded table and undo silently removes nothing - Grain is static per source — a dimension left unflagged cannot be pivoted on. Dynamic per-cut grain (intersect the viewer's field set with the eligible set) is the additive next step; see `pf_spec.md` → §Display-grain pre-aggregation ## Deferred (not in v1) Baseline replay (`replay: true` returns 501), approval workflow, territory filtering, export, version comparison, multi-DB connections. Live server-side aggregation (Path A / DuckDB virtual server) is parked on branch `spike/duckdb-virtual-server`.