CLAUDE.md said is_key "marks dimensions used in slice WHERE clauses", which is not true and never was: assertSelective and buildWhere validate against filterCols, which is every dimension plus every date column regardless. What it actually drives is the key of a dim_group, value completion, the sibling autofill trigger, and the dim_period anchor -- and the first of those wants exactly one column per group while the others want several. When a group has more than one, find() silently takes the lowest opos, which is how segment_new outranked part and a master-data refresh keyed on a column that is null throughout: zero members, reported as success. Written down because the failure gives no signal at all -- the refresh succeeds, the list is simply empty, and every symptom appears somewhere else entirely. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
314 lines
21 KiB
Markdown
314 lines
21 KiB
Markdown
# Pivot Forecast — CLAUDE.md
|
||
|
||
## What this app is
|
||
|
||
A web app for building named forecast scenarios against any PostgreSQL table. The workflow: load historical actuals as a baseline (optionally date-shifted into the forecast period), then apply incremental adjustments (scale, recode, clone) to build a plan. All changes are append-only, fully audited, and reversible by log entry.
|
||
|
||
Full spec: `pf_spec.md`
|
||
Data transport architecture options: `pf_perspective_options.md`
|
||
|
||
---
|
||
|
||
## Tech stack
|
||
|
||
- **Backend:** Node.js / Express (`server.js`)
|
||
- **Database:** PostgreSQL — isolated `pf` schema
|
||
- **Frontend:** React + Vite + Tailwind CSS in `ui/`; built output lands in `public/app/`
|
||
- **Pivot:** [Perspective](https://github.com/perspective-dev/perspective) (`@perspective-dev/*` distribution, **not** FINOS `@finos/perspective`) 5.2.0, **bundled inline via the `/inline` entrypoints — never from a CDN** (the 4.x CDN bundle resolves its server WASM to an unversioned path and silently pulls whatever is newest). See `PERSPECTIVE.md`.
|
||
- **Dev:** `npm run dev` (nodemon) in root; `npm run build` in `ui/`
|
||
|
||
---
|
||
|
||
## Project layout
|
||
|
||
```
|
||
server.js Express entry point; pg pool; session; type parsers for bigint/numeric
|
||
routes/
|
||
auth.js POST /api/login, /api/logout, GET /api/me; login throttle
|
||
tables.js GET /api/tables, /api/tables/:schema/:tname/preview
|
||
sources.js Source registration, col_meta, SQL generation
|
||
versions.js Version CRUD, baseline/reference load, data stream
|
||
operations.js scale, recode, clone, undo — the core forecast ops
|
||
log.js GET /api/versions/:id/log, DELETE /api/log/:logid
|
||
lib/
|
||
sql_generator.js buildFilterClause, token substitution helpers
|
||
auth.js scrypt hash/verify, requireAuth, sessionUser; `node lib/auth.js hash` CLI
|
||
utils.js
|
||
setup_sql/
|
||
01_schema.sql pf schema DDL — run once to install
|
||
02_auth.sql pf.app_user + pf.session
|
||
ui/src/
|
||
observerShim.js Wraps window.Intersection/ResizeObserver; MUST be main.jsx's first import
|
||
auth.jsx AuthProvider/useAuth; wraps fetch so any 401 returns to login
|
||
views/
|
||
Login.jsx Sign-in form
|
||
Setup.jsx DB browser, source registration, col_meta editor
|
||
Baseline.jsx Version management, baseline workbench, reference load
|
||
Forecast.jsx Perspective pivot, selection handling, operation dispatch
|
||
components/
|
||
OperationPanel.jsx The adjustment workbench — ledger + scale/recode/clone forms
|
||
BridgeView.jsx Baseline → current waterfall by tag (exports buildSteps/layoutSteps)
|
||
Sidebar.jsx 3-step collapsible nav
|
||
StatusBar.jsx Source · version · write target · row counts · theme
|
||
Timeline.jsx Date-range preview bar for baseline segments
|
||
```
|
||
|
||
---
|
||
|
||
## Database schema (`pf`)
|
||
|
||
- **`pf.source`** — registered source tables
|
||
- **`pf.col_meta`** — column roles: `dimension` | `value` | `units` | `date` | `filter` | `ignore`; `dim_group` groups functionally dependent columns (e.g. a part and its attributes, or a date and its derived year/month dimensions); `dim_period_col` maps a dimension to a `pf.dim_period` column so date-adjacent values are derived at load time rather than copied raw; `in_grain` flags dimension/date columns that define the **display grain** (see below); `is_key` is described under §`is_key` below
|
||
- **`pf.version`** — named forecast scenarios; `exclude_iters` (default `["reference"]`) blocks those iter values from all operations
|
||
- **`pf.fc_{tname}_{version_id}`** — one forecast table per version; contains both operational rows (`pf_iter = baseline|scale|recode|clone`) and reference rows (`pf_iter = reference`). Indexed on `pf_logid`, which is how undo, the change-log aggregate and the grain key all find their rows — without it each is a sequential scan of the whole table. Tables created before that index was added do not have it.
|
||
- **`pf.log`** — audit log; every write gets one entry; `slice` + `params` stored as jsonb
|
||
- **`pf.sql`** — generated SQL templates per source/operation; tokens substituted at request time
|
||
- **`pf.app_user`** — login accounts; scrypt `pass_hash`, `is_active`, `last_login_at`
|
||
- **`pf.session`** — express-session store (connect-pg-simple layout)
|
||
- **`pf.dim_period`** — calendar lookup table (2018–2035); one row per month keyed on `sdat` (month start date); provides cal/fiscal year, quarter, and month columns; populated by `setup_sql/gen_dim_period.sql` with a configurable fiscal year start month
|
||
|
||
### `is_key`
|
||
|
||
Read in four places, and **not** the one it sounds like — slices are validated
|
||
against `filterCols`, which is every `dimension` plus every `date` column
|
||
regardless of `is_key`.
|
||
|
||
1. **The key of a `dim_group`** — `resolveGroup()` in `routes/sources.js` takes
|
||
`members.find(c => c.is_key)`, the column every other member is keyed on for
|
||
`pf.dim_member`
|
||
2. **Value completion** — only `is_key` columns get a dropdown, and
|
||
`/sources/:id/values/:col` refuses anything else
|
||
3. **Sibling autofill** — fires on blur only when `is_key && dim_group`
|
||
4. **The `dim_period` anchor** — `role === 'date' && is_key && dim_group` picks the
|
||
date whose siblings are derived from the calendar
|
||
|
||
Uses 2 and 3 want several columns flagged; use 1 needs exactly one per group.
|
||
**When a group has more than one, `.find()` silently takes the lowest `opos`.**
|
||
That is not hypothetical: `segment_new` (opos 12) outranked `part` (opos 15) in
|
||
the `part` group, so a refresh keyed on a column that is null throughout, matched
|
||
nothing, and reported success with zero members. `customer`'s group has four keys
|
||
and picks the right one only by `opos` luck.
|
||
|
||
The two meanings want separating — a per-group key choice, or a rule that the
|
||
group key is the column named by the group (which these groups nearly follow
|
||
already, except `sdate` → `sdate_e`). Until then, a refresh that finds two keys
|
||
should refuse rather than guess.
|
||
|
||
### Key token substitution tokens
|
||
`{{fc_table}}`, `{{where_clause}}`, `{{exclude_clause}}`, `{{logid}}`, `{{pf_user}}`, `{{value_incr}}`, `{{units_incr}}`, `{{pct}}`, `{{set_clause}}`, `{{scale_factor}}`, `{{date_offset}}`, `{{filter_clause}}`
|
||
|
||
---
|
||
|
||
## Core data flow
|
||
|
||
### Initial load (Forecast view)
|
||
`Forecast.jsx` fetches col_meta first, then picks the endpoint:
|
||
|
||
- **grain mode** (any `in_grain` column) — `GET /api/versions/:id/agg`, rows pre-aggregated to the grain, table indexed on `pf_gkey`
|
||
- **raw mode** (no grain) — `GET /api/versions/:id/data`, raw forecast rows, table indexed on `pf_id`
|
||
|
||
Either way: Arrow IPC binary stream → `worker.table(buffer)` in Perspective WASM. `fetchArrow()` handles both.
|
||
|
||
**Why one batch (not streaming):** pg returns `bigint`/`numeric` as strings by default — type parsers in `server.js` coerce them to numbers. Per-batch Arrow encoding creates independent dictionaries that cause Perspective WASM to crash on dictionary replacement messages. Server accumulates all rows, emits one record batch.
|
||
|
||
### Display grain
|
||
Aggregating to the grain the pivot actually displays is the load-time fix — measured 534,902 → 6,154 rows on `osm_stack`. It keeps the **native** Perspective engine, so expand/collapse/depth/sort/filter all still work. Set the grain in Setup (`in_grain` per column); it is baked into `pf.sql` at Generate SQL time so load and operations agree. `grainOf()` in `lib/sql_generator.js` is the single definition of what the grain is — `Setup.jsx` and `routes/log.js` mirror it. Full design: `pf_spec.md` → §Display-grain pre-aggregation. Why not a DuckDB virtual server: `pf_perspective_options.md` → §Spike findings.
|
||
|
||
### Segment and note columns
|
||
`/data` and `/agg` both LEFT JOIN `pf.log` and emit two columns the forecast table
|
||
does not itself carry:
|
||
|
||
- **`pf_segment`** — for a baseline or reference row, that load's label (`tag`, else
|
||
`note`); `'(adjustment)'` for everything else
|
||
- **`pf_note`** — the free text on a scale/recode/clone; null on loads
|
||
|
||
They are deliberately separate: commingling a segment name with an adjustment note
|
||
makes neither pivotable. In grain mode `pf_logid` is part of the grain, so the join
|
||
adds no rows. The operation routes stamp the same two fields onto the rows they push
|
||
back incrementally, since those come from `RETURNING *` and would otherwise arrive
|
||
without them.
|
||
|
||
### Forecast operations
|
||
POST to `/api/versions/:id/{scale|recode|clone}` → SQL executed with `RETURNING *` → new rows returned as JSON → `pspTable.update(rows)` — no full reload. In grain mode the operation's final CTE aggregates its own new rows to grain first; since `pf_logid` is part of `pf_gkey` those keys are always new, so `update()` **appends** and the view re-sums.
|
||
|
||
### Price or volume (`plug`)
|
||
A sales figure alone does not say which of price or volume moved, so scale takes
|
||
`plug` — `'price'` (default, volume holds) or `'volume'` (price holds, units scale
|
||
in proportion). Resolved in `resolveIncrs()` in `routes/operations.js`; the panel
|
||
only offers it when the edit is dollars-only, because naming units or price has
|
||
already answered it. The semantics come from the predecessor Excel model,
|
||
`/opt/forecast_api/VBA/fpvt.frm` → `calc_val` / `calc_price`:
|
||
|
||
```
|
||
plug volume: pchange = fVal/(pVal+bVal); fVol = (pVol+bVol)*pchange
|
||
plug price: fVol = pVol + bVol
|
||
```
|
||
|
||
A `target_price` with a `target_units` alongside is that form's Edit Price mode,
|
||
where both are inputs and dollars fall out. The ledger's **Result** line previews
|
||
value, units and price together using the same rules, so what you see is what
|
||
gets written.
|
||
|
||
### Undo
|
||
`DELETE /api/log/:logid` → removes rows by logid → `table.remove()` of the affected index values (`pf_gkeys` in grain mode, `pf_ids` in raw mode); the view re-sums. No full reload.
|
||
|
||
---
|
||
|
||
## Row depth and the observer shim
|
||
|
||
`set_depth()` and per-node expand/collapse live on the **view**, not in `ViewConfig`.
|
||
So every view rebuild starts fully expanded, and nothing in Perspective restores it —
|
||
the app has to re-apply depth itself.
|
||
|
||
The engine has no notion of focus or tab visibility. It re-renders because
|
||
`IntersectionObserver` or `ResizeObserver` told it the element is visible again, and
|
||
*that* is what discards the view. `ui/src/observerShim.js` wraps both constructors so
|
||
the re-apply fires on the actual callback rather than on a window `focus` guess.
|
||
|
||
**It must stay the first import in `main.jsx`.** Perspective's viewer captures the
|
||
constructors at module-evaluation time (`var it=window.ResizeObserver; var
|
||
st=window.IntersectionObserver`), so any import that reaches `perspective-viewer`
|
||
first leaves the shim with nothing to intercept. `focus`/`visibilitychange`/`pageshow`
|
||
remain as a backstop for when that happens.
|
||
|
||
Two things learned the hard way, both easy to repeat:
|
||
|
||
- Around a rebuild the viewer throws `No table set` from `getView()` while
|
||
`getTable()` resolves happily. Gating on the table does not help;
|
||
`applyDepthWhenReady()` retries `applyDepth` until it stops throwing.
|
||
- `getView()` returns a **fresh wrapper object every call**, so object identity
|
||
cannot be used to detect a rebuild. Two consecutive calls always look different.
|
||
|
||
Tracing lives behind `localStorage.pf_debug = '1'`, prefixed `[pf-depth]` and
|
||
`[pf-obs]`.
|
||
|
||
**Not solved:** per-node `+/−` is still lost on rebuild. The view exposes
|
||
`expand(row)` / `collapse(row)` with no getter, so the state cannot be read back —
|
||
it would have to be shadowed by intercepting the grid's own calls.
|
||
|
||
## Slice mechanics
|
||
|
||
When the user clicks a pivot cell, `perspective-click` fires. The handler in `Forecast.jsx` extracts `[col, '==', value]` filters from `detail.config.filter` — only `role = dimension` and `role = date` columns are kept as the slice. A plain click replaces the selection; ctrl/⌘/shift-click toggles a slice in or out of it, so the panel holds a **list** of slices sent as `slices` in operation POST bodies (the single `slice` object is still accepted server-side).
|
||
|
||
Dragging across a block of cells selects a region. The datagrid runs in `edit_mode: SELECT_REGION` (forced on restore, so a saved layout can't switch it off) and reports the region as a `perspective-select` event carrying a Perspective **ViewWindow** — `{ start_row, end_row, start_col, end_col }`, *not* the per-row `insertConfigs` payload an older API used. It fires on every mouseover as the region grows, so the handler only records the latest window and a window-level `mouseup` commits it. A single-cell region is ignored there: `perspective-click` already owns plain and modifier clicks, and handling it in both places would undo a ctrl-click toggle.
|
||
|
||
Turning a region back into slices re-derives, per cell, the same filters Perspective attaches to a click — row dimensions from the view's `__ROW_PATH__` (raw values, so dates stay epoch millis rather than whatever the grid formatted them as), column dimensions from the split_by segments of the column name. The grand-total row resolves to no dimension at all and is skipped; that would mean "the whole version".
|
||
|
||
**Selection highlight.** The datagrid highlights whatever sits in its own `model._selection_state.selected_areas`, and wipes that list on every mousedown — so a multi-slice selection built up over several ctrl-clicks would only ever show the last cell. `Forecast.jsx` keeps `areasRef`, a `sliceKey -> rectangles` map parallel to `slices`, and an effect pushes the full set back and redraws after every change. Deselecting anywhere (ctrl-click, the panel's ×, Clear selection) prunes the map by live slice key, so the grid and the panel can't disagree.
|
||
|
||
`pf_iter` is not a col_meta column, so it is stripped when a slice is built: two cells differing only by iter band produce the same effective slice. Duplicates are collapsed before the request — without that, `apply_mode: each` would apply the same change twice.
|
||
|
||
**Limitation:** computed columns created by Perspective's split_by (e.g. Month, YearDate) don't map back to raw rows — only native dimension columns work for slice extraction.
|
||
|
||
---
|
||
|
||
## Column hierarchy (collapse / expand)
|
||
|
||
The two pivot axes collapse by completely different mechanisms, and the asymmetry is a
|
||
Perspective constraint, not a choice:
|
||
|
||
- **Rows.** The `GROUP BY ROLLUP` view holds every level at once; `view.set_depth()` — which
|
||
lives on the view, not the config — hides the deeper ones. That is what the `EXPAND 0 1 2 3`
|
||
buttons drive, via `applyDepth()`.
|
||
- **Columns.** There is no equivalent. `expand()` / `collapse()` take a **row index**,
|
||
`ViewConfig` has `group_by_depth` but no `split_by_depth`, and `split_rollup_mode`
|
||
(`'flat' | 'rollup'`) only chooses whether subtotal column groups are *emitted* — it is a
|
||
view shape, not an interaction. So `applySplitDepth(n)` collapses by restoring a
|
||
**truncated `split_by`**, which rebuilds the view.
|
||
|
||
Three things follow from the rebuild, and each is handled:
|
||
|
||
1. The full hierarchy has to be remembered separately — once collapsed, `viewer.save()`
|
||
only reports the short `split_by`. `splitFullRef` / `splitFull` hold it, and it is
|
||
persisted into the saved layout as `split_full` so a reload while collapsed can still
|
||
expand back. `adoptSplit()` is the single place it is set.
|
||
2. `perspective-config-update` fires for our own restore as well as the user rearranging
|
||
the pivot. `collapsingRef` distinguishes them — without it, a collapse would overwrite
|
||
the full hierarchy with the truncated one and the deeper levels would be unreachable.
|
||
3. Row depth lives on the discarded view, so `applyDepth(expandDepthRef.current)` is
|
||
re-applied afterwards — the same wart as the refocus re-apply.
|
||
|
||
The selection is cleared on every change: slices name the split_by dimensions they were
|
||
cut from, and the highlight is keyed on grid coordinates. Neither survives a column axis
|
||
that just changed shape.
|
||
|
||
**Limitation:** this is whole-axis, not per-branch. Excel can collapse 2025 while 2026
|
||
stays expanded; truncating `split_by` collapses every column group at that level together.
|
||
Per-branch is not reachable — `columns` selects which *measures* appear, not individual
|
||
split combinations.
|
||
|
||
---
|
||
|
||
## Operation SQL patterns
|
||
|
||
All three operations follow the same structure: insert a `pf.log` row in a CTE, then insert forecast rows referencing its id. `{{where_clause}}` is built from the slice; `{{exclude_clause}}` blocks `exclude_iters` rows.
|
||
|
||
- **Scale** — distributes `value_incr`/`units_incr` proportionally across rows in the slice using window functions
|
||
- **Recode** — inserts negative rows (zero out original) + positive rows with `{{set_clause}}` dimension overrides; both share the same logid
|
||
- **Clone** — copies the slice with `{{set_clause}}` overrides and `{{scale_factor}}` multiplier; original untouched
|
||
|
||
`build_where()` validates every slice key against col_meta (only `role = dimension` allowed). Values are escaped but not parameterized — consistent with existing patterns, debuggable in pg logs.
|
||
|
||
---
|
||
|
||
## Authentication
|
||
|
||
Everything under `/api` except the auth routes sits behind a session; the React
|
||
app is only mounted once there is one (`Gate` in `main.jsx`), because its load
|
||
effects call the API immediately.
|
||
|
||
- **Accounts:** `pf.app_user` — scrypt hashes from `lib/auth.js`, never plaintext.
|
||
Managed with `./pf.sh add-user | passwd | list-users | disable-user | enable-user`;
|
||
the password is read on stdin and hashed before it reaches psql.
|
||
- **Sessions:** `express-session` + `connect-pg-simple` in `pf.session`, so a
|
||
restart doesn't sign everyone out and a session can be revoked by deleting its
|
||
row (`disable-user` does exactly that). Cookie `pf.sid`: httpOnly, SameSite=Lax,
|
||
Secure unless `COOKIE_SECURE=false`, 12h rolling.
|
||
- **Config:** `SESSION_SECRET` is required — the server exits at boot without one.
|
||
`TRUST_PROXY` (default 1) makes `req.ip` and secure-cookie detection correct
|
||
behind the TLS proxy. `CORS_ORIGIN` is the only way CORS is enabled at all; a
|
||
wildcard origin plus a session cookie would be cross-site request forgery by
|
||
construction.
|
||
- **Login hardening:** `routes/auth.js` throttles to 10 failures per IP per 15
|
||
minutes (in-memory), returns one message for unknown/wrong/disabled alike, and
|
||
regenerates the session id on success.
|
||
|
||
**Identity is server-side.** `pf_user`, `created_by` and `closed_by` come from
|
||
`sessionUser(req)`, never from the request body — the UI used to send a hardcoded
|
||
`pf_user: 'admin'`, which any client could have set to anything. The audit log
|
||
now names the account that made the change.
|
||
|
||
## Light / dark mode
|
||
|
||
Theme state lives in `ui/src/theme.jsx` — a React context (`ThemeContext`) with a `ThemeProvider` that wraps the app in `main.jsx`.
|
||
|
||
- **Storage key:** `pf_dark` in `localStorage`; falls back to `window.matchMedia('(prefers-color-scheme: dark)')` on first visit
|
||
- **Toggle:** `setDark(d => !d)` in `StatusBar.jsx`; effect writes `localStorage` and toggles the `.dark` class on `<html>`
|
||
- **CSS:** `ui/src/index.css` defines CSS custom properties under `:root` (light) and `.dark`. All Tailwind color overrides are written as `.dark .bg-white { ... }` etc. — no Tailwind dark-mode config needed
|
||
- **Palette:** dark mode uses Perspective's "Pro Dark" colours (`--bg-primary: #242526`, panels `#2a2c2f`, gridlines `#3b3f46`, text `#c5c9d0`)
|
||
- **Perspective viewer:** `Forecast.jsx` calls `viewer.setAttribute('theme', dark ? 'Pro Dark' : 'Pro Light')` both on initial load and in a `useEffect([dark, versionId])` so the viewer stays in sync when the toggle fires
|
||
- **Consuming the theme:** `import useTheme from '../theme.jsx'` then `const { dark, setDark } = useTheme()`
|
||
|
||
## Known issues / active work
|
||
|
||
- **Load time is dominated by row count, not payload size.** On `fc_osm_skinny_29`
|
||
(2.56M raw rows) a 24-column grain still yields 285,685 rows: ~2s to aggregate in
|
||
pg, but ~15s to serialise those rows out of Postgres and parse them into JS, then
|
||
~3s to build Arrow. Halving the payload (the `pf_gkey` md5) barely moved it. The
|
||
remaining lever is **dynamic grain** — group by the fields the current pivot
|
||
actually uses rather than every `in_grain` column; see `pf_spec.md` →
|
||
§Display-grain pre-aggregation, "the dynamic variant"
|
||
- Per-node row expand/collapse is lost whenever the view rebuilds; see §Row depth and
|
||
the observer shim
|
||
- Default pivot layout should be configurable per source (currently hardcodes first 2 dimensions)
|
||
- Source/version selection persists in `localStorage` (`pf_sourceId` / `pf_versionId`,
|
||
`App.jsx`). It is re-validated against the live list whenever that list changes, so a
|
||
deregistered source or deleted version re-points at the first remaining one instead of
|
||
leaving a dead id that 404s every call
|
||
- Col_meta / version schema drift: if col_meta roles change after a version's forecast table is created, SQL and DDL go out of sync — workaround is to delete and recreate the version
|
||
- Grain drift: changing `in_grain` after a load requires Generate SQL + a page reload, since the loaded table's index and columns are fixed at load time. `routes/log.js` derives the grain from live col_meta, so a grain changed mid-session yields `pf_gkeys` that don't match the loaded table and undo silently removes nothing
|
||
- Grain is static per source — a dimension left unflagged cannot be pivoted on. Dynamic per-cut grain (intersect the viewer's field set with the eligible set) is the additive next step; see `pf_spec.md` → §Display-grain pre-aggregation
|
||
|
||
## Deferred (not in v1)
|
||
Baseline replay (`replay: true` returns 501), approval workflow, territory filtering, export, version comparison, multi-DB connections. Live server-side aggregation (Path A / DuckDB virtual server) is parked on branch `spike/duckdb-virtual-server`.
|